Kandinsky 6 — ComfyUI INT8 ConvRot
English | 简体中文
Three single-file INT8 ConvRot DiTs for joint video/audio generation with ComfyUI Kandinsky T8. The complete video, audio and internal text branches remain in each DiT. Shared text encoders and decoders are separate, unmodified files.
Downloads
Place the selected model in ComfyUI/models/diffusion_models:
| Model | Download | Size | Workflow |
|---|---|---|---|
| Lite | k6_lite_int8_convrot.safetensors | 4.87 GB | Base: 50 steps, CFG 5 |
| Lite distill | k6_lite_distill_int8_convrot.safetensors | 3.77 GB | PiFlow: 10 steps, CFG 1 |
| Pro distill | k6_pro_distill_int8_convrot.safetensors | 34.97 GB | PiFlow: 10 steps, CFG 1 |
For example, download Lite distill and the independent components directly into your ComfyUI models directory:
hf download t8star/Kandinsky-Comfy diffusion_models/k6_lite_distill_int8_convrot.safetensors \
text_encoders/qwen_2.5_vl_7b.safetensors text_encoders/clip_l.safetensors \
vae/hunyuan_video_vae_bf16.safetensors audio_vae/v1-44.pth \
audio_vae/bigvgan_vocoder/config.json audio_vae/bigvgan_vocoder/bigvgan_generator.pt \
--local-dir ComfyUI/models
The repository folders match ComfyUI's model folders. Keep audio_vae/bigvgan_vocoder/config.json beside bigvgan_generator.pt. Import a workflow, select the model in UNETLoader, and use weight_dtype=default.
PiFlow requires CFG=1, denoise=1 and one global, full-range conditioning; masks are limited to the Image workflow's clean reference tail.
Download the six ready GUI workflows, unzip and drag a root-level JSON onto the ComfyUI canvas. Each selects the corresponding model. Use Kandinsky T8 0.1.4+ and restart ComfyUI; the image example is installed as ComfyUI/input/kandinsky6_i2va_portrait.png. The pack includes this image and actual run verification records.
Requirements and validation
Requires ComfyUI 0.39.0+, comfy-kitchen 0.2.37+, the T8 nodes and an NVIDIA CUDA GPU with INT8 Tensor Core support. Default output is 864×480, 121 frames, 24 fps, with audio. Pro distill was tested using dynamic VRAM/offloading on RTX 5090 Laptop 24 GB / 64 GB RAM.
All three variants completed full text and image workflows; all six outputs passed audio/video decoding checks. See validation. Results cover the listed hardware and example prompts.
ConvRot uses normalized grouped Hadamard rotation with per-output-channel INT8 weights and dynamically quantized activations. Lite has 920 quantized Linear layers; Pro distill has 1728. Sensitive projections, modulation, norms, biases and embeddings retain their original tensors. SHA256SUMS, conversion manifests and MODEL_INDEX.json record checksums and pinned sources.
Credits and licenses
The converted Kandinsky DiTs are derived from Kandinsky Lab and retain the MIT license. INT8 conversion by T8star. Repository-level MIT metadata applies to the Kandinsky DiTs; independent components retain their own licenses:
| Component | Source | License |
|---|---|---|
| Qwen 2.5 VL 7B | Comfy-Org / Qwen | Apache 2.0 |
| CLIP-L | Comfy-Org / OpenAI CLIP | MIT |
| HunyuanVideo VAE | Comfy-Org / Tencent | Tencent Hunyuan Community |
TOD audio decoder v1-44.pth |
MMAudio / Sony Research | CC BY-NC 4.0 |
| BigVGAN 44 kHz | NVIDIA | MIT |
Shared components are redistributed unchanged. License texts and references are retained in licenses and NOTICE. TOD weights are licensed for noncommercial use; HunyuanVideo VAE retains its upstream territory and usage restrictions.
T8star
Bilibili · YouTube · API · Free gallery
Model tree for t8star/Kandinsky-Comfy
Base model
kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers