Kandinsky 6 recipes¶
Maintained family · Inference
Kandinsky 6 inference recipes
Kandinsky 6 from the Kandinsky Lab generates five-second video with synchronized audio from text or an image, with base and distilled pi-Flow checkpoints, and upscales existing clips with base or distilled video super-resolution.
Pick a recipe and runtime
Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.
Loading recipe details...
Exact device and memory details appear only when a recorded run supports them.
Loading...
- Model
- Loading...
- Workload
- Loading...
- Source configuration
- Loading...
- Expected output
- Loading...
Loading... Setup
The generated commands expect a local clone:
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo Use Configuration for supported Python and CLI settings, Optimizations for attention and memory tradeoffs, and the support matrix for the supported model and optimization surface.
For model-specific behavior, see Kandinsky 6 video with audio and Kandinsky 6 video super-resolution.
Troubleshooting
- Text and image conditioning use the same TI2VA pipeline and example. Set
IMAGE_PATHin the script to condition on an image; both base and distilled outputs include synchronized audio. - For pi-Flow, set
KANDINSKY6_MODEL_PATHto the distilled model ID when runningbasic_kandinsky6_ti2va.py. Its registered preset uses 10 inference steps, guidance 1.0,eps=1e-6,final_step_size_scale=0.5, andnum_policy_substeps=128. - For VSR, set
INPUT_VIDEOin the generated command. The pipeline accepts x2, x2.25 and x4 scales, processes up to 121 frames at 24 fps, and preserves source audio. - Checkpoint key-layout errors indicate that the checkpoint and FastVideo checkout target different Kandinsky 6 Diffusers revisions.
- Gated or missing checkpoints: run
huggingface-cli loginand confirm you accepted the model's license on Hugging Face.
Evidence status
All recipes on this page are Source-backed: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are Unknown and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.