Kandinsky 6 Video Super-Resolution¶
Kandinsky6SRPipeline upscales a low-resolution video by x2, x2.25 or x4 (video-to-video, no text prompt). The clip is encoded once with the SR VAE (KVAE); its latent is cut into overlapping tiles, each tile is enlarged by the latent upscaler, refined by a text-free SR DiT in a few denoising steps and decoded, and the tiles are blended back together. The source audio is kept. The clip can come from anywhere, for example from the Kandinsky 6 T2IVA pipeline.
Models¶
| Variant | Hub repo | Scheduler | Default steps per tile |
|---|---|---|---|
| Flow matching | kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers | FlowMatchEulerDiscreteScheduler (shift 5.0) | 4 |
| Distilled | kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers | PiflowScheduler (shift 3.5, n_grid 10) | 2 |
Both repos have the same layout and share vae/ and latent_upscaler/:
model_index.json _class_name = Kandinsky6SRPipeline
transformer/ Kandinsky6SRTransformer3DModel (sr_params: trained base resolution, RoPE scale, noise level)
vae/ Kandinsky6SRVAE
latent_upscaler/ Kandinsky6SRLatentUpscalerBank (x2 and x4 models)
scheduler/ FlowMatchEulerDiscreteScheduler or PiflowScheduler
The scheduler component drives the denoising loop, and each repo resolves to its own preset (4 or 2 steps). The distilled transformer's head holds n_grid predictions per latent channel, which PiflowScheduler integrates.
Usage¶
python examples/inference/basic/basic_kandinsky6_sr.py --video-path input.mp4 --scale 2.25
python examples/inference/basic/basic_kandinsky6_sr.py --video-path input.mp4 \
--model-path kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
from fastvideo import VideoGenerator
generator = VideoGenerator.from_pretrained("kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers", num_gpus=1)
result = generator.generate({
"inputs": {"video_path": "input.mp4"},
"output": {"output_path": "outputs_video/sr", "return_frames": False},
"extensions": {"sr_resolution_scale": 2.25, "sr_target_resolution": "fullhd"},
})
The output geometry and frame rate follow the input clip; the request's height / width / num_frames are not used. Clips are resampled to 24 fps by a fixed stride and only the first 121 frames (5 s, 1 + 8k-aligned) are processed. The source audio (mono, 44.1 kHz) is trimmed to the processed span and muxed into the output.
Request parameters¶
The sr_* options belong only to Kandinsky6 SR. Pass them in request.extensions (as above), or under request.stage_overrides.sr. The SR pipeline reads them from ForwardBatch.extra; they are not fields of the shared SamplingParam or ForwardBatch. Other model families reject these options. Existing generate_video(..., sr_resolution_scale=...) calls remain supported for Kandinsky6 SR; replace direct SamplingParam(sr_...=...) construction with request extensions or these keyword arguments. For the config-based CLI, use dotted overrides such as --request.extensions.sr_resolution_scale 4, rather than shared --sr-* flags. The model-specific example above keeps its --scale and --tiles-batch-size flags.
| Field | Default | Meaning |
|---|---|---|
num_inference_steps | 4 / 2 | Denoising steps (DiT calls) per tile. The upstream Diffusers pipeline counts grid points instead (its 5 is 4 steps here). The distilled model was trained for 2; other values run with a warning. |
sr_resolution_scale | 2.25 | Total upscale: 2, 4 or 2.25 (x1.125 pixel pre-upscale, then x2). |
sr_tiles_batch_size | 1 | Tiles denoised per DiT call (raise it only if memory allows). |
sr_tile_min_overlap | 0.20 | Minimum overlap between neighbouring tiles, as a fraction of the tile. |
sr_target_resolution | None | Downscale the result to hd, fullhd, 2k or WxH. |
sr_target_resize_mode | fit | fit keeps the aspect ratio, exact uses the bucket dimensions. |
seed | 42 | Tile group k (of sr_tiles_batch_size tiles) is seeded with seed + index of its first tile. |
Two Python-only inputs take raw tensors and are passed as keyword arguments of generate_video():
sr_lr_latent: an unscaled KVAE latent[T, C, H, W]of the source, instead ofvideo_path(skips decoding and encoding the video; scale 2 or 4 only, since 2.25 needs the pixel pre-upscale).sr_audio/sr_audio_sample_rate: a mono waveform in[-1, 1]to mux instead of the source's own audio.
Limitations¶
- NABLA block-sparse attention, which both repos request for 512-pixel tiles, is not wired; the DiT runs dense attention and logs a warning.
- No sequence or tensor parallelism: the DiT runs on one GPU. Tiles are processed one group after another.
- One clip per request.