kandinsky6_sr ¶
Stages of the Kandinsky6 video super-resolution (SR) pipeline.
video encoding -> latent preparation -> denoising -> decoding
The source clip is encoded once. Its latent is cut into overlapping tiles whose upscaled size is a resolution the SR DiT was trained on (basic/kandinsky6_sr/tiling.py); every tile is upscaled by the latent upscaler, noised, denoised with the bundle's scheduler (flow-matching Euler or the distilled pi-Flow) and decoded, and the decoded tiles are blended into the output video. The denoising loop runs per tile, which is why these stages replace the generic latent-preparation / denoising / decoding stages.
Classes¶
fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRDecodingStage ¶
Bases: PipelineStage
KVAE-decode every tile, blend the tiles and apply the optional delivery resize.
Tiles stay float through blending and are quantized to uint8 once, inside stitch_tiles -- matching the diffusers reference's decode_latents/__call__ (it accumulates video_acc += tile * window in float and only rounds the final video_acc / weight_acc), not the lossier quantize-per-tile-then-blend-already-rounded-values approach. batch.output stores code v as (v + 0.5) / 255 so the truncating * 255 -> uint8 conversion in VideoGenerator gives back exactly v.
Source code in fastvideo/pipelines/stages/kandinsky6_sr.py
fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRDenoisingStage ¶
Bases: PipelineStage
Denoise every tile with the bundle's scheduler; batch.latents becomes [num_tiles, T', H, W, C].
FlowMatchEulerDiscreteScheduler bundles run num_inference_steps Euler steps over the shifted linspace(1, 0) grid; PiflowScheduler bundles integrate the distilled policy (the DiT head then holds n_grid predictions per channel). The latent state stays fp32 across steps. Only the first C channels are denoised; the anchor channels are fixed conditioning.
num_inference_steps is the number of DiT calls per tile for both schedulers, as elsewhere in FastVideo. The upstream Diffusers pipeline counts timestep grid points instead, so its default 5 is 4 steps here.
Source code in fastvideo/pipelines/stages/kandinsky6_sr.py
fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRLatentPreparationStage ¶
Bases: PipelineStage
Upscale every latent tile and build the SR DiT input [noised upscaled latent | anchor | anchor mask].
The released checkpoints are trained with an HR anchor channel group; SR of a real low-quality clip has no anchor, so it is zero with a zero mask. The starting point is the upscaled latent mixed with Gaussian noise (variance-preserving, lq_noise_scale). Noise is drawn per group of sr_tiles_batch_size tiles from a generator seeded with seed + first tile index, so results do not depend on how many tile groups ran before. Output: batch.latents [num_tiles, T', H, W, 2C + 1] fp32.
Source code in fastvideo/pipelines/stages/kandinsky6_sr.py
fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRVideoEncodingStage ¶
Kandinsky6SRVideoEncodingStage(vae: Any)
Bases: PipelineStage
Read the source video (and its audio), apply the x2.25 pre-upscale and KVAE-encode it into lq_latents.