Skip to content

kandinsky6_sr

Stages of the Kandinsky6 video super-resolution (SR) pipeline.

video encoding -> latent preparation -> denoising -> decoding

The source clip is encoded once. Its latent is cut into overlapping tiles whose upscaled size is a resolution the SR DiT was trained on (basic/kandinsky6_sr/tiling.py); every tile is upscaled by the latent upscaler, noised, denoised with the bundle's scheduler (flow-matching Euler or the distilled pi-Flow) and decoded, and the decoded tiles are blended into the output video. The denoising loop runs per tile, which is why these stages replace the generic latent-preparation / denoising / decoding stages.

Classes

fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRDecodingStage

Kandinsky6SRDecodingStage(vae: Any, transformer: Any)

Bases: PipelineStage

KVAE-decode every tile, blend the tiles and apply the optional delivery resize.

Tiles stay float through blending and are quantized to uint8 once, inside stitch_tiles -- matching the diffusers reference's decode_latents/__call__ (it accumulates video_acc += tile * window in float and only rounds the final video_acc / weight_acc), not the lossier quantize-per-tile-then-blend-already-rounded-values approach. batch.output stores code v as (v + 0.5) / 255 so the truncating * 255 -> uint8 conversion in VideoGenerator gives back exactly v.

Source code in fastvideo/pipelines/stages/kandinsky6_sr.py
def __init__(self, vae: Any, transformer: Any) -> None:
    super().__init__()
    self.vae = vae
    self.transformer = transformer

fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRDenoisingStage

Kandinsky6SRDenoisingStage(transformer: Any, scheduler: Any)

Bases: PipelineStage

Denoise every tile with the bundle's scheduler; batch.latents becomes [num_tiles, T', H, W, C].

FlowMatchEulerDiscreteScheduler bundles run num_inference_steps Euler steps over the shifted linspace(1, 0) grid; PiflowScheduler bundles integrate the distilled policy (the DiT head then holds n_grid predictions per channel). The latent state stays fp32 across steps. Only the first C channels are denoised; the anchor channels are fixed conditioning.

num_inference_steps is the number of DiT calls per tile for both schedulers, as elsewhere in FastVideo. The upstream Diffusers pipeline counts timestep grid points instead, so its default 5 is 4 steps here.

Source code in fastvideo/pipelines/stages/kandinsky6_sr.py
def __init__(self, transformer: Any, scheduler: Any) -> None:
    super().__init__()
    self.transformer = transformer
    self.scheduler = scheduler

fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRLatentPreparationStage

Kandinsky6SRLatentPreparationStage(latent_upscaler: Any, vae: Any, transformer: Any)

Bases: PipelineStage

Upscale every latent tile and build the SR DiT input [noised upscaled latent | anchor | anchor mask].

The released checkpoints are trained with an HR anchor channel group; SR of a real low-quality clip has no anchor, so it is zero with a zero mask. The starting point is the upscaled latent mixed with Gaussian noise (variance-preserving, lq_noise_scale). Noise is drawn per group of sr_tiles_batch_size tiles from a generator seeded with seed + first tile index, so results do not depend on how many tile groups ran before. Output: batch.latents [num_tiles, T', H, W, 2C + 1] fp32.

Source code in fastvideo/pipelines/stages/kandinsky6_sr.py
def __init__(self, latent_upscaler: Any, vae: Any, transformer: Any) -> None:
    super().__init__()
    self.latent_upscaler = latent_upscaler
    self.vae = vae
    self.transformer = transformer

fastvideo.pipelines.stages.kandinsky6_sr.Kandinsky6SRVideoEncodingStage

Kandinsky6SRVideoEncodingStage(vae: Any)

Bases: PipelineStage

Read the source video (and its audio), apply the x2.25 pre-upscale and KVAE-encode it into lq_latents.

Source code in fastvideo/pipelines/stages/kandinsky6_sr.py
def __init__(self, vae: Any) -> None:
    super().__init__()
    self.vae = vae

Functions: