Skip to content

kandinsky6

Classes

fastvideo.pipelines.stages.kandinsky6.Kandinsky6AudioDecodingStage

Kandinsky6AudioDecodingStage(audio_vae, vocoder)

Bases: PipelineStage

Decode Kandinsky6 audio latents into a waveform via two separate components -- audio_vae (mel-VAE decoder) and vocoder (BigVGAN-v2, mel->waveform) -- like LTX-2's audio_vae/vocoder pair. Writes batch.extra["audio"] / ["audio_sample_rate"], which VideoGenerator already knows how to mux into the saved video file.

Source code in fastvideo/pipelines/stages/kandinsky6.py
def __init__(self, audio_vae, vocoder) -> None:
    super().__init__()
    self.audio_vae = audio_vae
    self.vocoder = vocoder

fastvideo.pipelines.stages.kandinsky6.Kandinsky6CFGResolutionStage

Kandinsky6CFGResolutionStage(scheduler)

Bases: PipelineStage

Recomputes batch.do_classifier_free_guidance before text encoding using Kandinsky6's CFG contract. The diffusers reference (pipeline_kandinsky6_ti2va.py)'s do_classifier_free_guidance property gates on the standard guidance_scale > 1.0 (any guidance_scale <= 1.0, including below 1, runs cond-only with no negative-prompt encoding); this stage mirrors that. A PiFlow (distilled) scheduler never uses CFG regardless of the requested guidance_scale (its cond-only contract, and the guidance==1.0 requirement, are enforced by Kandinsky6DenoisingStage instead), so this also short-circuits to False for it and skips encoding a negative prompt that would otherwise go unused.

Source code in fastvideo/pipelines/stages/kandinsky6.py
def __init__(self, scheduler) -> None:
    super().__init__()
    self.scheduler = scheduler

fastvideo.pipelines.stages.kandinsky6.Kandinsky6DecodingStage

Kandinsky6DecodingStage(vae: ParallelTiledVAE, pipeline=None)

Bases: DecodingStage

Channel-last [B,T,H,W,C] -> channel-first [B,C,T,H,W] then the generic VAE decode, matching Kandinsky5DecodingStage.

Source code in fastvideo/pipelines/stages/kandinsky6.py
def __init__(self, vae: ParallelTiledVAE, pipeline=None) -> None:
    super().__init__(vae=vae, pipeline=pipeline)

fastvideo.pipelines.stages.kandinsky6.Kandinsky6DenoisingStage

Kandinsky6DenoisingStage(transformer, scheduler)

Bases: PipelineStage

Joint video/audio denoising with independent scheduler state per modality.

Mirrors the diffusers reference's denoise_loop (pipeline_kandinsky6_ ti2va.py): one joint transformer call per step for both modalities, CFG via uncond + w*(cond-uncond), video advanced through the shared scheduler's step() (which owns the sigma index), audio advanced with a manual Euler update using the same per-step sigma delta so the scheduler's step index is never double-advanced.

Source code in fastvideo/pipelines/stages/kandinsky6.py
def __init__(self, transformer, scheduler) -> None:
    super().__init__()
    self.transformer = transformer
    self.scheduler = scheduler

fastvideo.pipelines.stages.kandinsky6.Kandinsky6ImageEncodingStage

Kandinsky6ImageEncodingStage(vae: ParallelTiledVAE)

Bases: PipelineStage

Optional IT2VA conditioning: no-op for a pure T2VA call.

When batch.pil_image is supplied, encodes it through the shared video VAE and appends it as one extra "clean" reference frame at the end of the video latent sequence (the diffusers reference's default tail_cond_first_frame scheme), tagged via a token-type id so the transformer's visual_token_type_embeddings can distinguish it from generated frames. Kandinsky6DenoisingStage re-pins this frame every step and strips it back out after the loop.

Source code in fastvideo/pipelines/stages/kandinsky6.py
def __init__(self, vae: ParallelTiledVAE) -> None:
    self.vae = vae

fastvideo.pipelines.stages.kandinsky6.Kandinsky6LatentPreparationStage

Kandinsky6LatentPreparationStage(scheduler, transformer)

Bases: PipelineStage

Draw initial video and audio noise latents.

Runs for both T2VA and IT2VA calls -- the conditioning image (if any) is applied afterward by Kandinsky6ImageEncodingStage, which appends an extra reference frame rather than overwriting one drawn here, so the RNG order here does not depend on whether an image was supplied.

Source code in fastvideo/pipelines/stages/kandinsky6.py
def __init__(self, scheduler, transformer) -> None:
    super().__init__()
    self.scheduler = scheduler
    self.transformer = transformer

fastvideo.pipelines.stages.kandinsky6.Kandinsky6TextEncodingStage

Kandinsky6TextEncodingStage(text_encoders, tokenizers)

Bases: TextEncodingStage

Maps the request's max_sequence_length onto the Qwen encoder only.

Matches the diffusers reference, whose Qwen max_length = max_sequence_length + 129 (the prompt-template prefix crop, ENCODE_START_IDX). The CLIP pooled encoder always keeps its fixed 77-token tokenizer config regardless of the request: applying the same override to it (the generic TextEncodingStage's default behavior, one max_length for every encoder) makes a CLIP tokenizer with only 77 positions try to pad to whatever length was requested for Qwen and crash.

Source code in fastvideo/pipelines/stages/text_encoding.py
def __init__(self, text_encoders, tokenizers) -> None:
    """
    Initialize the prompt encoding stage.

    Args:
        enable_logging: Whether to enable logging for this stage.
        is_secondary: Whether this is a secondary text encoder.
    """
    super().__init__()
    self.tokenizers = tokenizers
    self.text_encoders = text_encoders
    self._last_audio_embeds: list[torch.Tensor] | None = None

Functions:

fastvideo.pipelines.stages.kandinsky6.audio_latent_duration

audio_latent_duration(video_latent_frames: int, *, fps: float, audio_sample_rate: int, audio_downsample_factor: int) -> int

Audio latent length matching the diffusers Kandinsky6 T2VA reference: ceil(pixel_frames / fps * audio_sample_rate / audio_downsample_factor), where pixel_frames = (video_latent_frames - 1) * 4 + 1 is the causal video VAE's temporal-compression convention (matches HunyuanVAEConfig.temporal_compression_ratio == 4).

Source code in fastvideo/pipelines/stages/kandinsky6.py
def audio_latent_duration(video_latent_frames: int, *, fps: float, audio_sample_rate: int,
                          audio_downsample_factor: int) -> int:
    """Audio latent length matching the diffusers Kandinsky6 T2VA reference:
    ``ceil(pixel_frames / fps * audio_sample_rate / audio_downsample_factor)``,
    where ``pixel_frames = (video_latent_frames - 1) * 4 + 1`` is the causal
    video VAE's temporal-compression convention (matches
    HunyuanVAEConfig.temporal_compression_ratio == 4).
    """
    pixel_frames = (video_latent_frames - 1) * 4 + 1
    return int(math.ceil(pixel_frames / fps * audio_sample_rate / audio_downsample_factor))