kandinsky6 ¶
Classes¶
fastvideo.pipelines.stages.kandinsky6.Kandinsky6AudioDecodingStage ¶
Bases: PipelineStage
Decode Kandinsky6 audio latents into a waveform via two separate components -- audio_vae (mel-VAE decoder) and vocoder (BigVGAN-v2, mel->waveform) -- like LTX-2's audio_vae/vocoder pair. Writes batch.extra["audio"] / ["audio_sample_rate"], which VideoGenerator already knows how to mux into the saved video file.
Source code in fastvideo/pipelines/stages/kandinsky6.py
fastvideo.pipelines.stages.kandinsky6.Kandinsky6CFGResolutionStage ¶
Bases: PipelineStage
Recomputes batch.do_classifier_free_guidance before text encoding using Kandinsky6's CFG contract. The diffusers reference (pipeline_kandinsky6_ti2va.py)'s do_classifier_free_guidance property gates on the standard guidance_scale > 1.0 (any guidance_scale <= 1.0, including below 1, runs cond-only with no negative-prompt encoding); this stage mirrors that. A PiFlow (distilled) scheduler never uses CFG regardless of the requested guidance_scale (its cond-only contract, and the guidance==1.0 requirement, are enforced by Kandinsky6DenoisingStage instead), so this also short-circuits to False for it and skips encoding a negative prompt that would otherwise go unused.
Source code in fastvideo/pipelines/stages/kandinsky6.py
fastvideo.pipelines.stages.kandinsky6.Kandinsky6DecodingStage ¶
Kandinsky6DecodingStage(vae: ParallelTiledVAE, pipeline=None)
Bases: DecodingStage
Channel-last [B,T,H,W,C] -> channel-first [B,C,T,H,W] then the generic VAE decode, matching Kandinsky5DecodingStage.
Source code in fastvideo/pipelines/stages/kandinsky6.py
fastvideo.pipelines.stages.kandinsky6.Kandinsky6DenoisingStage ¶
Bases: PipelineStage
Joint video/audio denoising with independent scheduler state per modality.
Mirrors the diffusers reference's denoise_loop (pipeline_kandinsky6_ ti2va.py): one joint transformer call per step for both modalities, CFG via uncond + w*(cond-uncond), video advanced through the shared scheduler's step() (which owns the sigma index), audio advanced with a manual Euler update using the same per-step sigma delta so the scheduler's step index is never double-advanced.
Source code in fastvideo/pipelines/stages/kandinsky6.py
fastvideo.pipelines.stages.kandinsky6.Kandinsky6ImageEncodingStage ¶
Kandinsky6ImageEncodingStage(vae: ParallelTiledVAE)
Bases: PipelineStage
Optional IT2VA conditioning: no-op for a pure T2VA call.
When batch.pil_image is supplied, encodes it through the shared video VAE and appends it as one extra "clean" reference frame at the end of the video latent sequence (the diffusers reference's default tail_cond_first_frame scheme), tagged via a token-type id so the transformer's visual_token_type_embeddings can distinguish it from generated frames. Kandinsky6DenoisingStage re-pins this frame every step and strips it back out after the loop.
Source code in fastvideo/pipelines/stages/kandinsky6.py
fastvideo.pipelines.stages.kandinsky6.Kandinsky6LatentPreparationStage ¶
Bases: PipelineStage
Draw initial video and audio noise latents.
Runs for both T2VA and IT2VA calls -- the conditioning image (if any) is applied afterward by Kandinsky6ImageEncodingStage, which appends an extra reference frame rather than overwriting one drawn here, so the RNG order here does not depend on whether an image was supplied.
Source code in fastvideo/pipelines/stages/kandinsky6.py
fastvideo.pipelines.stages.kandinsky6.Kandinsky6TextEncodingStage ¶
Bases: TextEncodingStage
Maps the request's max_sequence_length onto the Qwen encoder only.
Matches the diffusers reference, whose Qwen max_length = max_sequence_length + 129 (the prompt-template prefix crop, ENCODE_START_IDX). The CLIP pooled encoder always keeps its fixed 77-token tokenizer config regardless of the request: applying the same override to it (the generic TextEncodingStage's default behavior, one max_length for every encoder) makes a CLIP tokenizer with only 77 positions try to pad to whatever length was requested for Qwen and crash.
Source code in fastvideo/pipelines/stages/text_encoding.py
Functions:¶
fastvideo.pipelines.stages.kandinsky6.audio_latent_duration ¶
audio_latent_duration(video_latent_frames: int, *, fps: float, audio_sample_rate: int, audio_downsample_factor: int) -> int
Audio latent length matching the diffusers Kandinsky6 T2VA reference: ceil(pixel_frames / fps * audio_sample_rate / audio_downsample_factor), where pixel_frames = (video_latent_frames - 1) * 4 + 1 is the causal video VAE's temporal-compression convention (matches HunyuanVAEConfig.temporal_compression_ratio == 4).