kandinsky6_audio_vae ¶
Kandinsky6 audio VAE: a thin nesting wrapper, not a new architecture.
FastVideo's existing MMAudioVAE (the mel<->latent codec) nested directly under vae.*, plus a mel_converter waveform->mel STFT frontend, matching the checkpoint's audio_vae layout. Only the module nesting is new; the layers are reused unmodified.
The BigVGAN-v2 mel->waveform vocoder is a separate pipeline component (vocoder, loaded by VocoderLoader as BigVGANV2; see Kandinsky6AudioDecodingStage). The checkpoint's audio_vae/*.safetensors also bundles a copy of the vocoder weights under vocoder.*; AudioDecoderLoader drops those keys before the strict state-dict load, since this class has no vocoder submodule.
mel_converter is only needed to encode real audio, which T2VA/IT2VA generation never does: it only decodes audio latents into a waveform.
Classes¶
fastvideo.models.audio.kandinsky6_audio_vae.Kandinsky6AudioVAE ¶
Bases: Module
Nests the reused MMAudioVAE to match the checkpoint's exact parameter names. Decodes latents to a mel spectrogram only -- the separate vocoder pipeline component turns that into a waveform (see module docstring).
Source code in fastvideo/models/audio/kandinsky6_audio_vae.py
Methods:¶
fastvideo.models.audio.kandinsky6_audio_vae.Kandinsky6AudioVAE.decode ¶
latents: [B, embed_dim, A] (1D-conv channel-first) -> mel: [B, num_mels, T].
fastvideo.models.audio.kandinsky6_audio_vae.Kandinsky6AudioVAE.encode_audio ¶
waveform -> mel -> VAE posterior. Not used by generation (see the module docstring).
fastvideo.models.audio.kandinsky6_audio_vae.Kandinsky6AudioVAE.remove_weight_norm ¶
Called by the loader after load_state_dict: MMAudioVAE's custom post-load weight renormalization (needs real loaded values to compute from).
Source code in fastvideo/models/audio/kandinsky6_audio_vae.py
fastvideo.models.audio.kandinsky6_audio_vae.Kandinsky6AudioVAE.wrapped_encode ¶
waveform -> mean audio latent in one call, matching the diffusers reference's MMAudioVAE.wrapped_encode.
fastvideo.models.audio.kandinsky6_audio_vae.MelConverter ¶
MelConverter(*, sampling_rate: int = 44100, n_fft: int = 2048, num_mels: int = 128, hop_size: int = 512, win_size: int = 2048, fmin: int = 0, fmax: int | None = 22050)
Bases: Module
Waveform -> log-mel-spectrogram STFT frontend, matching the diffusers reference's MelConverter/get_mel_converter("44k"). See module docstring: not exercised by the current (decode-only) inference path.