audio_encoding ¶
Audio conditioning stage for speech-driven video (Wan2.2-S2V).
Text has a tokenizer because language is discrete -- a finite vocabulary you can look words up in. Audio has none: a waveform is a continuous stream, so the "tokens" are manufactured by slicing time and learning features, not by lookup. This stage does that: load -> resample to wav2vec2's 16kHz -> encode -> resample the feature stream to the video frame rate -> bucket into a per-frame window.
The resampling is the load-bearing part. If the audio feature stream and the video latents disagree on frame count, nothing crashes -- the lips just drift out of sync. verify_output pins the frame count for that reason.
The waveform itself is also kept on batch.extra["audio"] so the saved MP4 carries the speech track, the same convention the LTX-2 and MagiHuman audio stages use.
Classes¶
fastvideo.pipelines.stages.audio_encoding.AudioEncodingStage ¶
Bases: PipelineStage
Waveform -> per-frame audio embeddings for the DiT's audio injector.
Writes batch.audio_embeds with shape [B, num_layers, C_a, num_frames]. All wav2vec2 hidden states are kept, not just the last: the model learns its own weighting over encoder depth (casual_audio_encoder.weights).