s2v_stages ¶
Wan-S2V-specific pipeline stages and the clip plan the pipeline loops over.
These replicate the steps of the official runner (wan/speech2video.py) that the shared stages do differently:
- The reference image is VAE-encoded alone as one frame. The shared
ImageVAEEncodingStageinstead builds an I2V-style zero-padded video and encodes all of it, which is a different conditioning format entirely. - Decoding prepends temporal context so the causal VAE does not start cold: the reference latent on the first clip, the previous clip's motion latents afterwards. Only the generated span is kept, and the first clip also drops
WARMUP_FRAMESframes that decode with too little context. - Long audio is covered by generating several clips back to back. Each clip re-encodes the last
motion_framespixels of the video so far as its motion history (plan_clipsdecides how many clips a request needs).
Classes¶
fastvideo.pipelines.basic.wan.s2v_stages.S2VClipPlan dataclass ¶
How one request splits into clips.
num_frames is what the caller asked for and what they get back: the usual FastVideo 4k+1 count. infer_frames is what the transformer generates per clip (4n, the official runner's infer_frames). The first clip shows infer_frames - WARMUP_FRAMES of those, later clips all of them, and the concatenation is cut down to num_frames.
fastvideo.pipelines.basic.wan.s2v_stages.S2VDecodingStage ¶
Bases: DecodingStage
Decode with temporal context prepended, official-runner style.
The Wan VAE is causal in time: the first frames decode with less context and come out degraded. The official runner therefore decodes [context | generated latents] and keeps the trailing infer_frames pixels. The context is the previous clip's motion latents when there are any, else the reference latent -- and in that first-clip case 3 more warm-up frames are dropped. With output_type='latent' nothing is decoded, so the prepended context is sliced back off instead.
Source code in fastvideo/pipelines/stages/decoding.py
fastvideo.pipelines.basic.wan.s2v_stages.S2VRefImageEncodingStage ¶
S2VRefImageEncodingStage(vae: ParallelTiledVAE)
Bases: ImageVAEEncodingStage
VAE-encode the reference image as a single latent frame.
Writes batch.image_latent with shape [B, C, 1, h, w]. Deterministic (distribution mode, not a sample): the reference is ground truth to preserve, and the official runner's native VAE encode is deterministic too. encode_pixels is shared with the pipeline's motion-history encode so both conditioning latents go through exactly the same normalisation.
Source code in fastvideo/pipelines/stages/image_encoding.py
Methods:¶
fastvideo.pipelines.basic.wan.s2v_stages.S2VRefImageEncodingStage.encode_pixels ¶
encode_pixels(pixels: Tensor, fastvideo_args: FastVideoArgs) -> Tensor
[B, 3, T, H, W] pixels in [-1, 1] -> normalised latents [B, C, t, h, w].
Source code in fastvideo/pipelines/basic/wan/s2v_stages.py
Functions:¶
fastvideo.pipelines.basic.wan.s2v_stages.plan_clips ¶
plan_clips(num_frames: int, clip_frames: int) -> S2VClipPlan
Split num_frames output frames into clips of at most clip_frames.
A single clip covers clip_frames - WARMUP_FRAMES visible frames, so the default 84-frame clip yields exactly the 81-frame default request. Shorter requests shrink the clip instead of generating frames that get thrown away.