animate_stages ¶
Wan-Animate-specific pipeline stages.
These replicate the conditioning assembly of the official runner (wan/animate.py) and diffusers' WanAnimatePipeline, which the shared stages do differently:
- The denoised sequence carries an extra leading latent frame (the reference slot), so latent preparation allocates
T_lat + 1frames. - The conditional latent
yis 20 channels -- a 4-channel folded I2V mask plus the 16-channel VAE encoding of [reference | zeros-or-background] -- temporally concatenated as[ref (1 frame) | target (T_lat frames)]and channel-concatenated with the noise by the shared DenoisingStage (channel order: noise 16 | mask 4 | cond 16 = the DiT's in_dim 36). - The pose skeleton video is VAE-encoded onto the same latent grid.
- The face video stays in pixel space (the DiT's motion encoder consumes raw 512x512 crops).
- Decoding drops the reference slot (
vae.decode(latents[:, :, 1:])).
Every numeric convention here (mask folding, mask inversion, zeros-video encoding, argmax sampling, reflect padding) is transcribed from pipeline_wan_animate.py -- none of it crashes when wrong, it just produces subtly broken video, so deviations are bugs even when output "looks fine".
v1 scope: a single 77-frame segment. Multi-segment chaining (temporal-guidance frames from the previous segment) is pipeline-loop work on top of these stages.
Classes¶
fastvideo.pipelines.basic.wan.animate_stages.AnimateConditioningLatentsStage ¶
AnimateConditioningLatentsStage(vae: ParallelTiledVAE)
Bases: ImageVAEEncodingStage
Assemble the 20-channel conditional latent y into batch.image_latent.
Layout (channel-first inside each frame group, ref frame first in time):
[ mask 4ch | cond latent 16ch ] x [ ref (1 frame) | target (T_lat) ]
Animation mode: the target's cond video is a zeros video encoded through the VAE -- raw zeros, i.e. mid-gray in the VAE's [-1,1] pixel space (NOT black; the diffusers reference also feeds raw zeros). VAE(zeros) != zero latents, so encoding it is load-bearing and must stay zeros to preserve the checkpoint's conditioning statistics -- do not "fix" it to fill(-1.0) (true black), which would break parity with the reference. The target mask is all zeros ("generate everything"). Replace mode: the target's cond video is the background video and the mask is the inverted character mask (input convention: white = generate), nearest-downsampled to the latent grid -- 1 on preserved background, 0 in the person-shaped hole.
Source code in fastvideo/pipelines/stages/image_encoding.py
fastvideo.pipelines.basic.wan.animate_stages.AnimateDecodingStage ¶
Bases: DecodingStage
Decode without the reference slot.
The denoiser emits T_lat + 1 latent frames; the leading one is the reference-image slot, generated only as conditioning context. The official runner and diffusers both decode latents[:, :, 1:].
Source code in fastvideo/pipelines/stages/decoding.py
fastvideo.pipelines.basic.wan.animate_stages.AnimateFaceVideoStage ¶
Bases: PipelineStage
Load the preprocessed face-crop video as raw pixels.
The DiT's motion encoder consumes pixel crops directly (no VAE): resized to motion_encoder_size (512), normalised to [-1, 1]. num_frames pixel frames yield exactly one motion-vector group per latent frame after the model's causal 4x funnel (num_frames = 4 * T_lat - 3).
fastvideo.pipelines.basic.wan.animate_stages.AnimateLatentPreparationStage ¶
AnimateLatentPreparationStage(scheduler, transformer, use_btchw_layout: bool = False)
Bases: LatentPreparationStage
Allocate noise for T_lat + 1 latent frames: the extra slot is the reference frame, denoised alongside the video and dropped at decode.
Source code in fastvideo/pipelines/stages/latent_preparation.py
fastvideo.pipelines.basic.wan.animate_stages.AnimatePoseVideoEncodingStage ¶
AnimatePoseVideoEncodingStage(vae: ParallelTiledVAE)
Bases: ImageVAEEncodingStage
VAE-encode the preprocessed skeleton video onto the target latent grid.
Writes batch.pose_latents with T_lat frames -- one fewer than the denoised sequence, because the reference slot carries no pose (the DiT adds pose tokens to frames 1..T only).