Skip to content

animate_stages

Wan-Animate-specific pipeline stages.

These replicate the conditioning assembly of the official runner (wan/animate.py) and diffusers' WanAnimatePipeline, which the shared stages do differently:

  • The denoised sequence carries an extra leading latent frame (the reference slot), so latent preparation allocates T_lat + 1 frames.
  • The conditional latent y is 20 channels -- a 4-channel folded I2V mask plus the 16-channel VAE encoding of [reference | zeros-or-background] -- temporally concatenated as [ref (1 frame) | target (T_lat frames)] and channel-concatenated with the noise by the shared DenoisingStage (channel order: noise 16 | mask 4 | cond 16 = the DiT's in_dim 36).
  • The pose skeleton video is VAE-encoded onto the same latent grid.
  • The face video stays in pixel space (the DiT's motion encoder consumes raw 512x512 crops).
  • Decoding drops the reference slot (vae.decode(latents[:, :, 1:])).

Every numeric convention here (mask folding, mask inversion, zeros-video encoding, argmax sampling, reflect padding) is transcribed from pipeline_wan_animate.py -- none of it crashes when wrong, it just produces subtly broken video, so deviations are bugs even when output "looks fine".

v1 scope: a single 77-frame segment. Multi-segment chaining (temporal-guidance frames from the previous segment) is pipeline-loop work on top of these stages.

Classes

fastvideo.pipelines.basic.wan.animate_stages.AnimateConditioningLatentsStage

AnimateConditioningLatentsStage(vae: ParallelTiledVAE)

Bases: ImageVAEEncodingStage

Assemble the 20-channel conditional latent y into batch.image_latent.

Layout (channel-first inside each frame group, ref frame first in time):

[ mask 4ch | cond latent 16ch ] x [ ref (1 frame) | target (T_lat) ]

Animation mode: the target's cond video is a zeros video encoded through the VAE -- raw zeros, i.e. mid-gray in the VAE's [-1,1] pixel space (NOT black; the diffusers reference also feeds raw zeros). VAE(zeros) != zero latents, so encoding it is load-bearing and must stay zeros to preserve the checkpoint's conditioning statistics -- do not "fix" it to fill(-1.0) (true black), which would break parity with the reference. The target mask is all zeros ("generate everything"). Replace mode: the target's cond video is the background video and the mask is the inverted character mask (input convention: white = generate), nearest-downsampled to the latent grid -- 1 on preserved background, 0 in the person-shaped hole.

Source code in fastvideo/pipelines/stages/image_encoding.py
def __init__(self, vae: ParallelTiledVAE) -> None:
    self.vae: ParallelTiledVAE = vae

fastvideo.pipelines.basic.wan.animate_stages.AnimateDecodingStage

AnimateDecodingStage(vae, pipeline=None)

Bases: DecodingStage

Decode without the reference slot.

The denoiser emits T_lat + 1 latent frames; the leading one is the reference-image slot, generated only as conditioning context. The official runner and diffusers both decode latents[:, :, 1:].

Source code in fastvideo/pipelines/stages/decoding.py
def __init__(self, vae, pipeline=None) -> None:
    self.vae: ParallelTiledVAE = vae
    self.pipeline = weakref.ref(pipeline) if pipeline else None

fastvideo.pipelines.basic.wan.animate_stages.AnimateFaceVideoStage

Bases: PipelineStage

Load the preprocessed face-crop video as raw pixels.

The DiT's motion encoder consumes pixel crops directly (no VAE): resized to motion_encoder_size (512), normalised to [-1, 1]. num_frames pixel frames yield exactly one motion-vector group per latent frame after the model's causal 4x funnel (num_frames = 4 * T_lat - 3).

fastvideo.pipelines.basic.wan.animate_stages.AnimateLatentPreparationStage

AnimateLatentPreparationStage(scheduler, transformer, use_btchw_layout: bool = False)

Bases: LatentPreparationStage

Allocate noise for T_lat + 1 latent frames: the extra slot is the reference frame, denoised alongside the video and dropped at decode.

Source code in fastvideo/pipelines/stages/latent_preparation.py
def __init__(self, scheduler, transformer, use_btchw_layout: bool = False) -> None:
    super().__init__()
    self.scheduler = scheduler
    self.transformer = transformer
    self.use_btchw_layout = use_btchw_layout

fastvideo.pipelines.basic.wan.animate_stages.AnimatePoseVideoEncodingStage

AnimatePoseVideoEncodingStage(vae: ParallelTiledVAE)

Bases: ImageVAEEncodingStage

VAE-encode the preprocessed skeleton video onto the target latent grid.

Writes batch.pose_latents with T_lat frames -- one fewer than the denoised sequence, because the reference slot carries no pose (the DiT adds pose tokens to frames 1..T only).

Source code in fastvideo/pipelines/stages/image_encoding.py
def __init__(self, vae: ParallelTiledVAE) -> None:
    self.vae: ParallelTiledVAE = vae

Functions: