wan_animate ¶
Arch config for Wan2.2-Animate-14B (character animation / replacement).
Every value is transcribed verbatim from Wan-AI/Wan2.2-Animate-14B-Diffusers/transformer/config.json -- do not "tidy" them. Field names match that file's keys one-for-one because update_model_arch overlays any matching key from the checkpoint's config.json at load time; a renamed field silently stops receiving its checkpoint value.
The tower itself is the Wan-I2V one (this class only switches on the I2V knobs image_dim/added_kv_proj_dim); why that is the right base is explained on WanAnimate14BConfig in models/wan/pipeline_config.py. What Animate adds on top:
in_channels = 36 = 16 (noise) + 4 (mask) + 16 (conditional latent).pose_patch_embedding-- a second patchifier whose output is added to the video tokens (skipping the reference latent frame).- A LIA-style motion encoder (
motion_*fields) turning 512x512 face crops into 20-dim motion codes. The checkpoint stores its conv/linear weights unit-scale: the model must apply the StyleGAN2 runtime factor (1/sqrt(fan_in)) in forward. Loading these into vanillann.Conv2d/nn.Linearforwards succeeds and is silently wrong. - A face encoder +
face_adaptercross-attention blocks. The checkpoint indexes the adapters densely (face_adapter.0 .. .7) with no record of which transformer block each serves; adapteriserves blocki * inject_face_latents_blocks. That correspondence exists only here, so it is asserted in__post_init__-- a wrong value loads cleanly and steers the wrong blocks.
Naming contract with the model (fastvideo/models/dits/wan_animate.py): the Animate-specific modules keep checkpoint-identical parameter names (motion_encoder.*, face_encoder.*, face_adapter.*) so the loader passes them through verbatim; only pose_patch_embedding needs a mapping entry, mirroring how the base patch_embedding is wrapped in .proj.