mmaudio_validation ¶
Validation loss and audio sampling for MMAudio feature training.
Classes¶
fastvideo.train.callbacks.mmaudio_validation.MMAudioValidationCallback ¶
MMAudioValidationCallback(*, data_path: str, every_steps: int = 5000, max_batches: int = 0, batch_size: int = 8, num_data_workers: int = 2, run_at_start: bool = False, use_ema: bool = False, inference_every_steps: int = 20000, inference_model_path: str = '', inference_num_samples: int = 16, inference_num_steps: int = 25, inference_guidance_scale: float = 4.5, inference_seed: int = 14159265, inference_save_video: bool = True, inference_log_to_tracker: bool = True, output_dir: str | None = None)
Bases: Callback
Evaluate cached val features and periodically run native V2A inference.
Validation loss follows the official MMAudio val_fn: sample a VAE posterior latent, a logit-normal flow time, prior noise, and independent video/text CFG masks. The RNG is reset for every pass, making values at different training steps directly comparable.
Optional inference reuses the live FSDP transformer in :class:MMAudioPipeline, while frozen VAE/vocoder weights are loaded from inference_model_path. Precomputed CLIP, Synchformer, and text features go directly into the pipeline, so validation does not decode source video or run feature encoders again.