mmaudio_feature_dataset ¶
Memory-mapped precomputed feature dataset for audio generation training.
Classes¶
fastvideo.dataset.mmaudio_feature_dataset.MMAudioFeatureDataset ¶
MMAudioFeatureDataset(root: str | Path, *, latent_seq_len: int, latent_dim: int, clip_seq_len: int, clip_dim: int, sync_seq_len: int, sync_dim: int, text_seq_len: int, text_dim: int, include_metadata: bool = False)
Bases: Dataset
Read MMAudio-compatible TensorDict memmaps without upstream imports.
Required tensors are mean, std, and text_features. Video caches additionally contain clip_features and sync_features. Audio-only caches omit both video tensors; fixed-size placeholders and a false video_exists flag are returned so video and audio datasets can be mixed by the same dataloader.
Source code in fastvideo/dataset/mmaudio_feature_dataset.py
Functions:¶
fastvideo.dataset.mmaudio_feature_dataset.build_mmaudio_feature_dataloader ¶
build_mmaudio_feature_dataloader(data_path: str | Sequence[str] | dict[str, int], *, batch_size: int, num_data_workers: int, seed: int, pin_memory: bool, feature_shapes: dict[str, int], include_metadata: bool = False) -> StatefulDataLoader
Build a distributed/stateful loader over one or more feature caches.
Source code in fastvideo/dataset/mmaudio_feature_dataset.py
fastvideo.dataset.mmaudio_feature_dataset.build_mmaudio_feature_dataset ¶
build_mmaudio_feature_dataset(data_path: str | Sequence[str] | dict[str, int], *, feature_shapes: dict[str, int], include_metadata: bool = False) -> Dataset
Build one map-style dataset over MMAudio TensorDict cache shards.
Unlike :func:build_mmaudio_feature_dataloader, this helper does not add a distributed sampler. Dataset-scale inference can therefore assign indices explicitly (for example range(rank, len(dataset), world_size)) without DistributedSampler padding the tail with duplicate samples.
Source code in fastvideo/dataset/mmaudio_feature_dataset.py
fastvideo.dataset.mmaudio_feature_dataset.compute_mmaudio_latent_stats ¶
compute_mmaudio_latent_stats(data_path: str | Sequence[str] | dict[str, int], *, latent_seq_len: int, latent_dim: int, chunk_size: int = 32) -> tuple[Tensor, Tensor]
Compute official MMAudio normalization stats from the first cache.
The reference trainer uses the posterior means from its first video dataset and reduces over sample and sequence dimensions. This chunked implementation preserves that contract without materializing the complete VGGSound tensor in RAM.