MMAudio recipes¶
Maintained family · Inference
MMAudio inference recipes
MMAudio adds synchronized audio to video or generates audio from text. The checkpoint must be in Diffusers layout; the recipe loads the converted FastVideo repo through the `MMAUDIO_MODEL_PATH` environment variable the example reads.
Pick a recipe and runtime
Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.
Loading recipe details...
Exact device and memory details appear only when a recorded run supports them.
Loading...
- Model
- Loading...
- Workload
- Loading...
- Source configuration
- Loading...
- Expected output
- Loading...
Loading... Setup
The generated commands expect a local clone:
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo Use Configuration for supported Python and CLI settings, Optimizations for attention and memory tradeoffs, and the support matrix for the supported model and optimization surface.
Troubleshooting
- If the example reports a missing local
converted_weights/mmaudio/large_44k_v2path, theMMAUDIO_MODEL_PATHenv var from the cookbook command was not applied; export it in the same shell. - Alternatively convert upstream weights yourself with
scripts/checkpoint_conversion/convert_mmaudio_to_diffusers.pyand point the env var at the result.
Evidence status
All recipes on this page are Source-backed: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are Unknown and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.