TL;DR:
FastVideo, is open sourcing FastH3 Preview v1 for text-to-video-and-audio (T2VA), post-trained on Minimax H3.
We collaborated closely Nuva Lab,
NVIDIA FastGen (Julius Berner, Chao Liu, Arash Vahdat), and NVIDIA Enterprise Products (Pengcheng Li, Cliff Woolley) on FastH3.
We took production readiness, quality, user experience and openness seriously. We hope this joint effort will lead to a solid foundation for people who would love to use it in real commerical workload beyond an academic experimentation.
Speed
- FastH3 can generate 15s 768p video in less than 13s with sub-realtime generation on 8xB200 GPUs.
- Up to 14x speedup on a single Nvidia Blackwell GPU
Quality
- We used 1k+ B200 training hours, paired with real world multi-shot, visual audio synced input distribution and output formats for best possible quality preservation.
- FastH3 natively supports variable resolution, aspect ratio, and duration. In a single checkpoint.
Openness
- Start with the 4-step VSA / Data-Free checkpoint, our recommended FastH3 Preview v1 release. We provide full weights and a pre-extracted LoRA, plus dense and synthetic-data ablations.
- Fully open source with training (coming soon!) and inference code recipe for your customization.
What’s Next
- Follow us along for image ref (FL2VA) and full omni ref (Ref2VA) coming in the next a few weeks
- Motion and more generation quality improvements
- Nvfp4 and GPU memory reduction.
- Optimizations targeting local AI devices including RTX, DGX Sparks, and Apple MLX.
- New training runs using FastGen team’s new Parallel Decoding Distillation (PDD) method!
Why open H3 matters
The strongest video systems were mostly closed until MiniMax released the H3-Base weights. With downloadable weights, the community can inspect H3, post-train it, replace kernels, and run it on its own hardware.
This continues our work on FastWan sparse distillation and FastWan-QAD. Both releases paired checkpoints with their FastVideo training and inference stacks. FastH3 does the same for H3.
FastH3 Preview v1
FastH3 distills H3’s base transformer for text-to-video-and-audio (T2VA) and reuses the H3-Base text encoder, video VAE, audio VAE, tokenizers, and schedulers.
FastH3 is not limited to the 1344×768 benchmark resolution. The checkpoints were trained and validated at multiple aspect ratios, including square, portrait, landscape, and ultrawide 768p video. FastVideo accepts custom heights and widths in multiples of 32.
Our recommended Preview v1 checkpoint is 4-step VSA / Data-Free. It trains from prompts without target videos and is the checkpoint to try first.
We provide both full weights and a pre-extracted LoRA for the recommended checkpoint.
| Checkpoint | Pre-extracted LoRA | Training source | Attention | Training step |
|---|---|---|---|---|
| VSA / Data-Free | LoRA folder | Prompts only, mixed shapes | VSA, 90% sparse, tile 64 | 1300 |
The validation gallery and VSA performance rows below use this checkpoint.
The recommended checkpoint and its LoRA require FastVideo’s Video Sparse Attention (VSA-H3) backend and kernel to achieve the reported speed and quality. Dense attention is not a drop-in substitute.
All Preview v1 checkpoints currently support T2VA only. H3 uses the base transformer for both T2VA and first/last-frame-to-video-and-audio (FL2VA), but these students were not trained with first/last-frame conditioning. Reference-to-video-and-audio (Ref2VA) uses a separate reference transformer and needs its own distilled checkpoint. FL2VA and Ref2VA checkpoints are in development.
Validation samples
Fifteen samples from the recommended VSA / Data-Free checkpoint, with generated stereo audio and full prompts.
Prompt
integrated_multimodal_description: [Shot 1] A 9-second 16:9 widescreen educational documentary tutorial in clean paper-textured motion-graphics design, following one illustrated paper-craft instructor at a neatly gridded tabletop. The instructor places a coral square of paper on the grid, aligns its corners, and folds it diagonally into a sharp crease; a Japanese title card reading "折り紙ランタン" appears at the upper left, while a thin animated guide line and the label "谷折り" appear beside the crease. The overhead camera makes a measured push-in as the instructor smooths the fold, with small geometric accents tracking the paper edges. [Shot 2] At 00:04.500, the same instructor unfolds the paper, rotates it a quarter turn, and presses the intersecting creases into a compact lantern shape; the camera cuts to a close three-quarter tabletop view and makes a short lateral track to reveal the dimensional form. Animated arrows trace the final fold, and the Japanese completion card "完成" settles along the lower edge as the instructor places the lantern in a small cardboard tray. overall_soundscape: Close ASMR foley records the paper's dry flex and crisp crease, fingertip taps on the matte work surface, and a soft sleeve rustle, with a quiet studio room tone underneath. A small cardboard tray clicks when the instructor sets the finished lantern down, while gentle breathing remains audible. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] 3D CG game-cinematic rendering in a 16:9 widescreen composition frames an original music-dance stage performance on a circular black platform surrounded by cyan light strips. Four principal dancers stand in a tight diamond: the copper-jacketed lead dancer (S1), the teal-haired dancer (S2), the silver-sleeved dancer (S3), and the orange-accented dancer (S4). The stage floor is blank and unlettered, with no readable signs, captions, logos, labels, or subtitles. The four dancers launch into rapid synchronized footwork, shoulder hits, and a low crouching sweep while keeping their hands empty; S1 shouts in a clear on-screen voice, <d>[English] Three, two, one—break!</d> The camera begins a fast push in with medium amplitude as their shoes strike the platform together. [Shot 2] At 00:02.000, the camera cuts to a low side angle as S2 vaults into a twisting aerial jump, S3 slides beneath the movement, and S1 and S4 pivot around them without breaking formation. S2 lands and calls, <d>[Spanish] ¡Gira conmigo!</d> The camera tracks right at fast speed, matching the dancers as they sweep across the stage. [Shot 3] At 00:04.000, the camera cuts to a close three-quarter view as S1 and S4 perform opposing backflips while S2 and S3 execute rapid heel-toe steps and snap into a mirrored pose. S3 shouts, <d>[English] Cross left!</d> The camera arcs around the group with large amplitude at fast speed, revealing the cyan floor lights streaking beneath their feet. [Shot 4] At 00:06.000, the camera cuts to a high overhead view as all four dancers rotate into a moving diamond, trade positions through two fast spins, and rebound from a simultaneous knee drop into a vertical leap. S4 calls, <d>[Spanish] ¡Ahora, salta!</d> The camera rolls slightly clockwise while descending in a controlled arc toward the center of the formation. [Shot 5] At 00:08.000, the camera cuts to a tight frontal angle as S1 and S2 exchange a rapid mirrored footwork sequence, then S3 and S4 cross behind them in a synchronized sliding pass. S1 snaps, <d>[English] Switch!</d> The camera tracks backward with large amplitude at fast speed, keeping all four dancers visible as they accelerate into the final combination. All four performers keep empty hands throughout, and no additional performers enter the stage. [Shot 6] At 00:10.000, the camera cuts to a wide frontal composition as the four dancers sprint toward the platform center, perform a coordinated four-person jump with separated spins, and land in a staggered diamond pose. S2 shouts, <d>[Spanish] ¡Juntos!</d> The camera pulls out with medium amplitude at slow speed while S1 raises one arm, S2 drops into a lunge, S3 kneels, and S4 holds a sharp standing angle; their final foot impacts and breathing settle by the end of the 12.00-second event. overall_soundscape: A cavernous indoor stage ambience surrounds the performance with steady ventilation and a faint electrical hum from the floor lights. Hard shoe impacts, sliding soles, fabric snaps, and brief handclaps synchronize with the choreography. Rapid breathing, exertion grunts, and the resonance of each landing continue beneath the spoken calls. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] 2D anime action fantasy rendered as kinetic motion graphics in a 16:9 widescreen frame, with crisp cel-shaded figures, layered vector-like energy planes, hard-edged color blocks, and luminous geometric effects. In a suspended observatory above a storm of violet clouds, two rival sky-mages face each other on a circular transparent platform: the young woman in a white-and-cobalt coat (S1) grips a silver staff, while the young man in a black-and-crimson mantle (S2) crouches opposite her with a glowing amber gauntlet. S1 sweeps her staff across the floor, launching three blue energy crescents; S2 braces and punches forward, shattering the crescents into angular fragments with a concussive flash. The camera executes a fast high-angle arc around both figures, revealing the platform rotating above the clouds as their loose fabric and hair whip in the wind. S1 shouts in a clear, urgent young woman's voice: <d>[Chinese] 让开!星门要塌了!</d> [Shot 2] At 00:03.000, the camera cuts to a low close-up as S2 drives his gauntlet into the platform, sending a crackling amber shockwave outward before the transparent floor buckles into rising polygonal slabs. S1 leaps from slab to slab, twists over the expanding fracture, and slashes downward with her staff, drawing a vertical blue beam that pins the collapsing geometry in place. S2 looks up through the beam and answers in a tense young man's voice: <d>[Chinese] 我不会把钥匙交给你!</d> The camera whip-pans from S2's fist to S1's airborne silhouette, with graphic streaks compressing the motion between viewpoints. [Shot 3] At 00:05.800, the camera cuts to an overhead view as the two mages are pulled toward the observatory's rotating central aperture; S1 plants her staff into a seam while S2 catches a floating metal ring and swings around it, then releases himself toward her. Their weapons collide at the aperture, producing a rapidly spinning lattice of blue and amber planes that bends the cloud vortex into a funnel. S1 reaches through the lattice and grabs S2's wrist as the platform fragments accelerate around them. The camera plunges vertically through the rotating energy lattice, then rolls clockwise to keep both bodies centered while the physical debris whips past. [Shot 4] At 00:08.700, the camera cuts to a tight side view as S1 and S2 pull in opposite directions, their combined strain forcing the aperture shut in a single explosive fold. The remaining slabs fly upward, the cloud vortex collapses into calm violet layers, and both characters land back-to-back on the restored circular platform. S2 lowers his gauntlet and S1 releases her staff; they exchange a wary glance while small blue and amber sparks orbit their hands, then the camera pulls out with large amplitude at fast speed to reveal the stabilized observatory and the two figures silhouetted against the clearing sky through the 11.00-second endpoint. overall_soundscape: Violent wind rushes around the suspended observatory throughout, with fabric snaps, hair flutter, and layered cloud turbulence beneath the action. Staff sweeps, gauntlet impacts, crystalline fractures, flying metal rings, and rapid energy discharges produce sharply synchronized physical sounds. Heavy landings thud against the platform, followed by falling fragments skittering and the final sparks crackling softly. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic. The camera holds a static medium close-up on a beautiful woman with long flowing brown hair and expressive green eyes, sitting alone in the corner of a dimly lit room. She wears a simple, elegant black dress that hugs her curves; her posture is slumped, her head bowed slightly, her shoulders drooping, and her eyes glisten under the low light with her mouth turned downward. Behind her, the room is sparsely furnished with a few pieces of vintage furniture, including a worn wooden dresser and an old upholstered armchair, while a single table lamp casts soft shadows across the walls. She draws a slow, trembling breath, then lowers her gaze toward the floor. overall_soundscape: The room is nearly silent, filled only with the woman's slow, shallow breathing and a soft sigh escaping her lips. Her black dress rustles faintly as her shoulders shift, and a vintage clock ticks quietly somewhere in the room. non_diegetic_music: A sparse solo piano plays slow, widely spaced notes in a low register at a soft volume, with long pauses between phrases and no crescendo.
Prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic. A sweeping wide shot captures a large wooden sailing ship navigating the vast open ocean beneath a starlit night sky, the camera performing an Arc Shot around the vessel at slow speed. The ship's tall sails billow gently in the wind, illuminated by the soft warm glow of lanterns hanging from the rigging, while moonlight and starlight shimmer across the subtly undulating waves. Weathered wooden planks, taut ropes, and towering masts are rendered in fine detail as the hull rocks with the swell, and in the background the deep blue of the sea blends into the dark expanse of a night sky densely dotted with stars, the lone ship small against the horizon. overall_soundscape: Wind hums steadily through the rigging and fills the canvas sails with soft flapping and billowing sounds. Wooden timbers and ropes creak rhythmically as the ship rolls, and waves lap and splash gently against the hull. non_diegetic_music: A slow orchestral score of sustained strings and low brass builds gradually in volume, layered with a sparse piano melody in a gentle, unhurried rhythm.
Prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic. A medium close-up frames a cozy study room seen through a window, the dark wooden window frame forming a natural border around the view. Inside, tall well-stocked bookshelves packed with books and small decorative objects line the wall, a wooden desk sits cluttered with pens, notebooks, loose papers and an open laptop, and a soft reading lamp casts a warm golden glow across the desktop. Framed photos and potted green plants add personal touches to the shelves and desk corners, while the dim evening light outside the glass contrasts with the warm interior. The camera pushes in at slow speed toward the window, drawing slightly closer to the serene study space. overall_soundscape: A faint wind stirs outside the window, occasionally pressing softly against the glass. The laptop emits a low, steady hum, and the wooden desk and shelves give off subtle creaks in the stillness of the room. non_diegetic_music: A soft piano piece plays at a slow tempo with sparse, evenly spaced notes. Gentle sustained strings join underneath at a low, steady volume that remains consistent throughout.
Prompt
integrated_multimodal_description: [Shot 1] 2D-animated in a whimsical storybook style. A wide shot of the night sky: a small, glowing moon with a friendly face, big expressive eyes and a gentle smile, surrounded by soft radiant light, floats gracefully across the dark blue sky above a cozy village below. Twinkling stars dot the sky and fluffy clouds drift past as the camera follows the moon in a Tracking Shot at slow speed. [Shot 2] At 00:02.600, the camera cuts to a medium close-up of the little moon as it gently descends toward the sleepy village, the camera Tilting Down with it at slow speed. Its warm glow spreads over quaint houses with thatched roofs, a winding river that reflects the moonlight, and a few villagers waking up: one opens wooden shutters, and another stretches and yawns in a doorway. overall_soundscape: A gentle night wind drifts through the air and crickets chirp softly around the village, while the winding river laps quietly against its banks. Wooden shutters creak open and a villager yawns. non_diegetic_music: A soft ambient score of sustained string pads and slow piano arpeggios plays at slow speed, with a light celesta melody entering in the second half. The dynamics stay quiet throughout, swelling slightly as the moon descends.
Prompt
integrated_multimodal_description: [Shot 1] 3D CG, medium shot of Shrek, a large green ogre with trumpet-shaped ears, a round belly, and a wide grin, wearing a cream long-sleeved shirt, a brown vest, and a brown checked kilt, bouncing on a round trampoline in a rustic forest clearing surrounded by tall trees, a wooden fence, and a clear blue sky. He springs off the mat, spreads his arms and legs wide in mid-air, lands, and bounces up again, his belly jiggling with each impact. The camera shakes slightly each time he lands on the trampoline. overall_soundscape: The trampoline springs creak and twang with each bounce, punctuated by soft thuds as his feet hit the mat. A gentle breeze rustles the tree leaves and birds chirp in the background. Shrek lets out hearty laughter and short grunts of effort as he jumps. non_diegetic_music: A light orchestral score with pizzicato strings, tuba, and xylophone plays at a fast tempo with a bouncy, syncopated rhythm. The volume rises slightly with each bounce and settles as he lands.
Prompt
integrated_multimodal_description: [Shot 1] 3D CG. In a static wide shot, an ogre mage with green skin and wild dark hair stands in the midst of a vibrant alien landscape, wearing a tattered earth-toned robe and gripping a wooden staff adorned with glowing blue runes. Lush, brightly colored flora surrounds him, sparkling waterfalls cascade over rocks nearby, and rolling hills stretch toward a distant horizon beneath a sky filled with bioluminescent clouds. His face is contorted in concentration as he conjures a swirling mass of blue energy above his raised staff, his gaze fixed upward on the glowing spell, whose radiant light washes over the plants and terrain around him. overall_soundscape: Waterfalls rush and splash over the rocks, blending with a soft wind that rustles the dense alien foliage. The swirling blue spell emits a low, crackling hum of energy. The ogre mage draws slow, strained breaths as he holds the spell. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] Realistic, aesthetically refined 4K studio performance video, opening in a wide full-body composition of one male Kazakh dancer centered on a stage. He wears a colorful traditional shapan with intricate embroidered designs and a matching embroidered hat; the fabric’s vibrant colors and decorative patterns remain clearly visible. Behind him, a scenic backdrop depicts the Kazakh steppe with rolling hills beneath a vast blue sky. Soft natural-style studio lighting evenly illuminates the dancer, stage, and landscape backdrop. The camera begins a smooth, slow forward tracking move as he glides across the stage with graceful traditional Kazakh dance gestures, extending and shaping his arms while his embroidered sleeves sway. He gathers momentum into a dynamic leap, landing lightly with a soft footfall, then continues into fluid spins; the shapan flares outward in a colorful circle as the camera tracks laterally with him at a steady pace. In the final seconds, the camera gently arcs closer around his movement while he completes the last controlled spin and settles into a poised traditional dance stance, facing forward against the open steppe landscape. overall_soundscape: Quiet studio ambience surrounds the performance. Soft rhythmic footfalls and subtle fabric swishes accompany the dancer’s leaps, gliding steps, and spins. A faint natural spaciousness complements the steppe imagery. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] Cinematic fantasy realism, a wide aerial composition follows a majestic emerald-green dragon soaring above a dense misty forest at sunset. Its large leathery wings beat in slow, powerful cycles, and its long sinuous tail trails gracefully behind its regal body. The camera glides forward and slightly downward alongside the dragon as it moves effortlessly between the upper canopy and drifting pale mist. Warm patches of sunlight filter through the tall trees below while the sky behind it shifts from vibrant orange and pink near the horizon to deep purple and blue overhead. The dragon’s piercing amber eyes catch the fading light as it banks gently, wingbeats creating visible ripples in the mist and soft gusts through the treetops. [Shot 2] At 00:08.500, a low-angle view from beneath the forest canopy looks upward as the same dragon passes overhead in a grand, unhurried ascent. The camera tilts up smoothly, framing its broad scaled chest, outstretched leathery wings, and sweeping tail against the layered sunset sky. Its wings give several resonant, powerful flaps, stirring loose leaves and mist below, before it glides onward above the tall trees and recedes into the deepening purple-blue distance. overall_soundscape: Continuous high-altitude wind moves through the forest canopy and around the dragon’s wings. Each heavy wingbeat produces a deep leathery whoosh, accompanied by soft gusts that rustle leaves and shift the mist. Distant forest ambience remains subdued beneath the passing flight. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] Bright, polished three-dimensional CG game rendering in a vertical 9:16 creator-phone composition, presented as a single unbroken UGC comedy meme take lasting from 00:00.000 to 00:12.000. Two principal characters, a green-haired hoodie-wearing game creator and a compact orange robot sidekick, attempt a ridiculous kitchen challenge: the creator flips an oversized pancake toward a plate while the robot races underneath carrying the plate, slips on a rolling lemon, rebounds off a stool, and launches the plate upward just as the pancake bounces off the creator's helmet and lands perfectly. Their movements are high-energy and physics-sensitive, with exaggerated squash-and-stretch, frantic footwork, sliding props, and expressive surprised faces. A large floating creator caption appears above them reading "今天一定成功". The virtual phone camera begins at arm's length on the creator and robot, then smoothly tracks backward as they sprint, swings low alongside the robot's skid, rises with the airborne plate, and arcs around the creator for a close celebratory reveal, visibly coordinating its speed and framing with each physical action while remaining one continuous take. Do not introduce any third character. Do not show brand logos or English lettering. overall_soundscape: The diegetic kitchen ambience includes a faint appliance hum, rubbery game-world footsteps, the lemon rolling and tapping across the floor, stool creaks, plate clacks, pancake flops, and a soft successful landing thump. The two characters produce startled grunts, breathy exertion, and exuberant non-verbal laughter-like chirps. non_diegetic_music: An audience-only 142 BPM electro-pop score with clipped synth stabs, punchy handclap percussion, and a rising bass pulse drives the montage-like comedic pacing and camera timing, dropping suddenly at the successful landing; the characters neither perform nor hear this music.
Prompt
integrated_multimodal_description: [Shot 1] A square 1:1 live-action photoreal studio tabletop becomes a text-free motion-graphics title infographic, lit by soft neutral panels with crisp practical shadows and realistic acrylic reflections. Three clearly distinct adult designers, the principal subjects, are visible from the chest down around a matte white work surface, each wearing a different solid-color glove: cobalt blue, warm amber, and deep green. The blue-gloved designer places a translucent blue ring at center, the amber-gloved designer aligns three blank white bars into a clean radial layout, and the green-gloved designer rotates a small amber disc into the open gap, creating deliberate moderate coordinated motion. After these actions, the only camera movement is one smooth, controlled 10-centimeter overhead push-in at a slow, even pace. No extra hands, faces, logos, labels, readable text, captions, or subtitles appear; every graphic panel remains completely unmarked, and the live-action materials never become cartoon-like or digitally rendered. [Shot 2] At 00:03.800, the three designers withdraw their hands in sequence, then simultaneously tap the completed abstract arrangement once so the rings, bars, and disc settle into a precise balanced emblem for the final title-infographic frame. The camera holds the closer overhead composition without further movement, showing photoreal surface texture, tiny edge highlights, and stable graphic alignment through the 00:06.000 endpoint. No new subject enters, no object jumps position, and there is no camera shake or distracting reflection. overall_soundscape: Close ASMR foley captures the soft nitrile-glove friction, faint sleeve rustle, quiet controlled breathing, translucent acrylic sliding on matte board, delicate bar-to-table taps, and the final synchronized fingertip clicks. The room remains nearly silent and dry, with no speech, singing, voice, or other human vocalization. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] A dense-production, single-take 3:4 portrait motion-graphics design opens on a deep cobalt field with layered translucent circles and angular amber shards, framing one lone flat-vector lantern keeper (S1) at center. The character cradles a tiny fading amber star in both hands, looks down with worried concentration, then slowly lifts it toward the chest as the surrounding shapes pulse with moderate, elastic motion. The performer (S1) sings an intimate diegetic melody in English: <d>[English]I carry one small light through the blue</d>. As the star dims, S1 presses one palm against it, and a ring of warm geometric light expands outward; the performer (S1) continues singing: <d>[English]Wake, little star, and lead me home</d>. The star brightens, reflected amber shapes sweep across the character’s face, and S1 turns slightly toward the glow with a relieved expression. After the subject’s opening gesture, the only camera movement is a smooth, continuous slow push-in from a medium portrait framing to a close-up, increasing gently in amplitude across the full 14-second event from 00:00.000 to 00:14.000. Keep the graphic palette limited to cobalt, charcoal, ivory, and amber, with crisp vector edges, layered parallax planes, and a text-free frame throughout. overall_soundscape: A quiet synthetic room tone surrounds the performance, with soft breath from S1, faint paper-like rustles from the layered shapes, and delicate glassy pulses synchronized to the star’s fading and brightening. The final flare produces a warm low shimmer and a gentle exhale. non_diegetic_music: N/A
Prompt
integrated_multimodal_description: [Shot 1] Live-action photorealism in a 4:3 landscape composition frames a solo contemporary dancer centered on an amber-lit theater stage, backed by black curtains and narrow beams of white light. The dancer, a young woman in a cobalt-blue cropped jacket, silver sash, and black trousers, grips a closed red folding fan; she is the only principal performer and uses the stable speaker ID (S1). She snaps her head toward the audience and says sharply: <d>[Japanese]始めます!</d> Then she launches into high-speed percussive footwork, sweeps the fan across her body, and spins twice while the camera tracks backward with large amplitude at fast speed, arcing to keep her centered as she advances diagonally toward stage right. [Shot 2] At 00:03.100, the camera cuts to a low three-quarter angle and continues a fast tracking movement alongside her as she drives into a running leap, rotates in midair, lands with one knee bent, and snaps the red fan open above her head. She rises through a final sharp turn, thrusts the fan toward the lens, and says breathlessly: <d>[Japanese]これで終わりです!</d> The camera arcs around her with large amplitude at fast speed, settling on her extended-arm finishing pose as the stage lights tighten around her at the end of the six-second performance. overall_soundscape: Rapid barefoot footfalls strike the wooden stage in tightly spaced rhythms, accompanied by fabric swishes, the metallic snap of the folding fan, and heavy landing thumps. Her breathing grows audible during the leap while the theater's low ventilation hum continues beneath the performance. non_diegetic_music: N/A
Performance
Each local test uses FastVideo’s optimized inference at 1344×768 and 24 FPS with audio. The 5s, 10s, and 15s shapes contain 124, 243, and 345 frames. We report the median of three timed requests after one full warmup. Model loading and compilation are excluded. End-to-end time includes encoding, denoising, decoding, audio, muxing, and file output.
| Model / runtime | Duration | 1× B200 E2E (s) | 4× B200 E2E (s) | 8× B200 E2E (s) | Speedup over Base H3 (1× / 4×) |
|---|---|---|---|---|---|
| Base H3 · Dense FA4 | 5s | 132.5 | 40.6 | 1.0× / 1.0× | |
| 10s | 377.4 | 108.7 | 1.0× / 1.0× | ||
| 15s | 678.7 | 193.1 | 1.0× / 1.0× | ||
| Preview v1 VSA / Data-Free · 90% sparse | 5s | 16.2 | 6.1 | 6.84 | 8.16× / 6.65× |
| 10s | 31.1 | 12.0 | 11.66 | 12.13× / 9.03× | |
| 15s | 47.2 | 15.5 | 12.88 | 14.38× / 12.48× | |
| Preview v1 Dense / Data-Free · Dense FA4 | 5s | 18.3 | 6.8 | 7.24× / 5.97× | |
| 10s | 50.2 | 15.0 | 7.52× / 7.25× | ||
| 15s | 91.3 | 25.6 | 7.43× / 7.54× |
Speedup uses the unrounded timings for the same duration and GPU count; no 8× speedup is claimed without a matched Base H3 run.
VSA / Data-Free is our recommended release and default performance path. On B200, this path uses FastVideo’s optimized tile-64 CUDA VSA kernel, regional DiT compilation, H3 fusions, and compiled video VAE. The 8× B200 measurements also use SP8 parallel VAE decoding.
How FastH3 works
Base H3 calls its 33B audio-video diffusion transformer 49 times. FastH3 lowers that cost in two ways: four calls instead of 49, and less attention work inside each call.
Distribution Matching Distillation (DMD2) cuts the number of calls. DMD2 trains the student with a frozen Base H3 teacher and a learned critic. The difference between their score estimates supplies the training signal. For prompt-only runs, backward simulation exposes the student to the few-step states it will see at inference. The synthetic-video runs instead start from forward-noised Base-H3 video-and-audio latents.
VSA makes each call cheaper. The VSA paper introduces trainable sparse attention for video diffusion. In FastH3, the student keeps about 10% of eligible video-to-video tiles (90% sparsity) with 64-token blocks. Text and audio remain dense. The teacher and critic also use dense attention, giving the sparse student a full-attention target. This extends the FastVideo sparse-distillation recipe to H3.
Other runs and ablations
We publish three comparison runs covering training source, training duration, and dense attention. All use four DiT calls. Their pre-extracted LoRAs are grouped in one FastH3 Preview LoRA repository.
| Checkpoint | Purpose | Pre-extracted LoRA | Training source | Attention | Training step |
|---|---|---|---|---|---|
| VSA / Synthetic / Step 1300 | Training source at matched step | LoRA folder | Synthetic Base-H3 videos | VSA, 90% sparse, tile 64 | 1300 |
| VSA / Synthetic / Step 1900 | Longer synthetic training | LoRA folder | Synthetic Base-H3 videos | VSA, 90% sparse, tile 64 | 1900 |
| Dense / Data-Free | Dense-attention reference | LoRA folder | Prompts only, mixed shapes | Dense FA4 | 1000 |
Here, “data-free” means prompt-only training; the synthetic runs use videos generated by Base H3. Dense / Data-Free is the full-attention reference.
The two synthetic checkpoints use the same VSA architecture and runtime as the recommended checkpoint, so we do not repeat their latency in the performance table. Both require the VSA-H3 backend and tile-64 kernel; Dense / Data-Free uses FA4.
Try FastH3
The setup below targets four NVIDIA B200 GPUs with CUDA 13. On first use, the launcher downloads Base H3 and the VSA / Data-Free adapter. Each new process loads them and compiles the fast inference path; warmup and measured generations in one process reuse that work. For other supported platforms, start with the FastVideo installation guide.
For guided environment setup, use FastVideo’s
agent-guided installation
and ask the agent to finish with the FastH3-specific install command below.
For a manual install, first install
uv:
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
uv venv --python 3.12 --seed
source .venv/bin/activate
UV_TORCH_BACKEND=cu130 uv pip install \
--no-sources-package fastvideo-kernel \
-e ".[fasth3]"
The --no-sources-package option installs the published
fastvideo-kernel wheel, which includes the B200 sm100a VSA kernel, instead
of compiling the kernel locally. Base H3 and the FastH3 adapters are public, so
no Hugging Face login is required. Review the MiniMax H3 Community License
before use.
Use the pre-extracted
VSA / Data-Free LoRA
with the
FastVideo launcher.
The launcher downloads the adapter, loads it on top of MiniMaxAI/MiniMax-H3,
and enables the required VSA-H3 backend and tile-64 kernel.
Try VSA / Data-Free:
PROMPT='integrated_multimodal_description: A red fox runs through fresh snow at dawn. overall_soundscape: Fast pawsteps in snow, winter wind, and distant birds.'
bash examples/inference/basic/run_fasth3_lora_preview_vsa_datafree.sh \
--prompt "$PROMPT" \
--no-warmup \
--repeats 1
On other multi-GPU CUDA systems, add
--no-replicated-dit --vsa-kernel triton --no-fa4. The value of --num-gpus
must divide H3’s 56 attention heads.
Do not remove --vsa or run this LoRA through the dense path. The other three
checkpoints remain available in the ablation table above.
The launcher uses the shared
basic_fasth3_lora_preview.py
runner. Its defaults enable FastVideo’s H3 fusions, regional full-graph DiT
compile, FA4 for dense attention, compiled sequence-parallel VAE decoding,
replicated DiT weights, and pinned CPU offload. VSA variants also select 90%
sparsity, tile size 64, and the sm100a block-sparse kernel. Five scheduler
points mean exactly four DiT calls. Guidance stays at 1.0, matching training.
By default, the runner performs one compile warmup and then saves three measured
generations. Use --no-warmup --repeats 1 for one clip, as above. Use
--num-frames 243 for about 10 seconds or --num-frames 345 for about 15
seconds.
Adapter strength is also available:
bash examples/inference/basic/run_fasth3_lora_preview_vsa_datafree.sh \
--prompt "$PROMPT" \
--lora-strength 0.5 \
--no-warmup \
--repeats 1
Strength 1.0 applies the published rank-64 adapter at its trained scale.
Strength 0 removes its weight deltas but keeps the selected attention backend,
so a VSA run remains sparse. FastH3 adapters also contain exact parameter deltas
and, for VSA, compression-gate weights. FastVideo therefore applies the adapter
while building the pipeline and rejects unsafe runtime switching. Create a new
generator when changing variants or strength.
For a local adapter file or custom flags, call the shared Python runner directly.
For a JSONL prompt set, use
minimax_h3_lora_inference.py.
What’s coming next
Preview v1 is an early checkpoint family. FastVideo’s next priorities are:
1. Improve motion and offer an eight-step option
We will finish the four-step low-noise A/B and 8-step model post-training. We will then test stronger final-step training, more low-noise critic samples, motion-sensitive losses, and learned timestep placement. The four-step model targets minimum latency; eight steps may be a better quality setting. We will release an 8-step checkpoint only after a matched comparison.
2. Add FL2VA and Ref2VA
T2VA is only one H3 workflow. FL2VA uses the base transformer but needs new conditioning training. Ref2VA uses the separate reference transformer, so FastVideo must distill it separately. We will evaluate mixed reference types, long clips, and reference fidelity before releasing either workflow.
3. Apply Parallel Decoding Distillation methods to H3
We will continue to collaborate with NVIDIA’s FastGen team to try their new PDD algorithm and obtain the highest quality step-distilled H3 possible!
4. Make H3 easier to run for local AI and to extend
At four steps, encoding, VAE decode, audio, and file output become a larger share of latency. FastVideo will keep improving VAE compilation and parallelism, sparse compilation, portable kernels, cold start, and multi-GPU serving. We are also exploring FP8 and NVFP4 variants.
Help us test more hardware
Our published latency numbers use NVIDIA B200 GPUs because that is our controlled benchmark platform, not because the checkpoints require B200. The weights are hardware-independent and can run anywhere with enough memory and a compatible H3 runtime. The VSA checkpoints additionally need a compatible FastVideo VSA kernel.
We are preparing optimized FastVideo recipes for NVIDIA RTX GPUs, NVIDIA DGX Spark, and Apple Silicon through MLX. Stay tuned, and help us test hardware we do not have locally.
Share results, unsupported hardware, regressions, and new ideas on GitHub or in the FastVideo Slack. Start with the recommended VSA / Data-Free checkpoint and its optimized FastVideo launcher. Use the full collection for the ablations, and share your videos with us!
Acknowledgements
FastVideo FastH3 builds on Minimax H3. We thank the Minimax team for releasing its weights and code.
We thank the NVIDIA Enterprise Products team (Pengcheng Li, Cliff Woolley) for the amazing work on the Video Sparse Attention (VSA) kernel used by FastH3!
We thank Nuva Lab for bringing production grounding to FastH3 through its experience with real-world creative video-agent workloads. Its production-aligned post-training insights help bridge open-source research to practical data-assisted distillation for commercial video workflows, with Omni Ref as the next focus.
We thank the NVIDIA FastGen (Julius Berner, Chao Liu, Arash Vahdat) team for the DMD2 framework and H3 reference experiment that helped us align the score clock, modality shifts, and backward simulation.
We also thank MiniMax for releasing H3-Base, and the vLLM project, NVIDIA, and MBZUAI for their continued sponsorship and support of FastVideo.
Citations
Method background: the DMD2 paper, the VSA paper, our earlier FastWan sparse-distillation blog, and the FastWan-QAD blog. If you build on FastH3, please cite this release, DMD2, VSA, and FastVideo.
@misc{fastvideo_fasth3_2026,
title = {FastH3 Preview v1: Four Open-Weight H3 Models in Four Steps},
author = {FastVideo Team},
year = {2026},
howpublished = {\url{https://haoailab.com/blogs/fasth3-preview/}},
}
@misc{yin2024improved,
title = {Improved Distribution Matching Distillation for Fast Image Synthesis},
author = {Tianwei Yin and Michaël Gharbi and Taesung Park and Richard Zhang and Eli Shechtman and Fredo Durand and William T. Freeman},
year = {2024},
eprint = {2405.14867},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
}
@article{zhang2025vsa,
title = {VSA: Faster Video Diffusion with Trainable Sparse Attention},
author = {Peiyuan Zhang and Yongqi Chen and Haofeng Huang and Will Lin and Zhengzhong Liu and Ion Stoica and Eric Xing and Hao Zhang},
journal = {arXiv preprint arXiv:2505.13389},
year = {2025},
}
@software{fastvideo2024,
title = {FastVideo: A Unified Framework for Accelerated Video Generation},
author = {The FastVideo Team},
url = {https://github.com/hao-ai-lab/FastVideo},
year = {2024},
}