exact ¶
Bit-exact Triton replacements for MiniMax-H3's eager RoPE, AdaLN and SwiGLU.
Each kernel performs exactly the BF16-rounded operation sequence the eager PyTorch expression performs, in the same order, so its output equals the eager output bit for bit. That is the difference from the Sol-Engine fusions in this package (FASTVIDEO_MINIMAX_H3_FUSIONS), which keep the whole expression in FP32 and round once at the store: faster still, but not the eager bits.
The kernels gain by fusing the eager expression's separate elementwise launches (and the index_select gathers of the AdaLN tables) into one pass over the activations, not by changing the arithmetic. Two details keep the rounding exact:
- every FP32 multiply, add and divide is issued as explicit round-to-nearest PTX (
mul.rn/add.rn/div.rn), so the compiler cannot contract a multiply and an add into one FMA, or fold the FP32 result and the BF16 downcast into one BF16 instruction, either of which rounds once where eager rounds twice; - every intermediate eager materializes as a BF16 tensor is rounded to BF16 (round-to-nearest-even) at the same point.
Selected with FASTVIDEO_MINIMAX_H3_EXACT_KERNELS. Every entry point has a supports_* predicate; callers fall back to the eager expression for inputs outside it.
Functions:¶
fastvideo.models.dits.minimax_h3_fusions.exact.gate_residual ¶
hidden + gate.index_select(0, indices) * update.
Source code in fastvideo/models/dits/minimax_h3_fusions/exact.py
fastvideo.models.dits.minimax_h3_fusions.exact.modulate ¶
normed * (1.0 + scale.index_select(0, indices)) + shift.index_select(0, indices).
Source code in fastvideo/models/dits/minimax_h3_fusions/exact.py
fastvideo.models.dits.minimax_h3_fusions.exact.rope_prefix ¶
rope_prefix(hidden_states: Tensor, rotary_emb: tuple[Tensor, Tensor]) -> Tensor
Eager MiniMaxH3Attention._apply_rotary_emb for contiguous BF16 [1, S, H, D].
Source code in fastvideo/models/dits/minimax_h3_fusions/exact.py
fastvideo.models.dits.minimax_h3_fusions.exact.supports_rope ¶
Whether rope_prefix reproduces eager _apply_rotary_emb on these inputs.
Source code in fastvideo/models/dits/minimax_h3_fusions/exact.py
fastvideo.models.dits.minimax_h3_fusions.exact.supports_rowwise ¶
supports_rowwise(*tensors: Tensor) -> bool
Whether modulate/gate_residual/swiglu reproduce eager on these tensors.
Source code in fastvideo/models/dits/minimax_h3_fusions/exact.py
fastvideo.models.dits.minimax_h3_fusions.exact.swiglu ¶
value, gate = packed.chunk(2, -1); value * F.silu(gate).