minimax_h3_vsa_fp4 ¶
Inference fast path: VSA-H3 attention on the block-sparse FP4 kernel.
Opt-in with FASTVIDEO_H3_VSA_FP4=1 (no-grad, single sequence-parallel rank, fastvideo-kernel built with attn_qat_infer). The selection is VSA-H3's own: tile pooling, top-k block mask and the gated compression branch are unchanged; only the block-sparse attention itself runs on SageAttention3's FP4 kernel (BF16 Triton otherwise), with 64-token tiles carried by quadrant masks on the kernel's 128x128 blocks.
The attention input is gathered into tile order once per block (one hidden_size-wide pass; pad rows stay zero, so q/k/v pad rows are exactly zero through the bias-free projections, RMSNorm and RoPE). That replaces the generic path's concat, four tile scatters and three transposes, and lets q, k and v share one activation quantization. The output returns to packed order with one gather before to_out.
Functions:¶
fastvideo.models.dits.minimax_h3_vsa_fp4.vsa_fp4_attention ¶
vsa_fp4_attention(attn: Any, hidden_states: Tensor, rotary_emb: tuple[Tensor, Tensor], meta: MiniMaxH3VSAMetadata, use_fused_rope: bool) -> Tensor
Attention core for MiniMaxH3Attention; returns the pre-to_out [B, L, H*D].
Source code in fastvideo/models/dits/minimax_h3_vsa_fp4.py
fastvideo.models.dits.minimax_h3_vsa_fp4.vsa_fp4_attention_sp ¶
vsa_fp4_attention_sp(attn: Any, hidden_states: Tensor, rotary_emb: tuple[Tensor, Tensor], meta: MiniMaxH3VSAMetadata, use_fused_rope: bool, sp_group: Any) -> Tensor
Ulysses-SP attention core on local sequence rows [1, rows, C]; returns pre-to_out [1, rows, H*D].
Source code in fastvideo/models/dits/minimax_h3_vsa_fp4.py
fastvideo.models.dits.minimax_h3_vsa_fp4.vsa_tile_first_attention ¶
vsa_tile_first_attention(attn: Any, hidden_states: Tensor, rotary_emb: tuple[Tensor, Tensor], meta: MiniMaxH3VSAMetadata, use_fused_rope: bool) -> Tensor
Single-rank BF16 VSA with one input scatter instead of a Q/K/V/gate stack.
The existing backend computes the same tile-64 mask, valid-key handling, fine attention and compression branch. Bias-free projections keep pad rows zero. This path is inference-only and keeps the checkpoint layout.