video_sparse_attn_h3_scatter ¶
Heads-first tile scatter for VSA-H3 (FASTVIDEO_H3_VSA_HEADS_FIRST_TILE).
MiniMaxH3VSAImpl.tile scatters packed [B, S, H, D] rows into a padded [B, S_pad, H, D] tile buffer, and the 64/128-token kernels then copy that buffer to [B, H, S_pad, D] with transpose(1, 2).contiguous() before every call, for query, key and value. This kernel scatters the rows straight into a [B, H, S_pad, D] buffer instead, and tile returns its transpose(1, 2) view: the same logical tensor, so every consumer reads the same values, and the kernels' .contiguous() becomes free. It only moves bytes, so the result is bit-identical.
Functions:¶
fastvideo.attention.backends.video_sparse_attn_h3_scatter.scatter_rows_heads_first ¶
buffer.transpose(1, 2)[:, slots] = x for x [B, S, H, D] and buffer [B, H, S_pad, D].
Source code in fastvideo/attention/backends/video_sparse_attn_h3_scatter.py
fastvideo.attention.backends.video_sparse_attn_h3_scatter.supports_heads_first_scatter ¶
supports_heads_first_scatter(x: Tensor, slots: Tensor) -> bool
Whether scatter_rows_heads_first handles x [B, S, H, D] and its row slots.