Skip to content

minimax_h3_qwen3_vl

Native Qwen3-VL conditioner used by MiniMax H3.

Classes

fastvideo.models.encoders.minimax_h3_qwen3_vl.MiniMaxH3Qwen3VLConditioner

MiniMaxH3Qwen3VLConditioner(config: MiniMaxH3Qwen3VLConfig)

Bases: TextEncoder[Tensor]

H3 conditioner returning the unnormalized layer-50 hidden tensor.

Source code in fastvideo/models/encoders/minimax_h3_qwen3_vl.py
def __init__(self, config: MiniMaxH3Qwen3VLConfig) -> None:
    super().__init__(config)
    self.visual = MiniMaxH3Qwen3VLVisionModel(config)
    self.language_model = MiniMaxH3Qwen3VLLanguageModel(config)

Methods:

fastvideo.models.encoders.minimax_h3_qwen3_vl.MiniMaxH3Qwen3VLConditioner.prepare_layerwise_offload
prepare_layerwise_offload(device: device) -> None

Stream language layers for text-only CUDA inference, retaining embeddings on CPU.

Source code in fastvideo/models/encoders/minimax_h3_qwen3_vl.py
def prepare_layerwise_offload(self, device: torch.device) -> None:
    """Stream language layers for text-only CUDA inference, retaining embeddings on CPU."""
    if getattr(self, "_h3_encoder_layerwise_device", None) is not None:
        return
    if device.type != "cuda":
        raise ValueError("Layerwise H3 encoder requires CUDA")
    from fastvideo.distributed import get_tp_world_size
    from fastvideo.hooks.layerwise_offload import enable_layerwise_offload

    if get_tp_world_size() != 1:
        raise ValueError("Layerwise H3 encoder requires tensor parallel size 1")
    self.to("cpu")
    self.language_model.rotary_emb.to(device)
    if self.language_model.norm is not None:
        self.language_model.norm.to(device)
    enable_layerwise_offload(self.language_model, resident_blocks=0, cyclic=False)
    self._h3_encoder_layerwise_device = device

fastvideo.models.encoders.minimax_h3_qwen3_vl.MiniMaxH3Qwen3VLTextRotaryEmbedding

MiniMaxH3Qwen3VLTextRotaryEmbedding(config: MiniMaxH3Qwen3VLConfig)

Bases: Module

Shared Qwen3-VL interleaved temporal/height/width rotary embedding.

Source code in fastvideo/models/encoders/minimax_h3_qwen3_vl.py
def __init__(self, config: MiniMaxH3Qwen3VLConfig) -> None:
    super().__init__()
    head_dim = config.head_dim
    inv_freq = 1.0 / (config.rope_theta**(torch.arange(0, head_dim, 2, dtype=torch.float32) / head_dim))
    self.register_buffer("inv_freq", inv_freq, persistent=False)
    self.mrope_section = tuple(config.mrope_section)

fastvideo.models.encoders.minimax_h3_qwen3_vl.MiniMaxH3SerializedFP8Config

MiniMaxH3SerializedFP8Config(weight_block_size: tuple[int, int])

Bases: QuantizationConfig

Serialized 128x128 block-FP8 contract for the H3 text encoder.

Source code in fastvideo/models/encoders/minimax_h3_checkpoint_fp8.py
def __init__(self, weight_block_size: tuple[int, int]) -> None:
    super().__init__()
    if weight_block_size != (128, 128):
        raise ValueError("MiniMax-H3 serialized FP8 requires weight_block_size=[128, 128], "
                         f"got {list(weight_block_size)}")
    self.weight_block_size = weight_block_size
    self.is_checkpoint_fp8_serialized = True
    self.activation_scheme = "dynamic"

fastvideo.models.encoders.minimax_h3_qwen3_vl.MiniMaxH3SerializedNVFP4Config

MiniMaxH3SerializedNVFP4Config(bf16_projections: tuple[str, ...] = (), pre_quant_scale: bool = False)

Bases: QuantizationConfig

Serialized 16-group NVFP4 contract for the H3 text encoder.

The group size and scale layout are fixed by the loader's parameter shapes, so the only state is which projection kinds the checkpoint kept in bf16.

Source code in fastvideo/models/encoders/minimax_h3_checkpoint_nvfp4.py
def __init__(self, bf16_projections: tuple[str, ...] = (), pre_quant_scale: bool = False) -> None:
    super().__init__()
    unknown = sorted(set(bf16_projections) - set(LANGUAGE_PROJECTIONS))
    if unknown:
        raise ValueError(f"MiniMax-H3 serialized NVFP4 cannot keep unknown projection kinds {unknown} in bf16; "
                         f"choose from {LANGUAGE_PROJECTIONS}")
    # Projection kinds the checkpoint kept in bf16 in every language layer,
    # e.g. ("mlp.down_proj",). Those linears load a plain weight.
    self.bf16_projections = tuple(bf16_projections)
    self.pre_quant_scale = pre_quant_scale

Functions: