encoder_split ¶
Component-level pipeline parallel for MiniMax H3: dedicated text-encoder ranks.
Enabled with FASTVIDEO_H3_ENCODER_SPLIT=1 (or FastVideoArgs.h3_encoder_split) on the Ray backend. The workers on the first FASTVIDEO_H3_ENCODER_NODES nodes (ranks 0..n-1) load only the Qwen3-VL conditioner, run the condition stages, and NCCL-broadcast prompt_embeds to the remaining denoising ranks over the world group. The denoising ranks never load the text encoder, so their peak resident memory drops by the encoder size (~48 GiB on GB10), which is what makes 720p fit in 128 GiB of unified memory.
The denoising ranks form the sequence-parallel group (sp_size is recomputed per rank as num_gpus - h3_encoder_workers); encoder ranks hold singleton SP/DP groups so every rank still belongs to exactly one group of every kind, keeping GroupCoordinator construction uniform.
Constraints and semantics: - The denoise group runs SP, so its size must divide the DiT attention head count (56 in the released FastH3 checkpoints). With 8 single-GPU nodes the legal encoder-node counts are therefore 1, 4, 6, 7 (SP = 7, 4, 2, 1). This is validated early — in the Ray executor against the placement-derived worker count, and again here before any weights load; the driver-side check in FastVideoArgs.check_fastvideo_args only rejects a split that leaves no denoise rank, because it sees nodes where the executor sees workers. - With tp_size > 1 the encoder group must also be a multiple of tp_size: TP groups are built from consecutive ranks and the Qwen3-VL conditioner is TP-sharded, so a group straddling the encoder/denoise boundary would all-reduce across ranks that never run the encoder. - One request packs a single prompt presentation, so world rank 0 is the only encoder that computes; encoder ranks 1..n-1 mirror the broadcast collectives and stay idle. Extra encoder nodes are reserved capacity (e.g. for future multi-prompt parallel encoding), not a speedup for single requests.
Classes¶
Functions:¶
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_broadcast_condition ¶
h3_broadcast_condition(batch: ForwardBatch, fastvideo_args: FastVideoArgs) -> None
Send the conditioning output to every rank over the world NCCL group.
Only world rank 0 encodes a request (one prompt presentation per request), so this runs on rank 0 as the source. Other encoder ranks mirror the same collectives through h3_receive_condition and drop the payload; skipping them entirely would desynchronize the world-group broadcasts for the denoise ranks.
Source code in fastvideo/pipelines/basic/minimax_h3/encoder_split.py
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_denoise_device_mesh ¶
h3_denoise_device_mesh(fastvideo_args: FastVideoArgs) -> DeviceMesh
FSDP mesh covering exactly the denoise ranks.
init_device_mesh always lays a mesh over ranks 0..numel-1, but the encoder group owns the leading ranks, so the mesh is built from an explicit rank map. Only the denoise ranks build it: the DiT exists nowhere else, and the groups it creates are local to those ranks.
Source code in fastvideo/pipelines/basic/minimax_h3/encoder_split.py
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_encoder_worker_count ¶
h3_encoder_worker_count(fastvideo_args: FastVideoArgs) -> int
Number of world ranks reserved for the encoder group.
Source code in fastvideo/pipelines/basic/minimax_h3/encoder_split.py
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_is_output_worker ¶
h3_is_output_worker(fastvideo_args: FastVideoArgs) -> bool
Whether this rank owns the decoded output (see h3_output_worker_rank).
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_is_primary_encoder_worker ¶
h3_is_primary_encoder_worker(fastvideo_args: FastVideoArgs) -> bool
Rank 0 encodes every request; ranks 1..n-1 only mirror the broadcast.
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_output_worker_rank ¶
h3_output_worker_rank(fastvideo_args: FastVideoArgs) -> int
World rank that owns the decoded output: world rank 0 normally, first denoise rank in split.
Source code in fastvideo/pipelines/basic/minimax_h3/encoder_split.py
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_prepare_split_worker_parallelism ¶
h3_prepare_split_worker_parallelism(fastvideo_args: FastVideoArgs, rank: int, world_size: int) -> tuple[int, list[list[int]], list[list[int]]]
Resolve per-rank SP size and the custom SP/DP group layout for split mode.
Called by the worker before maybe_init_distributed_environment_and_model_parallel, and stamped onto the worker-local fastvideo_args so the consistency check in ComposedPipelineBase.__init__ sees the same per-rank SP size.
Source code in fastvideo/pipelines/basic/minimax_h3/encoder_split.py
fastvideo.pipelines.basic.minimax_h3.encoder_split.h3_receive_condition ¶
h3_receive_condition(batch: ForwardBatch, fastvideo_args: FastVideoArgs) -> None
Materialize the encoder worker's conditioning output on a denoise rank.