preprocess_minimax_h3_ref2va ¶
Precompute a raw MiniMax H3 Ref2VA manifest into training Parquet shards.
Classes¶
Functions:¶
fastvideo.pipelines.preprocess.preprocess_minimax_h3_ref2va.encode_ref2va_conditioning ¶
encode_ref2va_conditioning(caption: str, references: list[MiniMaxH3PreparedReference], model_path: Path, model_index: dict[str, Any], fastvideo_args: FastVideoArgs) -> tuple[Tensor, Tensor]
Encode the exact ordered presentation without padding or truncation.
Source code in fastvideo/pipelines/preprocess/preprocess_minimax_h3_ref2va.py
fastvideo.pipelines.preprocess.preprocess_minimax_h3_ref2va.encode_ref_audio_anchor ¶
encode_ref_audio_anchor(references: list[MiniMaxH3PreparedReference], model_path: Path, model_index: dict[str, Any], fastvideo_args: FastVideoArgs) -> Tensor
Encode clean channel-major audio anchors in ordered-reference order.
Source code in fastvideo/pipelines/preprocess/preprocess_minimax_h3_ref2va.py
fastvideo.pipelines.preprocess.preprocess_minimax_h3_ref2va.encode_ref_visual_anchor ¶
encode_ref_visual_anchor(references: list[MiniMaxH3PreparedReference], model_path: Path, model_index: dict[str, Any], fastvideo_args: FastVideoArgs, patch_size: tuple[int, int, int]) -> Tensor
Encode and cache official 0.999-noised ordered visual condition rows.