Skip to content

preprocess_hunyuan15_overfit

Preprocess HunyuanVideo 1.5 overfit data into parquet format.

This script writes both the Qwen (text_embedding_*) and the ByT5 (text_embedding_2_*) embeddings, laid out in pyarrow_schema_t2v_dual_text -- the schema Hunyuan15Model trains on. The generic CLI preprocess workflow (PreprocessPipeline_T2V) only runs the primary text encoder, so HunyuanVideo 1.5 data must come from this script to keep glyph conditioning.

Classes

Functions:

fastvideo.pipelines.preprocess.preprocess_hunyuan15_overfit.build_parquet_table

build_parquet_table(records: list[dict[str, Any]]) -> Table

Lay the records out in the dual-text schema Hunyuan15Model reads.

The base pyarrow_schema_t2v has no ByT5 columns, so writing these records with it would drop the glyph stream without an error.

Source code in fastvideo/pipelines/preprocess/preprocess_hunyuan15_overfit.py
def build_parquet_table(records: list[dict[str, Any]]) -> pa.Table:
    """Lay the records out in the dual-text schema ``Hunyuan15Model`` reads.

    The base ``pyarrow_schema_t2v`` has no ByT5 columns, so writing these
    records with it would drop the glyph stream without an error.
    """
    columns = {field.name: [record.get(field.name) for record in records] for field in pyarrow_schema_t2v_dual_text}
    return pa.Table.from_pydict(columns, schema=pyarrow_schema_t2v_dual_text)