Skip to content

preprocess_workflow_v2a

Shared raw-media workflow for video-to-audio feature preprocessing.

Classes

fastvideo.workflow.preprocess.preprocess_workflow_v2a.PreprocessWorkflowV2A

PreprocessWorkflowV2A(fastvideo_args: FastVideoArgs)

Bases: PreprocessWorkflow

Run model-specific V2A encoders with shared data/cache orchestration.

Source code in fastvideo/workflow/workflow_base.py
def __init__(self, fastvideo_args: FastVideoArgs):
    """
    Initialize the workflow with configuration arguments.

    Args:
        fastvideo_args: Configuration object containing all parameters
                      needed for workflow and pipeline setup.
    """
    self.fastvideo_args = fastvideo_args

    # TODO: pipeline_config should be: dict[str, PipelineConfig]
    # pipeline_type should be included in the PipelineConfig
    # pipeline_config[pipeline_name] = (pipeline_type, fastvideo_args)
    self._pipeline_configs: dict[str, tuple[PipelineType, FastVideoArgs]] = {}
    self._pipelines: dict[str, ComposedPipelineBase] = {}
    self._components: dict[str, Any] = {}
    self.register_pipelines()
    self.register_components()

    self.prepare_system_environment()
    self.load_pipelines()

fastvideo.workflow.preprocess.preprocess_workflow_v2a.V2AForwardBatchBuilder

V2AForwardBatchBuilder(seed: int)

Translate standardized raw-media rows into a FastVideo batch.

Source code in fastvideo/workflow/preprocess/preprocess_workflow_v2a.py
def __init__(self, seed: int) -> None:
    self.seed = seed

Functions:

fastvideo.workflow.preprocess.preprocess_workflow_v2a.build_v2a_dataset

build_v2a_dataset(dataset_type: DatasetType, dataset_path: str, *, dataset_metadata_path: str = '', dataset_split: str = 'train')

Build a raw-media dataset for a V2A preprocessing workflow.

Source code in fastvideo/workflow/preprocess/preprocess_workflow_v2a.py
def build_v2a_dataset(
    dataset_type: DatasetType,
    dataset_path: str,
    *,
    dataset_metadata_path: str = "",
    dataset_split: str = "train",
):
    """Build a raw-media dataset for a V2A preprocessing workflow."""
    if isinstance(dataset_type, str):
        dataset_type = DatasetType.from_string(dataset_type)
    if dataset_type is DatasetType.VGGSOUND:
        metadata_path = dataset_metadata_path or None
        return VGGSoundDataset(
            dataset_path,
            split=dataset_split,
            metadata_path=metadata_path,
        )
    raise ValueError(f"V2A preprocessing currently supports dataset_type=vggsound; got {dataset_type.value!r}")