Skip to content

vggsound

Raw VGGSound metadata adapter for V2A preprocessing.

Classes

fastvideo.dataset.vggsound.VGGSoundDataset

VGGSoundDataset(root: str | Path, *, split: str = 'train', metadata_path: str | Path | None = None)

Bases: Dataset

Map VGGSound metadata rows to local MP4 paths and captions.

Loie/VGGSound stores clips as <youtube-id>_<start:06d>.mp4. The downloaded tar archives must be extracted before random-access GPU preprocessing; repeatedly seeking inside gzip archives is prohibitively expensive for a shuffled training dataset.

Source code in fastvideo/dataset/vggsound.py
def __init__(
    self,
    root: str | Path,
    *,
    split: str = "train",
    metadata_path: str | Path | None = None,
) -> None:
    super().__init__()
    self.root = Path(root).expanduser().resolve()
    metadata = (Path(metadata_path).expanduser().resolve() if metadata_path is not None else self.root /
                "vggsound.csv")
    if not metadata.is_file():
        raise FileNotFoundError(f"VGGSound metadata file does not exist: {metadata}")

    candidates = (self.root / "videos", self.root / "video", self.root)
    self.video_root = next((path for path in candidates if path.is_dir()), self.root)
    self.samples: list[tuple[str, str, Path]] = []
    if metadata.suffix.lower() == ".tsv":
        self._read_caption_manifest(metadata, split=split)
    else:
        self._read_vggsound_csv(metadata, split=split)

    if not self.samples:
        raise ValueError(f"No VGGSound samples found for split {split!r} in {metadata}")

Functions: