tiling ¶
Tile planning and stitching for Kandinsky6 video super-resolution.
The SR DiT was trained on fixed base resolutions, so a source frame is cut into overlapping tiles whose upscaled size is one of those bases, each tile is super-resolved independently and the decoded tiles are blended back with a Hanning window. Tile positions are aligned to the VAE spatial factor so the same grid addresses both the latent and the decoded pixels.
Classes¶
fastvideo.pipelines.basic.kandinsky6_sr.tiling.TileGrid dataclass ¶
Tile size and per-axis tile start positions, in pixels of the (pre-upscaled) source frame.
Functions:¶
fastvideo.pipelines.basic.kandinsky6_sr.tiling.plan_tiles ¶
plan_tiles(height: int, width: int, visual_size: int, scale: int, min_overlap: float, spatial_factor: int) -> TileGrid
Tile grid of a height x width source frame for an integer tiling scale.
The base resolution is the trained one closest in aspect ratio; each tile is base / scale source pixels, so its upscaled size is exactly the base resolution the DiT was trained on.
Source code in fastvideo/pipelines/basic/kandinsky6_sr/tiling.py
fastvideo.pipelines.basic.kandinsky6_sr.tiling.pre_upscale ¶
Bilinear pixel upscale of a [T, C, H, W] uint8 video, target size rounded to spatial_multiple.
Source code in fastvideo/pipelines/basic/kandinsky6_sr/tiling.py
fastvideo.pipelines.basic.kandinsky6_sr.tiling.resolve_scale ¶
Requested total upscale -> (tiling_scale, pixel_pre_upscale).
The latent upscaler only has x2 and x4 entries; the fractional x2.25 is an x1.125 pixel pre-upscale followed by x2 tiling.
Source code in fastvideo/pipelines/basic/kandinsky6_sr/tiling.py
fastvideo.pipelines.basic.kandinsky6_sr.tiling.stitch_tiles ¶
stitch_tiles(tiles: list[Tensor], grid: TileGrid, height: int, width: int, scale: int, frame_chunk: int = 16) -> Tensor
Blend row-major float [0, 255] tiles [C, T, tile_h * scale, tile_w * scale] into a single uint8 [C, T, H * scale, W * scale] video, quantizing once at the very end (not per tile): pre-quantizing each tile to uint8 before blending would blend already-rounded values instead of the continuous pixel values, compounding rounding error at every overlap -- matches the diffusers reference's decode_latents/ __call__, which accumulates video_acc += tile * window in float and rounds only the final result.
Each tile is weighted by a 2D Hanning window without zero endpoints (so border pixels covered by a single tile keep a non-zero weight) and the accumulation is normalised per pixel, which also handles uneven overlaps. Frames are blended in chunks to bound host memory; frames are independent, so chunking does not change the result.