Skip to content

kandinsky6_sr

Kandinsky6 SR latent upscaler (LU): a bank of convolutional upsamplers of scaled KVAE latents.

Each bank entry is a cascade of two 2x stages over [B, C, T, h, w] latents::

x4: input_proj -> pre_blocks -> upsample_1 -> mid_blocks -> upsample_2 -> post_blocks -> output_proj
x2: mid_input_proj -> x2_branch (adapter -> finisher -> private mid stage -> private second stage)

Every norm is an RMSNorm FiLM-modulated by the input latent itself (zq, nearest-resized to the feature grid), the 3x3x3 convs pad time by repeating the edge frame, and the 2x upsample is Conv1x1(up + Conv(1,3,3)(up)) with up the nearest-neighbour 2x resize. The bank holds one entry per served scale under _models.<index> in config.scales order.

Module names match the re-keyed Diffusers checkpoint; no load-time compatibility hook is required.

Classes

fastvideo.models.upsamplers.kandinsky6_sr.Kandinsky6SRLatentUpscaler

Kandinsky6SRLatentUpscaler(config: Kandinsky6SRLatentUpscalerEntryConfig)

Bases: Module

One bank entry: [B, C, T, h, w] -> [B, C, T, s*h, s*w] for s = 4, or 2 with enable_x2_entry.

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, config: Kandinsky6SRLatentUpscalerEntryConfig) -> None:
    super().__init__()
    c = config.in_channels
    w1, w2, w3 = config.stage_channels
    er = config.expand_ratio
    self.input_proj = _stem(c, w1)
    # The 2x deep-supervision head of training: part of the checkpoint, never run at inference.
    self.mid_output_head = nn.Sequential(RMSNorm(w2), nn.SiLU(), ReplicateTimeConv3d(w2, c, kernel_size=3))
    self.output_proj = OutputHead(w3, c, zq_dim=c)
    self.upsample_1 = PXSUpsample(w1)
    self.upsample_2 = PXSUpsample(w2)
    self.pre_blocks = _residual_stack(config.num_pre_blocks, w1, w1, er, c)
    self.mid_blocks = _residual_stack(config.num_mid_blocks, w1, w2, er, c)
    self.post_blocks = _residual_stack(config.num_post_blocks, w2, w3, er, c)
    self.mid_input_proj: nn.Sequential | None = None
    self.x2_branch: X2Branch | None = None
    if config.enable_x2_entry:
        self.mid_input_proj = _stem(c, config.hidden_channels)
        self.x2_branch = X2Branch(config)

fastvideo.models.upsamplers.kandinsky6_sr.Kandinsky6SRLatentUpscalerBank

Kandinsky6SRLatentUpscalerBank(config: Kandinsky6SRLatentUpscalerConfig)

Bases: Module

The x2 / x4 latent-upscaler bank of a Kandinsky6 SR bundle (latent_upscaler/).

forward(z, scale) upsamples an already scaled (latent * scaling_factor) KVAE latent [B, C, T, h, w] by scale in H and W with the entry serving that scale.

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, config: Kandinsky6SRLatentUpscalerConfig) -> None:
    super().__init__()
    if not config.models:
        raise ValueError("latent upscaler config declares no `models` entries")
    self.config = config
    self.scaling_factor = config.scaling_factor
    entries = {entry.target_scale: entry for entry in config.models}
    if set(entries) != set(config.scales):
        raise ValueError(f"Latent upscaler models {tuple(entries)} must match scales={config.scales}")
    self._models = nn.ModuleList([Kandinsky6SRLatentUpscaler(entries[scale]) for scale in config.scales])

Attributes

fastvideo.models.upsamplers.kandinsky6_sr.Kandinsky6SRLatentUpscalerBank.scales property
scales: tuple[int, ...]

The upscale factors this bank serves, ascending.

fastvideo.models.upsamplers.kandinsky6_sr.ModulatedRMSNorm

ModulatedRMSNorm(dim: int, zq_dim: int)

Bases: Module

RMSNorm(x) * conv_y(zq) + conv_b(zq) with 1x1x1 convs of the conditioning latent zq.

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, dim: int, zq_dim: int) -> None:
    super().__init__()
    self.norm = RMSNorm(dim)
    self.conv_y = nn.Conv3d(zq_dim, dim, kernel_size=1)
    self.conv_b = nn.Conv3d(zq_dim, dim, kernel_size=1)

fastvideo.models.upsamplers.kandinsky6_sr.OutputHead

OutputHead(channels: int, out_channels: int, zq_dim: int)

Bases: Module

Named norm/activation/conv projection used by current Diffusers weights.

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, channels: int, out_channels: int, zq_dim: int) -> None:
    super().__init__()
    self.norm = ModulatedRMSNorm(channels, zq_dim)
    self.activation = nn.SiLU()
    self.conv = ReplicateTimeConv3d(channels, out_channels, kernel_size=3)

fastvideo.models.upsamplers.kandinsky6_sr.PXSUpsample

PXSUpsample(channels: int)

Bases: Module

2x spatial upsample linear(up + spatial_conv(up)) with up = nearest_2x(x).

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, channels: int) -> None:
    super().__init__()
    self.spatial_conv = nn.Conv3d(channels, channels, kernel_size=(1, 3, 3), padding=(0, 1, 1))
    self.linear = nn.Conv3d(channels, channels, kernel_size=1)

fastvideo.models.upsamplers.kandinsky6_sr.RMSNorm

RMSNorm(dim: int)

Bases: Module

Channel RMS norm of [B, C, T, H, W] features with a learnable per-channel gain.

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, dim: int) -> None:
    super().__init__()
    self.scale = dim**0.5
    self.gamma = nn.Parameter(torch.ones(dim, 1, 1, 1))

fastvideo.models.upsamplers.kandinsky6_sr.ReplicateTimeConv3d

ReplicateTimeConv3d(in_channels: int, out_channels: int, kernel_size: int)

Bases: Conv3d

Conv3d with 'same' padding that repeats the edge frame along T and zero-pads H and W.

nn.Conv3d has a single padding mode for all axes, so T is padded explicitly before the convolution.

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, in_channels: int, out_channels: int, kernel_size: int) -> None:
    pad = kernel_size // 2
    super().__init__(in_channels, out_channels, kernel_size, padding=(0, pad, pad))
    self.temporal_pad = pad

fastvideo.models.upsamplers.kandinsky6_sr.ResidualBlock

ResidualBlock(in_channels: int, out_channels: int, mid_channels: int, zq_dim: int)

Bases: Module

Pre-activation block norm1 -> SiLU -> conv1 -> norm2 -> SiLU -> conv2 plus a (1x1x1 if narrowing) skip.

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, in_channels: int, out_channels: int, mid_channels: int, zq_dim: int) -> None:
    super().__init__()
    self.norm1 = ModulatedRMSNorm(in_channels, zq_dim)
    self.conv1 = ReplicateTimeConv3d(in_channels, mid_channels, kernel_size=3)
    self.norm2 = ModulatedRMSNorm(mid_channels, zq_dim)
    self.conv2 = ReplicateTimeConv3d(mid_channels, out_channels, kernel_size=3)
    self.shortcut: nn.Module = (nn.Identity() if in_channels == out_channels else ReplicateTimeConv3d(
        in_channels, out_channels, kernel_size=1))

fastvideo.models.upsamplers.kandinsky6_sr.X2Branch

Bases: Module

Weights exclusive to the x2 path: a residual adapter at the input grid, a finisher, and private copies of the mid stage and second stage (so the x2 path shares no weights with the x4 path).

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, config: Kandinsky6SRLatentUpscalerEntryConfig) -> None:
    super().__init__()
    c = config.in_channels
    w1, w2, w3 = config.stage_channels
    er = config.expand_ratio
    self.adapter = _residual_stack(config.x2_adapter_blocks, w1, w1, er, c)
    self.finisher = X2Finisher(w1)
    self.upsample = PXSUpsample(w2)
    self.blocks = _residual_stack(config.num_post_blocks, w2, w3, er, c)
    self.output_proj = OutputHead(w3, c, zq_dim=c)
    self.mid_blocks = _residual_stack(config.num_mid_blocks, w1, w2, er, c)

fastvideo.models.upsamplers.kandinsky6_sr.X2Finisher

X2Finisher(channels: int)

Bases: Module

The PXSUpsample conv pair applied on the unchanged grid (no resize).

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def __init__(self, channels: int) -> None:
    super().__init__()
    self.linear = nn.Conv3d(channels, channels, kernel_size=1)
    self.spatial_conv = nn.Conv3d(channels, channels, kernel_size=(1, 3, 3), padding=(0, 1, 1))

Functions:

fastvideo.models.upsamplers.kandinsky6_sr.nearest_2x

nearest_2x(x: Tensor) -> Tensor

Nearest-neighbour 2x resize of H and W of [B, C, T, H, W] (T folded into the batch).

Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
def nearest_2x(x: Tensor) -> Tensor:
    """Nearest-neighbour 2x resize of H and W of ``[B, C, T, H, W]`` (T folded into the batch)."""
    b, c, t, h, w = x.shape
    x = x.permute(0, 2, 1, 3, 4).reshape(b * t, c, h, w)
    x = F.interpolate(x, scale_factor=2, mode="nearest")
    return x.reshape(b, t, c, 2 * h, 2 * w).permute(0, 2, 1, 3, 4)