kandinsky6_sr ¶
Kandinsky6 SR latent upscaler (LU): a bank of convolutional upsamplers of scaled KVAE latents.
Each bank entry is a cascade of two 2x stages over [B, C, T, h, w] latents::
x4: input_proj -> pre_blocks -> upsample_1 -> mid_blocks -> upsample_2 -> post_blocks -> output_proj
x2: mid_input_proj -> x2_branch (adapter -> finisher -> private mid stage -> private second stage)
Every norm is an RMSNorm FiLM-modulated by the input latent itself (zq, nearest-resized to the feature grid), the 3x3x3 convs pad time by repeating the edge frame, and the 2x upsample is Conv1x1(up + Conv(1,3,3)(up)) with up the nearest-neighbour 2x resize. The bank holds one entry per served scale under _models.<index> in config.scales order.
Module names match the re-keyed Diffusers checkpoint; no load-time compatibility hook is required.
Classes¶
fastvideo.models.upsamplers.kandinsky6_sr.Kandinsky6SRLatentUpscaler ¶
Kandinsky6SRLatentUpscaler(config: Kandinsky6SRLatentUpscalerEntryConfig)
Bases: Module
One bank entry: [B, C, T, h, w] -> [B, C, T, s*h, s*w] for s = 4, or 2 with enable_x2_entry.
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.Kandinsky6SRLatentUpscalerBank ¶
Kandinsky6SRLatentUpscalerBank(config: Kandinsky6SRLatentUpscalerConfig)
Bases: Module
The x2 / x4 latent-upscaler bank of a Kandinsky6 SR bundle (latent_upscaler/).
forward(z, scale) upsamples an already scaled (latent * scaling_factor) KVAE latent [B, C, T, h, w] by scale in H and W with the entry serving that scale.
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.ModulatedRMSNorm ¶
Bases: Module
RMSNorm(x) * conv_y(zq) + conv_b(zq) with 1x1x1 convs of the conditioning latent zq.
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.OutputHead ¶
Bases: Module
Named norm/activation/conv projection used by current Diffusers weights.
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.PXSUpsample ¶
PXSUpsample(channels: int)
Bases: Module
2x spatial upsample linear(up + spatial_conv(up)) with up = nearest_2x(x).
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.ReplicateTimeConv3d ¶
Bases: Conv3d
Conv3d with 'same' padding that repeats the edge frame along T and zero-pads H and W.
nn.Conv3d has a single padding mode for all axes, so T is padded explicitly before the convolution.
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.ResidualBlock ¶
Bases: Module
Pre-activation block norm1 -> SiLU -> conv1 -> norm2 -> SiLU -> conv2 plus a (1x1x1 if narrowing) skip.
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.X2Branch ¶
X2Branch(config: Kandinsky6SRLatentUpscalerEntryConfig)
Bases: Module
Weights exclusive to the x2 path: a residual adapter at the input grid, a finisher, and private copies of the mid stage and second stage (so the x2 path shares no weights with the x4 path).
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
fastvideo.models.upsamplers.kandinsky6_sr.X2Finisher ¶
X2Finisher(channels: int)
Bases: Module
The PXSUpsample conv pair applied on the unchanged grid (no resize).
Source code in fastvideo/models/upsamplers/kandinsky6_sr.py
Functions:¶
fastvideo.models.upsamplers.kandinsky6_sr.nearest_2x ¶
Nearest-neighbour 2x resize of H and W of [B, C, T, H, W] (T folded into the batch).