minimax_h3_int8_convrot ¶
Comfy int8_tensorwise + ConvRot overlay for the MiniMax-H3 video VAE decoder.
The export stores decoder transformer linears as signed int8 with per-output channel scales and a JSON comfy_quant marker:
{"format": "int8_tensorwise", "convrot": true, "convrot_groupsize": 256}
Weights were rotated offline by a normalized regular Hadamard (group 256). Inference rotates activations with the same matrix, row-quantizes them, then runs int8 GEMM. Encoder convolutions stay dense; only the ViT decoder blocks are quantized.
Comfy names (to_qkv, ff.w1 / ff.w2, x_embedder) are remapped onto FastVideo's split Q/K/V and ff.net surface.
Classes¶
fastvideo.models.vaes.minimax_h3_int8_convrot.Int8ConvRotLinear ¶
Int8ConvRotLinear(in_features: int, out_features: int, *, bias: bool, convrot: bool, group_size: int)
Bases: Module
W8A8 linear matching Comfy int8_tensorwise (+ optional ConvRot).
Source code in fastvideo/models/vaes/minimax_h3_int8_convrot.py
Functions:¶
fastvideo.models.vaes.minimax_h3_int8_convrot.dense_vae_safetensors ¶
Drop the ConvRot overlay so it is not loaded as a dense VAE shard.
fastvideo.models.vaes.minimax_h3_int8_convrot.overlay_minimax_h3_int8_convrot_decoder ¶
Swap H3 VAE decoder transformer linears for Comfy int8-convrot weights.
Dense decoder tensors that the export still stores in float (embed, norms, scales, proj_out) are copied onto the matching FastVideo modules. Returns the number of quantized linears installed.
Source code in fastvideo/models/vaes/minimax_h3_int8_convrot.py
252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 | |
fastvideo.models.vaes.minimax_h3_int8_convrot.parse_comfy_quant_marker ¶
Decode the uint8 JSON marker Comfy stores next to each quantized linear.
Source code in fastvideo/models/vaes/minimax_h3_int8_convrot.py
fastvideo.models.vaes.minimax_h3_int8_convrot.regular_hadamard ¶
regular_hadamard(size: int, *, device: device, dtype: dtype) -> Tensor
Normalized regular Hadamard of order 4**k (ConvRot Theorem 3.3).
Source code in fastvideo/models/vaes/minimax_h3_int8_convrot.py
fastvideo.models.vaes.minimax_h3_int8_convrot.shared_int8_projections ¶
Reuse identical ConvRot/row quantization while retaining each projection's INT8 GEMM.