Skip to content

memo

Content-keyed memo for MiniMax-H3 reference encodes.

Serving the same reference images again (a re-roll, a new seed, the next turn of a session) re-runs the Qwen3-VL presentation and the VAE keyframe encode on identical inputs. Both are pure functions of their inputs — the keyframe posterior sample uses a fixed-seed generator — so their outputs can be reused exactly. Keys hash the pixel bytes, never object identity, so a mutated or re-decoded image misses. Entries are cloned on the way in and out, so callers may mutate what they get back.

Every sequence-parallel rank prepares the same references in the same order, so hits and misses are uniform across ranks and an encode that enters a collective is either skipped or run on all of them.

Classes

fastvideo.pipelines.basic.minimax_h3.memo.ContentMemo

ContentMemo(capacity: int | None = None)

A small LRU of cloned tensor (or tuple-of-tensor) results; capacity 0 disables it.

Source code in fastvideo/pipelines/basic/minimax_h3/memo.py
def __init__(self, capacity: int | None = None) -> None:
    self.capacity = max(0, envs.FASTVIDEO_H3_REF2VA_MEMO_ENTRIES.get() if capacity is None else capacity)
    self._items: OrderedDict[Any, Any] = OrderedDict()

Functions:

fastvideo.pipelines.basic.minimax_h3.memo.image_key

image_key(image: Any) -> tuple

Shape, dtype and a 128-bit digest of an image's pixel bytes.

Source code in fastvideo/pipelines/basic/minimax_h3/memo.py
def image_key(image: Any) -> tuple:
    """Shape, dtype and a 128-bit digest of an image's pixel bytes."""
    pixels = np.ascontiguousarray(np.asarray(image))
    return pixels.shape, str(pixels.dtype), hashlib.blake2b(pixels.tobytes(), digest_size=16).hexdigest()