Skip to content

lingbotworld_fast

Classes

fastvideo.configs.models.dits.lingbotworld_fast.LingBotWorldFastArchConfig dataclass

LingBotWorldFastArchConfig(stacked_params_mapping: list[tuple[str, str, str]] = list(), _fsdp_shard_conditions: list = (lambda: [is_blocks])(), _compile_conditions: list = list(), param_names_mapping: dict = dict(), reverse_param_names_mapping: dict = dict(), lora_param_names_mapping: dict = dict(), cast_prompt_embeds_to_dit_dtype: bool = False, _supported_attention_backends: tuple[AttentionBackendEnum, ...] = (SAGE_ATTN, FLASH_ATTN, TORCH_SDPA, VIDEO_SPARSE_ATTN, VMOBA_ATTN, SAGE_ATTN_THREE, ATTN_QAT_INFER, ATTN_QAT_TRAIN, SLA_ATTN, SAGE_SLA_ATTN), hidden_size: int = 0, num_attention_heads: int = 0, num_channels_latents: int = 0, in_channels: int = 0, out_channels: int = 0, exclude_lora_layers: list[str] = list(), boundary_ratio: float | None = None, model_type: str = 'i2v', patch_size: tuple[int, int, int] = (1, 2, 2), text_len: int = 512, in_dim: int = 36, dim: int = 5120, ffn_dim: int = 13824, freq_dim: int = 256, text_dim: int = 4096, out_dim: int = 16, num_heads: int = 40, num_layers: int = 40, qk_norm: bool = True, cross_attn_norm: bool = True, eps: float = 1e-06, local_attn_size: int = -1, sink_size: int = 9, chunk_size: int = 3, sample_shift: float = 10.0, num_train_timesteps: int = 1000, timesteps_index: tuple[int, int, int, int] = (0, 179, 358, 679), max_area: int = 480 * 832)

Bases: LingBotWorld2CausalFastArchConfig

Arch config for the released LingBot-World-Fast checkpoint.

The tensor layout is identical to the LingBot World 2 causal-fast model, so the parent's shapes and param_names_mapping are reused verbatim. Only the sampling/attention-window values released with this checkpoint differ.