helios ¶
Classes¶
fastvideo.configs.models.dits.helios.HeliosArchConfig dataclass ¶
HeliosArchConfig(stacked_params_mapping: list[tuple[str, str, str]] = list(), _fsdp_shard_conditions: list = (lambda: [_is_transformer_block])(), _compile_conditions: list = list(), param_names_mapping: dict = dict(), reverse_param_names_mapping: dict = dict(), lora_param_names_mapping: dict = dict(), cast_prompt_embeds_to_dit_dtype: bool = False, _supported_attention_backends: tuple[AttentionBackendEnum, ...] = (FLASH_ATTN, TORCH_SDPA), hidden_size: int = 0, num_attention_heads: int = 40, num_channels_latents: int = 0, in_channels: int = 16, out_channels: int = 16, exclude_lora_layers: list[str] = list(), boundary_ratio: float | None = None, patch_size: tuple[int, int, int] = (1, 2, 2), attention_head_dim: int = 128, text_dim: int = 4096, freq_dim: int = 256, ffn_dim: int = 13824, num_layers: int = 40, cross_attn_norm: bool = True, qk_norm: str = 'rms_norm_across_heads', eps: float = 1e-06, added_kv_proj_dim: int | None = None, rope_dim: tuple[int, int, int] = (44, 42, 42), rope_theta: float = 10000.0, guidance_cross_attn: bool = True, zero_history_timestep: bool = True, has_multi_term_memory_patch: bool = True, is_amplify_history: bool = False, history_scale_mode: str = 'per_head')
Bases: DiTArchConfig
Architecture fields for Helios-Distilled's history-aware DiT.