vllm_omni.diffusion.models.minimax_h3.encoder ¶
MiniMax H3 Qwen3-VL layer-50 text/vision encoder.
The encoder is reimplemented on top of vLLM-style tensor-parallel building blocks (column/row/vocab-parallel linears bound to the MiniMax H3 encoder process group) instead of the single-GPU transformers backbone. This makes the large Qwen3-VL text model TP-shardable across text_encoder_tp_size ranks, which removes the rank-0 memory hotspot that dominates no-offload peak memory (the retained 50-layer encoder is ~51.5 GB in BF16).
The computation contract is unchanged from the HF reference path:
- plain BF16 weights, only the first
MINIMAX_H3_QWEN3VL_SELECTED_LM_LAYERdecoder layers are kept (the checkpoint consumes the unnormalized hidden state after decoder layer 49); - cuDNN SDP is enabled during encode;
- the multimodal backbone runs with an all-ones attention mask and
mm_token_type_idsderived from the image/video token ids; - DeepStack visual features are injected at the first
len(deepstack_visual_indexes)decoder layers; - the output is
hidden_states[50]of shape[seq, 5120].
Tensor parallelism follows vLLM's TP semantics: vocab-parallel embeddings and row-parallel projections all-reduce so the hidden state stays fully replicated on every encoder rank, while column-parallel projections (QKV / MLP up) keep a local shard. After the final layer every encoder rank therefore holds the complete [seq, 5120] hidden state.
MiniMaxH3Qwen3VLEncoder ¶
Bases: Module
Frozen, TP-capable Qwen3-VL backbone returning layer-50 states.
encoder_group is a GroupCoordinator over the encoder tensor-parallel ranks (by default the first text_encoder_tp_size DiT ranks). Only ranks inside the group load weights; other ranks construct a parameter-free stub that never runs the forward.
text_model instance-attribute ¶
text_model = MiniMaxH3Qwen3VLTextModel(
encoder_group,
config.text_config,
MINIMAX_H3_QWEN3VL_SELECTED_LM_LAYER,
dtype,
quant_config=quant_config,
)
encode_ids ¶
encode_ids(
input_ids: Tensor,
*,
pixel_values: Tensor | None = None,
image_grid_thw: Tensor | None = None,
pixel_values_videos: Tensor | None = None,
video_grid_thw: Tensor | None = None,
) -> Tensor
forward ¶
forward(
input_ids: Tensor,
*,
pixel_values: Tensor | None = None,
image_grid_thw: Tensor | None = None,
pixel_values_videos: Tensor | None = None,
video_grid_thw: Tensor | None = None,
) -> Tensor
Hook-compatible entry point for model-level CPU offloading.
MiniMaxH3Qwen3VLMergedColumnParallelLinear ¶
Bases: LinearBase
Packed gate/up projection sharded along the output dimension.
intermediate_size_per_partition instance-attribute ¶
MiniMaxH3Qwen3VLQKVParallelLinear ¶
Bases: LinearBase
QKV projection with GQA head sharding along the output dimension.
MiniMaxH3Qwen3VLRMSNorm ¶
Bases: RMSNorm
Qwen3-VL RMSNorm using the common implementation.
The model-specific name keeps checkpoint weight keys stable, while the common RMSNorm dispatches to torch_npu.npu_rms_norm on Ascend. Gamma uses the checkpoint/model dtype, but native fallbacks keep the RMS variance reduction and scaling in float32.
MiniMaxH3Qwen3VLRowParallelLinear ¶
Bases: LinearBase
MiniMaxH3Qwen3VLTextAttention ¶
Bases: Module
head_dim instance-attribute ¶
head_dim = getattr(
config,
"head_dim",
config.hidden_size // config.num_attention_heads,
)
k_norm instance-attribute ¶
k_norm = MiniMaxH3Qwen3VLRMSNorm(
self.head_dim, eps=config.rms_norm_eps, dtype=dtype
)
o_proj instance-attribute ¶
o_proj = MiniMaxH3Qwen3VLRowParallelLinear(
group,
input_size=self.num_heads * self.head_dim,
output_size=self.hidden_size,
dtype=dtype,
input_is_parallel=True,
quant_config=quant_config,
prefix=f"{prefix}.o_proj",
)
q_norm instance-attribute ¶
q_norm = MiniMaxH3Qwen3VLRMSNorm(
self.head_dim, eps=config.rms_norm_eps, dtype=dtype
)
qkv_proj instance-attribute ¶
qkv_proj = MiniMaxH3Qwen3VLQKVParallelLinear(
group,
hidden_size=self.hidden_size,
num_heads=self.num_heads,
num_kv_heads=self.num_kv_heads,
head_dim=self.head_dim,
dtype=dtype,
quant_config=quant_config,
prefix=f"{prefix}.qkv_proj",
)
MiniMaxH3Qwen3VLTextDecoderLayer ¶
Bases: Module
input_layernorm instance-attribute ¶
input_layernorm = MiniMaxH3Qwen3VLRMSNorm(
config.hidden_size, eps=config.rms_norm_eps, dtype=dtype
)
mlp instance-attribute ¶
mlp = MiniMaxH3Qwen3VLTextMLP(
group,
config,
dtype,
quant_config=quant_config,
prefix=f"{prefix}.mlp",
)
post_attention_layernorm instance-attribute ¶
post_attention_layernorm = MiniMaxH3Qwen3VLRMSNorm(
config.hidden_size, eps=config.rms_norm_eps, dtype=dtype
)
self_attn instance-attribute ¶
self_attn = MiniMaxH3Qwen3VLTextAttention(
group,
config,
dtype,
quant_config=quant_config,
prefix=f"{prefix}.self_attn",
)
MiniMaxH3Qwen3VLTextMLP ¶
Bases: Module
down_proj instance-attribute ¶
down_proj = MiniMaxH3Qwen3VLRowParallelLinear(
group,
input_size=self.intermediate_size,
output_size=self.hidden_size,
dtype=dtype,
input_is_parallel=True,
quant_config=quant_config,
prefix=f"{prefix}.down_proj",
)
gate_up_proj instance-attribute ¶
gate_up_proj = MiniMaxH3Qwen3VLMergedColumnParallelLinear(
group,
input_size=self.hidden_size,
intermediate_size=self.intermediate_size,
dtype=dtype,
quant_config=quant_config,
prefix=f"{prefix}.gate_up_proj",
)
MiniMaxH3Qwen3VLTextModel ¶
Bases: Module
TP-sharded Qwen3-VL text model returning unnormalized layer-50 states.
embed_tokens instance-attribute ¶
embed_tokens = MiniMaxH3Qwen3VLVocabParallelEmbedding(
group,
num_embeddings=config.vocab_size,
embedding_dim=config.hidden_size,
dtype=dtype,
)
layers instance-attribute ¶
layers = nn.ModuleList(
[
MiniMaxH3Qwen3VLTextDecoderLayer(
group,
config,
dtype,
quant_config=quant_config,
prefix=f"{prefix}.layers.{layer_idx}",
)
for layer_idx in range(self.num_layers)
]
)
num_layers instance-attribute ¶
MiniMaxH3Qwen3VLTextRotaryEmbedding ¶
MiniMaxH3Qwen3VLVisionAttention ¶
Bases: Module
MiniMaxH3Qwen3VLVisionBlock ¶
Bases: Module
MiniMaxH3Qwen3VLVisionMLP ¶
MiniMaxH3Qwen3VLVisionModel ¶
Bases: Module
Qwen3-VL vision tower (patch embed + blocks + merger + DeepStack).
blocks instance-attribute ¶
blocks = nn.ModuleList(
[
MiniMaxH3Qwen3VLVisionBlock(config=config)
for _ in range(config.depth)
]
)
deepstack_merger_list instance-attribute ¶
deepstack_merger_list = nn.ModuleList(
[
MiniMaxH3Qwen3VLVisionPatchMerger(
config=config, use_postshuffle_norm=True
)
for _ in range(
len(config.deepstack_visual_indexes)
)
]
)
deepstack_visual_indexes instance-attribute ¶
merger instance-attribute ¶
merger = MiniMaxH3Qwen3VLVisionPatchMerger(
config=config, use_postshuffle_norm=False
)
pos_embed instance-attribute ¶
rotary_pos_emb instance-attribute ¶
rotary_pos_emb = MiniMaxH3Qwen3VLVisionRotaryEmbedding(
head_dim // 2
)
spatial_merge_unit instance-attribute ¶
MiniMaxH3Qwen3VLVisionPatchEmbed ¶
Bases: Module
proj instance-attribute ¶
proj = nn.Conv3d(
self.in_channels,
self.embed_dim,
kernel_size=kernel_size,
stride=kernel_size,
bias=True,
)