Skip to content

vllm_omni.model_executor.models.ming_image.model

vLLM-native MLLM stage for inclusionAI Ming-Image.

MingImageDummyInputsBuilder

Bases: MingFlashOmniThinkerDummyInputsBuilder

get_dummy_mm_data

get_dummy_mm_data(
    seq_len: int,
    mm_counts: Mapping[str, int],
    mm_options=None,
) -> MultiModalDataDict

get_dummy_text

get_dummy_text(mm_counts: Mapping[str, int]) -> str

MingImageForConditionalGeneration

Bases: MingFlashOmniThinkerForConditionalGeneration

Ming-Image MLLM stage with Qwen2.5-VL and direct-state capture.

audio instance-attribute

audio = None

capture_layers instance-attribute

capture_layers = (5, 12, llm_config.num_hidden_layers)

config instance-attribute

config = llm_config

have_multimodal_outputs instance-attribute

have_multimodal_outputs = True

hf_to_vllm_mapper class-attribute instance-attribute

hf_to_vllm_mapper = WeightsMapper(
    orig_to_new_prefix={
        "model.": "language_model.",
        "linear_proj.": "linear_proj.proj.",
    }
)

language_model instance-attribute

language_model = BailingMoeV2ForCausalLM(
    vllm_config=vllm_config.with_hf_config(llm_config),
    prefix=maybe_prefix(prefix, "llm"),
)

linear_proj instance-attribute

linear_proj = VisionProjector(
    vision_dim=thinker_config.vision_config.out_hidden_size,
    llm_dim=llm_config.hidden_size,
    mlp_depth=getattr(thinker_config, "mlp_depth", 2),
)

linear_proj_audio instance-attribute

linear_proj_audio = None

make_empty_intermediate_tensors instance-attribute

make_empty_intermediate_tensors = (
    self.language_model.make_empty_intermediate_tensors
)

query_tokens_dict instance-attribute

query_tokens_dict = torch.nn.ParameterDict()

thinker_config instance-attribute

thinker_config = thinker_config

vision instance-attribute

vision = Qwen2_5_VisionTransformer(
    thinker_config.vision_config,
    norm_eps=llm_config.rms_norm_eps,
    quant_config=vllm_config.quant_config,
    prefix=maybe_prefix(prefix, "vision"),
)

compute_logits

compute_logits(
    hidden_states: Tensor, sampling_metadata=None
) -> Tensor | None

extract_image_feature

extract_image_feature(
    pixel_values: Tensor, grid_thw: Tensor
) -> Tensor

forward

forward(
    input_ids: Tensor,
    positions: Tensor,
    intermediate_tensors=None,
    inputs_embeds: Tensor | None = None,
    **kwargs,
) -> OmniOutput

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

MingImageMultiModalProcessor

MingImageProcessingInfo

Bases: MingFlashOmniThinkerProcessingInfo

get_data_parser

get_data_parser()

get_hf_config

get_hf_config() -> BailingMM2Config

get_hf_processor

get_hf_processor(**kwargs: object)

get_mm_max_tokens_per_item

get_mm_max_tokens_per_item(seq_len, mm_counts)

get_supported_mm_limits

get_supported_mm_limits() -> Mapping[str, int | None]