Skip to content

vllm_omni.inputs.preprocess

Omni renderer: the upstream renderer class with Omni engine-input routing.

OmniRenderer is a mixin. omni_renderer_cls puts it in front of a concrete BaseRenderer subclass, so the Omni instance is an upstream renderer (tokenizer, executors, multimodal cache and every public entry point come from upstream unchanged) and only _process_singleton / _process_singleton_async are extended. build_omni_renderer resolves the concrete class the same way upstream renderer_from_config does.

OmniRenderer

Mixin that routes Omni prompts through the upstream renderer.

It has no base class of its own; omni_renderer_cls prepends it to a concrete BaseRenderer subclass, so self and super() are that renderer at runtime (hence the casts below).

Two extensions over upstream _process_singleton:

  • a token prompt that carries mm_processor_kwargs but no multi_modal_data still goes through the multimodal processor (AR image-generation models such as GLM-Image read their target size from those kwargs);
  • Omni-only prompt keys (prompt, cache_salt, additional_information, model_intermediate_buffer) are copied onto the engine input.

build_omni_renderer

build_omni_renderer(
    vllm_config: VllmConfig,
    *,
    renderer_cls: type[BaseRenderer] | None = None,
    tokenizer: TokenizerLike | None = None,
) -> BaseRenderer

Build the Omni renderer for vllm_config.

With renderer_cls unset this mirrors upstream renderer_from_config: the tokenizer and renderer mode come from the model config and the class from RENDERER_REGISTRY. Pass renderer_cls (and optionally tokenizer) to skip that resolution, e.g. for tokenizer-less stages.

omni_renderer_cls

omni_renderer_cls(
    base_cls: type[BaseRenderer],
) -> type[BaseRenderer]

Return the Omni subclass of base_cls (created once per base class).