vllm_omni.inputs.preprocess ¶
Omni renderer: the upstream renderer class with Omni engine-input routing.
OmniRenderer is a mixin. omni_renderer_cls puts it in front of a concrete BaseRenderer subclass, so the Omni instance is an upstream renderer (tokenizer, executors, multimodal cache and every public entry point come from upstream unchanged) and only _process_singleton / _process_singleton_async are extended. build_omni_renderer resolves the concrete class the same way upstream renderer_from_config does.
OmniRenderer ¶
Mixin that routes Omni prompts through the upstream renderer.
It has no base class of its own; omni_renderer_cls prepends it to a concrete BaseRenderer subclass, so self and super() are that renderer at runtime (hence the casts below).
Two extensions over upstream _process_singleton:
- a token prompt that carries
mm_processor_kwargsbut nomulti_modal_datastill goes through the multimodal processor (AR image-generation models such as GLM-Image read their target size from those kwargs); - Omni-only prompt keys (
prompt,cache_salt,additional_information,model_intermediate_buffer) are copied onto the engine input.
build_omni_renderer ¶
build_omni_renderer(
vllm_config: VllmConfig,
*,
renderer_cls: type[BaseRenderer] | None = None,
tokenizer: TokenizerLike | None = None,
) -> BaseRenderer
Build the Omni renderer for vllm_config.
With renderer_cls unset this mirrors upstream renderer_from_config: the tokenizer and renderer mode come from the model config and the class from RENDERER_REGISTRY. Pass renderer_cls (and optionally tokenizer) to skip that resolution, e.g. for tokenizer-less stages.