vllm_omni.metrics ¶
Modules:
| Name | Description |
|---|---|
definitions | Single source of truth for vLLM-Omni Prometheus + bench CLI metric naming. |
modality | OmniModalityMetrics — per-modality Prometheus families (audio path only). |
prometheus | |
realtime | |
stat_logger | OmniPrometheusStatLogger — wrap upstream PrometheusStatLogger. |
stats | |
transfer | OmniTransferMetrics — cross-stage transfer Prometheus families. |
utils | |
OmniPrometheusMetrics ¶
Label-bound wrapper around the raw Prometheus metrics.
Metric collectors use the vllm_omni: prefix, distinct from the upstream vllm:* families.
observe_stage_gen_time ¶
OmniRequestCounter ¶
OrchestratorAggregator ¶
transfer_events instance-attribute ¶
accumulate_diffusion_metrics ¶
Accumulate diffusion metrics for a request.
Engine emits *_ms timings; the accumulator converts them to _s keys via _MS_TO_S so downstream observers read a uniform seconds-bearing dict. Per-chunk timing keys are summed; non-timing keys (e.g. image_num / resolution from format_diffusion_outputs) preserve the existing += semantics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
req_id | Any | Request ID | required |
engine_outputs | Any | Engine output object containing metrics | required |
on_finalize_request ¶
on_forward ¶
on_forward(
from_stage: int,
to_stage: int,
req_id: Any,
size_bytes: int,
tx_ms: float,
used_shm: bool,
) -> None
on_stage_metrics ¶
on_stage_metrics(
stage_id: int,
req_id: Any,
metrics: StageRequestStats,
final_output_type: str | None = None,
) -> None
process_stage_metrics ¶
process_stage_metrics(
*,
result: dict[str, Any],
stage_type: str,
stage_id: int,
req_id: str,
engine_outputs: Any,
finished: bool,
final_output_type: str | None,
output_to_yield: Any | None,
event_cursor: int = 0,
) -> None
Process and record stage metrics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
result | dict[str, Any] | Result dict containing metrics from stage | required |
stage_type | str | Type of the stage (e.g., 'llm', 'diffusion') | required |
stage_id | int | Stage identifier | required |
req_id | str | Request identifier | required |
engine_outputs | Any | Engine output object | required |
finished | bool | Whether stage processing is finished | required |
final_output_type | str | None | Type of final output (e.g., 'text', 'audio') | required |
output_to_yield | Any | None | Output object to attach metrics to | required |
record_audio_generated_frames ¶
record_stage_postprocess_time ¶
StageRequestStats dataclass ¶
diffusion_metrics class-attribute instance-attribute ¶
image_time_to_first_output_ms class-attribute instance-attribute ¶
image_time_to_first_output_ms: float = 0.0
inter_output_latencies_ms class-attribute instance-attribute ¶
pipeline_timings class-attribute instance-attribute ¶
StageStats dataclass ¶
count_audio_chunk_frames ¶
Count frames (samples) in one audio tensor / chunk.
Audio chunks are concatenated on dim=-1 in the output processor, so the frame/sample axis is the last dim (e.g. [channels, frames]). Keep this aligned with serving_chat.py: audio tensors are consumed as (T,), (C, T), or (B, C, T). Flattening would corrupt multi-channel audio.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
audio_chunk | object | A single audio tensor/array-like, or a scalar-like value. | required |
Returns:
| Type | Description |
|---|---|
int | Frame count for this chunk. Uses |
int |
|
count_audio_frames ¶
Sum frame counts over all audio chunks in mm_out["audio"] or with other related keys.
For multi-dim tensors (e.g. shape [channels, samples]) the last axis is the sample dim; for 1-D tensors the only axis is the sample dim; scalars count as 1. Missing or empty audio yields 0.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mm_out | Mapping[str, Any] | A multimodal_output Mapping (plain | required |
Returns:
| Type | Description |
|---|---|
int | Total audio frames (samples) across all chunks. |
count_image_pixels ¶
Count pixels in one image value, or sum over a nested list/tuple.
Accepts PIL-like objects (size=(W, H)), tensors / arrays with a shape attribute, and nested list / tuple containers.
Shape heuristics (aligned with StagePool image metrics):
ndim >= 4(e.g.BCHW):B * H * Wviadims[0] * dims[-2] * dims[-1]ndim == 3anddims[0] in (1, 3, 4): CHW →H * Wndim == 3anddims[-1] in (1, 3, 4): HWC →H * W- otherwise:
dims[-2] * dims[-1]
Returns 0 when value is missing or cannot be interpreted.