Skip to content

vllm_omni.metrics.utils

DIFFUSION_METRICS_ONLY_REQUEST_ID module-attribute

DIFFUSION_METRICS_ONLY_REQUEST_ID = (
    "__vllm_omni_diffusion_metrics__"
)

FAILURE_REASONS module-attribute

FAILURE_REASONS = frozenset(
    {
        "client_abort",
        "client_disconnect",
        "stage_error",
        "unknown",
    }
)

coerce_bool

coerce_bool(
    value: object, *, default: bool = False
) -> bool

Coerce JSON / form / env-like values to bool.

Accepts bool, numeric 0/1, and common truthy strings ("1" / "true" / "yes" / "on", case-insensitive). None and unrecognized types return default.

coerce_positive_float_scalar

coerce_positive_float_scalar(value: object) -> float | None

Coerce a scalar-like value to a positive float.

Handles plain numeric values, one-element tensor/numpy-like values via .item(), and list/tuple wrappers. Returns None when no positive float can be extracted.

coerce_positive_int_scalar

coerce_positive_int_scalar(value: object) -> int | None

Coerce a value to a positive int without importing tensor libs.

Meant to pull positive integers such as sample rate out of the shapes that show up across omni stage outputs / configs:

  • plain int: 44100
  • torch.Tensor / numpy scalar: tensor(44100) (via .item())
  • list / tuple wrap: [tensor(44100)] or [44100]
  • Mapping field values: {"sr": tensor(44100)} → coerce the field value (key lookup is done by the caller, e.g. :func:resolve_int_by_sequential_keys)

Returns None when the value is missing, unparsable, or not positive.

count_audio_chunk_frames

count_audio_chunk_frames(audio_chunk: object) -> int

Count frames (samples) in one audio tensor / chunk.

Audio chunks are concatenated on dim=-1 in the output processor, so the frame/sample axis is the last dim (e.g. [channels, frames]). Keep this aligned with serving_chat.py: audio tensors are consumed as (T,), (C, T), or (B, C, T). Flattening would corrupt multi-channel audio.

Parameters:

Name Type Description Default
audio_chunk object

A single audio tensor/array-like, or a scalar-like value.

required

Returns:

Type Description
int

Frame count for this chunk. Uses shape[-1] when shaped; otherwise

int

len(...); scalars / unlenable values count as 1.

count_audio_frames

count_audio_frames(mm_out: Mapping[str, Any]) -> int

Sum frame counts over all audio chunks in mm_out["audio"] or with other related keys.

For multi-dim tensors (e.g. shape [channels, samples]) the last axis is the sample dim; for 1-D tensors the only axis is the sample dim; scalars count as 1. Missing or empty audio yields 0.

Parameters:

Name Type Description Default
mm_out Mapping[str, Any]

A multimodal_output Mapping (plain dict or MultimodalPayload) that may contain an audio or related key whose value is one chunk or a list of chunks.

required

Returns:

Type Description
int

Total audio frames (samples) across all chunks.

count_image_pixels

count_image_pixels(value: object) -> int

Count pixels in one image value, or sum over a nested list/tuple.

Accepts PIL-like objects (size=(W, H)), tensors / arrays with a shape attribute, and nested list / tuple containers.

Shape heuristics (aligned with StagePool image metrics):

  • ndim >= 4 (e.g. BCHW): B * H * W via dims[0] * dims[-2] * dims[-1]
  • ndim == 3 and dims[0] in (1, 3, 4): CHW → H * W
  • ndim == 3 and dims[-1] in (1, 3, 4): HWC → H * W
  • otherwise: dims[-2] * dims[-1]

Returns 0 when value is missing or cannot be interpreted.

count_tokens_from_outputs

count_tokens_from_outputs(engine_outputs: list[Any]) -> int

count_video_frames

count_video_frames(video: object) -> int | None

Return the frame count for common nested and tensor video layouts.

diffusion_exception_metrics

diffusion_exception_metrics(
    exc: BaseException,
) -> dict[str, Any]

Return metrics attached to a terminal diffusion error.

diffusion_scheduler_waiting_metrics

diffusion_scheduler_waiting_metrics(
    n_waiting: int,
) -> dict[str, int]

Build the diffusion scheduler snapshot consumed by the orchestrator.

extract_diffusion_denoise_ms

extract_diffusion_denoise_ms(output: Any) -> float | None

Extract denoise-loop time (transformer and per-step scheduler).

extract_diffusion_vae_decode_ms

extract_diffusion_vae_decode_ms(
    output: Any,
) -> float | None

extract_mm_output

extract_mm_output(mm_source: object) -> Mapping[str, Any]

Return the first non-empty multimodal_output Mapping on mm_source.

Lookup order
  1. mm_source.multimodal_output — top-level attribute / property
  2. mm_source.outputs[0].multimodal_output — CompletionOutput nesting (typical AR audio path)

Accepts both plain dict and MultimodalPayload (a Mapping). Returns {} when neither location yields a non-empty Mapping.

Parameters:

Name Type Description Default
mm_source object

Duck-typed container with optional multimodal_output and/or outputs (e.g. vLLM RequestOutput or OmniRequestOutput).

required

Returns:

Type Description
Mapping[str, Any]

The first non-empty multimodal Mapping found, or an empty dict.

extract_queue_wait_s

extract_queue_wait_s(
    pipeline_timings: Mapping[str, float] | None,
) -> float | None

iter_mm_outputs

iter_mm_outputs(
    mm_source: object,
) -> list[Mapping[str, Any]]

Collect all non-empty multimodal_output Mappings from mm_source.

Lookup order
  1. mm_source.multimodal_output — top-level attribute / property
  2. each mm_source.outputs[i].multimodal_output — every CompletionOutput entry that carries a non-empty multimodal Mapping

Accepts both plain dict and MultimodalPayload (a Mapping). Used by stage-pool metrics aggregation that needs to visit every mm payload.

Parameters:

Name Type Description Default
mm_source object

Duck-typed container with optional multimodal_output and/or outputs (e.g. vLLM RequestOutput or OmniRequestOutput).

required

Returns:

Type Description
list[Mapping[str, Any]]

A list of non-empty multimodal Mappings in discovery order. Empty when

list[Mapping[str, Any]]

none are present.

normalize_failure_reason

normalize_failure_reason(reason: str | None) -> str

Map request failure details to the bounded metrics taxonomy.

observe_audio_finalize

observe_audio_finalize(
    mod_metrics: OmniModalityMetrics,
    *,
    stage_id: int,
    replica_id: int,
    stage_metrics: Any,
    engine_outputs: Any,
) -> None

observe_diffusion_finalize

observe_diffusion_finalize(
    mod_metrics: OmniModalityMetrics,
    *,
    stage_id: int,
    replica_id: int,
    stage_metrics: Any,
) -> None

observe_stage_workload_metrics

observe_stage_workload_metrics(
    prom_metrics: OmniPrometheusMetrics,
    *,
    stage_type: str,
    stage_metrics: StageRequestStats,
) -> None

resolve_int_by_sequential_keys

resolve_int_by_sequential_keys(
    source: Mapping[str, object] | object | None,
    keys: Sequence[str],
) -> int | None

Return the first positive int found by trying keys in order.

For each key, looks up source[key] when source is a Mapping, otherwise getattr(source, key, None). Values are coerced via :func:coerce_positive_int_scalar. Returns None when source is empty/missing or no key yields a usable int.

sum_diffusion_stage_durations_ms

sum_diffusion_stage_durations_ms(
    output: Any, suffix: str
) -> float | None

Sum matching diffusion stage durations in milliseconds.