vllm_omni.metrics.utils ¶
DIFFUSION_METRICS_ONLY_REQUEST_ID module-attribute ¶
FAILURE_REASONS module-attribute ¶
FAILURE_REASONS = frozenset(
{
"client_abort",
"client_disconnect",
"stage_error",
"unknown",
}
)
coerce_bool ¶
Coerce JSON / form / env-like values to bool.
Accepts bool, numeric 0/1, and common truthy strings ("1" / "true" / "yes" / "on", case-insensitive). None and unrecognized types return default.
coerce_positive_float_scalar ¶
Coerce a scalar-like value to a positive float.
Handles plain numeric values, one-element tensor/numpy-like values via .item(), and list/tuple wrappers. Returns None when no positive float can be extracted.
coerce_positive_int_scalar ¶
Coerce a value to a positive int without importing tensor libs.
Meant to pull positive integers such as sample rate out of the shapes that show up across omni stage outputs / configs:
- plain
int:44100 torch.Tensor/ numpy scalar:tensor(44100)(via.item())list/tuplewrap:[tensor(44100)]or[44100]- Mapping field values:
{"sr": tensor(44100)}→ coerce the field value (key lookup is done by the caller, e.g. :func:resolve_int_by_sequential_keys)
Returns None when the value is missing, unparsable, or not positive.
count_audio_chunk_frames ¶
Count frames (samples) in one audio tensor / chunk.
Audio chunks are concatenated on dim=-1 in the output processor, so the frame/sample axis is the last dim (e.g. [channels, frames]). Keep this aligned with serving_chat.py: audio tensors are consumed as (T,), (C, T), or (B, C, T). Flattening would corrupt multi-channel audio.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
audio_chunk | object | A single audio tensor/array-like, or a scalar-like value. | required |
Returns:
| Type | Description |
|---|---|
int | Frame count for this chunk. Uses |
int |
|
count_audio_frames ¶
Sum frame counts over all audio chunks in mm_out["audio"] or with other related keys.
For multi-dim tensors (e.g. shape [channels, samples]) the last axis is the sample dim; for 1-D tensors the only axis is the sample dim; scalars count as 1. Missing or empty audio yields 0.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mm_out | Mapping[str, Any] | A multimodal_output Mapping (plain | required |
Returns:
| Type | Description |
|---|---|
int | Total audio frames (samples) across all chunks. |
count_image_pixels ¶
Count pixels in one image value, or sum over a nested list/tuple.
Accepts PIL-like objects (size=(W, H)), tensors / arrays with a shape attribute, and nested list / tuple containers.
Shape heuristics (aligned with StagePool image metrics):
ndim >= 4(e.g.BCHW):B * H * Wviadims[0] * dims[-2] * dims[-1]ndim == 3anddims[0] in (1, 3, 4): CHW →H * Wndim == 3anddims[-1] in (1, 3, 4): HWC →H * W- otherwise:
dims[-2] * dims[-1]
Returns 0 when value is missing or cannot be interpreted.
count_video_frames ¶
Return the frame count for common nested and tensor video layouts.
diffusion_exception_metrics ¶
diffusion_exception_metrics(
exc: BaseException,
) -> dict[str, Any]
Return metrics attached to a terminal diffusion error.
diffusion_scheduler_waiting_metrics ¶
Build the diffusion scheduler snapshot consumed by the orchestrator.
extract_diffusion_denoise_ms ¶
Extract denoise-loop time (transformer and per-step scheduler).
extract_mm_output ¶
Return the first non-empty multimodal_output Mapping on mm_source.
Lookup order
mm_source.multimodal_output— top-level attribute / propertymm_source.outputs[0].multimodal_output— CompletionOutput nesting (typical AR audio path)
Accepts both plain dict and MultimodalPayload (a Mapping). Returns {} when neither location yields a non-empty Mapping.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mm_source | object | Duck-typed container with optional | required |
Returns:
| Type | Description |
|---|---|
Mapping[str, Any] | The first non-empty multimodal Mapping found, or an empty |
extract_queue_wait_s ¶
iter_mm_outputs ¶
Collect all non-empty multimodal_output Mappings from mm_source.
Lookup order
mm_source.multimodal_output— top-level attribute / property- each
mm_source.outputs[i].multimodal_output— every CompletionOutput entry that carries a non-empty multimodal Mapping
Accepts both plain dict and MultimodalPayload (a Mapping). Used by stage-pool metrics aggregation that needs to visit every mm payload.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mm_source | object | Duck-typed container with optional | required |
Returns:
| Type | Description |
|---|---|
list[Mapping[str, Any]] | A list of non-empty multimodal Mappings in discovery order. Empty when |
list[Mapping[str, Any]] | none are present. |
normalize_failure_reason ¶
Map request failure details to the bounded metrics taxonomy.
observe_audio_finalize ¶
observe_audio_finalize(
mod_metrics: OmniModalityMetrics,
*,
stage_id: int,
replica_id: int,
stage_metrics: Any,
engine_outputs: Any,
) -> None
observe_diffusion_finalize ¶
observe_diffusion_finalize(
mod_metrics: OmniModalityMetrics,
*,
stage_id: int,
replica_id: int,
stage_metrics: Any,
) -> None
observe_stage_workload_metrics ¶
observe_stage_workload_metrics(
prom_metrics: OmniPrometheusMetrics,
*,
stage_type: str,
stage_metrics: StageRequestStats,
) -> None
resolve_int_by_sequential_keys ¶
resolve_int_by_sequential_keys(
source: Mapping[str, object] | object | None,
keys: Sequence[str],
) -> int | None
Return the first positive int found by trying keys in order.
For each key, looks up source[key] when source is a Mapping, otherwise getattr(source, key, None). Values are coerced via :func:coerce_positive_int_scalar. Returns None when source is empty/missing or no key yields a usable int.