Skip to content

vllm_omni.utils.forced_aligner

Forced-aligner config + word-timestamp decoding for TTS.

The aligner runs as a pooling stage appended to the pipeline by --forced-aligner (see :func:inject_forced_aligner_stage). Qwen-specific word segmentation / prompt building / marker repair live in :mod:vllm_omni.utils.qwen3_force_align_processor.

logger module-attribute

logger = logging.getLogger(__name__)

ForcedAlignerConfig dataclass

Plain-data config for the forced-aligner pipeline stage; fields are consumed by :func:inject_forced_aligner_stage.

architecture class-attribute instance-attribute

architecture: str | None = None

dtype class-attribute instance-attribute

dtype: str | None = None

gpu_memory_utilization class-attribute instance-attribute

gpu_memory_utilization: float | None = None

max_model_len class-attribute instance-attribute

max_model_len: int | None = None

model instance-attribute

model: str

pooling_task class-attribute instance-attribute

pooling_task: str | None = None

runner class-attribute instance-attribute

runner: str | None = None

trust_remote_code class-attribute instance-attribute

trust_remote_code: bool | None = None

WordTimestamp dataclass

Internal alignment record. Serialized to a plain JSON object ({"word", "start_ms", "end_ms"}) at the WebSocket boundary.

end_ms instance-attribute

end_ms: int

start_ms instance-attribute

start_ms: int

word instance-attribute

word: str

build_forced_aligner_config

build_forced_aligner_config(
    model: str | None, config_path: str | None = None
) -> ForcedAlignerConfig | None

Build a config from the CLI values, or None when the feature is off.

Precedence (lowest to highest): the packaged default YAML (:data:_DEFAULT_CONFIG_PATH, Qwen deploy defaults) -> a user YAML passed via --forced-aligner-config -> the --forced-aligner model path. The feature is off (returns None) unless a model resolves from this chain. Per-field overrides such as gpu_memory_utilization live in the YAML.

extract_word_timestamps

extract_word_timestamps(
    res: Any, text: str, language: str | None = None
) -> list[dict] | None

Read a forced-aligner stage's engine output into per-word timestamps.

Returns [{word, start_ms, end_ms}, ...] from the stage's pooling output (res.outputs.data is an int32 [n_words, 2] tensor). Words come from re-segmenting text; any mismatch or non-aligner input returns None.

inject_forced_aligner_stage

inject_forced_aligner_stage(
    pipeline: PipelineConfig,
    deploy: DeployConfig,
    cli_overrides: dict[str, Any],
) -> tuple[PipelineConfig, DeployConfig]

Append a forced-aligner pooling stage to the pipeline tail when --forced-aligner or --forced-aligner-config resolves an aligner model; no-op otherwise. The stage runs the aligner model with runner="pooling", consuming the previous stage's audio and emitting a terminal word-timestamps side-output.