vllm_omni.utils.forced_aligner ¶
Forced-aligner config + word-timestamp decoding for TTS.
The aligner runs as a pooling stage appended to the pipeline by --forced-aligner (see :func:inject_forced_aligner_stage). Qwen-specific word segmentation / prompt building / marker repair live in :mod:vllm_omni.utils.qwen3_force_align_processor.
ForcedAlignerConfig dataclass ¶
WordTimestamp dataclass ¶
build_forced_aligner_config ¶
build_forced_aligner_config(
model: str | None, config_path: str | None = None
) -> ForcedAlignerConfig | None
Build a config from the CLI values, or None when the feature is off.
Precedence (lowest to highest): the packaged default YAML (:data:_DEFAULT_CONFIG_PATH, Qwen deploy defaults) -> a user YAML passed via --forced-aligner-config -> the --forced-aligner model path. The feature is off (returns None) unless a model resolves from this chain. Per-field overrides such as gpu_memory_utilization live in the YAML.
extract_word_timestamps ¶
Read a forced-aligner stage's engine output into per-word timestamps.
Returns [{word, start_ms, end_ms}, ...] from the stage's pooling output (res.outputs.data is an int32 [n_words, 2] tensor). Words come from re-segmenting text; any mismatch or non-aligner input returns None.
inject_forced_aligner_stage ¶
inject_forced_aligner_stage(
pipeline: PipelineConfig,
deploy: DeployConfig,
cli_overrides: dict[str, Any],
) -> tuple[PipelineConfig, DeployConfig]
Append a forced-aligner pooling stage to the pipeline tail when --forced-aligner or --forced-aligner-config resolves an aligner model; no-op otherwise. The stage runs the aligner model with runner="pooling", consuming the previous stage's audio and emitting a terminal word-timestamps side-output.