Skip to content

vllm_omni.utils.qwen3_force_align_processor

Qwen3 forced-aligner text/timestamp processor.

This is the model-specific half of upstream's Qwen3ForceAlignProcessor: it turns text into the aligner's word units and prompt, and repairs the predicted timestamp bins. It feeds the forced-aligner pipeline stage (input glue in :mod:vllm_omni.model_executor.stage_input_processors.forced_aligner, timestamp decoding in :mod:vllm_omni.utils.forced_aligner).

Keeping the Qwen-specific pieces here marks the seam for the model-agnostic aligner the issue asks for: a different aligner family would supply its own processor exposing the same small surface — :func:segment_words, :func:build_prompt, :func:fix_timestamp.

Word segmentation prefers qwen_asr's official Qwen3ForceAlignProcessor when installed (full multilingual fidelity, incl. Japanese/Korean) and otherwise uses the faithful port below: exact for whitespace-delimited and Chinese-mixed text; Japanese/Korean degrade to whitespace splitting.

AUDIO_PLACEHOLDER module-attribute

AUDIO_PLACEHOLDER = (
    "<|audio_start|><|audio_pad|><|audio_end|>"
)

TIMESTAMP_TOKEN module-attribute

TIMESTAMP_TOKEN = '<timestamp>'

logger module-attribute

logger = logging.getLogger(__name__)

build_prompt

build_prompt(words: list[str]) -> str

Build the Qwen3 aligner prompt exactly as qwen_asr's Qwen3ForceAlignProcessor.encode_timestamp does: the audio placeholder followed by the words, each with two trailing <timestamp> markers (start + end) that the model classifies into audio time bins.

No chat template. The official aligner feeds this string straight to the tokenizer; wrapping it in <|im_start|>user ... <|im_start|>assistant puts three tokens in front of the audio and shifts almost every predicted marker one 80 ms bin later than the official output.

fix_timestamp

fix_timestamp(values: list[int]) -> list[int]

Repair non-monotonic timestamp bins via LIS + interpolation.

The model occasionally emits out-of-order bins; this snaps anomalies back onto the longest non-decreasing subsequence (interpolating longer runs) so paired start/end markers stay ordered. Port of Qwen3ForceAlignProcessor.fix_timestamp (Apache-2.0).

segment_words

segment_words(
    text: str, language: str | None = None
) -> list[str]

Split text into the aligner's word units.

Prefers qwen_asr's official processor (full multilingual fidelity); falls back to the built-in port otherwise. language follows the official naming — "japanese" / "korean" (or codes like ja / ko) trigger language-specific tokenisers; anything else (incl. None / "auto") uses the whitespace + Chinese-mixed path.