Text-To-Speech (Online Serving)¶
Source https://github.com/vllm-project/vllm-omni/tree/main/examples/online_serving/text_to_speech.
vLLM-Omni exposes TTS models through the OpenAI-compatible POST /v1/audio/speech endpoint, launched with vllm serve <model> --omni. Each TTS model has its own subdirectory containing client snippets, gradio demos, and helper scripts; this README is the single doc entry point for all of them.
For offline inference, see examples/offline_inference/text_to_speech. For the full list of supported architectures across all modalities, see Supported Models.
Supported Models¶
| Model | HuggingFace repo | Voice cloning | Streaming | Voice presets / upload | Gradio demo |
|---|---|---|---|---|---|
| Breeze-TTS-2 | BreezeBlue/Breeze-TTS-2 | ✓ (ref_audio+ref_text) | ✓ (PCM stream) | speaker tags (S0..S9, default S0) | — |
| Audio8 TTS Preview | Audio8/Audio8-TTS-Preview-0.6b | ✓ (ref_audio+ref_text) | ✓ (PCM stream) | uploaded audio voice only; no presets | ✓ |
| Fish Speech S2 Pro | fishaudio/s2-pro | ✓ (ref_audio+ref_text) | ✓ (PCM stream) | — | ✓ |
| Gepard-1.0 | nineninesix/gepard-1.0 | — (zero-shot default voice) | ✓ (PCM / WAV stream) | default only | — |
| GLM-TTS | zai-org/GLM-TTS | ✓ (ref_audio+ref_text, required) | ✓ (PCM stream) | — | ✓ |
| IndexTTS-2 | IndexTeam/IndexTTS-2 | ✓ (ref_audio or uploaded voice) | stream=true response, non-chunk | uploaded audio voice only; no presets | — |
| IndexTTS-2.5 | native checkpoints/ bundle | ✓ (ref_audio or uploaded voice) | stream=true response, non-chunk | uploaded audio voice only; no presets | — |
| Ming-omni-tts | inclusionAI/Ming-omni-tts-0.5B | ✓ (ref_audio / speaker_embedding) | ✓ (PCM stream) | IP labels + structured instructions | — |
| Ming-flash-omni-TTS | Jonathan1909/Ming-flash-omni-2.0 | — (caption-controlled) | — | caption fields (instructions) | — |
| MOSS-TTS-Nano | OpenMOSS-Team/MOSS-TTS-Nano | ✓ (ref_audio required) | ✓ (PCM stream) | — | ✓ |
| OmniVoice | k2-fsa/OmniVoice | ✓ | — | — | — |
| Qwen3-TTS | Qwen/Qwen3-TTS-12Hz-1.7B-{CustomVoice,VoiceDesign,Base} | ✓ (Base) | ✓ (PCM + WebSocket) | ✓ (presets + /v1/audio/voices upload) | ✓ (standard + FastRTC) |
| VoxCPM2 | openbmb/VoxCPM2 | ✓ | ✓ (AudioWorklet via gradio) | — | ✓ |
| Voxtral TTS | mistralai/Voxtral-4B-TTS-2603 | ✓ (gated upstream) | ✓ | ✓ (presets) | ✓ |
CosyVoice3 is intentionally absent: no online example exists for it yet. See its offline section instead.
Common Quick Start¶
Launch the server (defaults shown — adjust --port, --gpu-memory-utilization, etc. as needed):
Send a TTS request via curl. These generic snippets assume a model with a preset/default voice; voice-cloning-only models such as IndexTTS-2 require ref_audio or an uploaded audio voice (see model-specific sections below).
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Hello, how are you?",
"voice": "default",
"response_format": "wav"
}' --output output.wav
Or via Python httpx:
import httpx
response = httpx.post(
"http://localhost:8091/v1/audio/speech",
json={
"input": "Hello, how are you?",
"voice": "default",
"response_format": "wav",
},
timeout=300.0,
)
open("output.wav", "wb").write(response.content)
Or via the OpenAI SDK:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8091/v1", api_key="none")
response = client.audio.speech.create(
model="<hf-repo>",
voice="default",
input="Hello, how are you?",
)
response.stream_to_file("output.wav")
Streaming PCM output (where supported) — set stream=true, stream_format="audio", and response_format="pcm":
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Hello, how are you?",
"voice": "default",
"stream": true,
"stream_format": "audio",
"response_format": "pcm"
}' --no-buffer | play -t raw -r 24000 -e signed -b 16 -c 1 -
Adjust the player's sample rate to match the model (44.1 kHz for Fish Speech, 48 kHz for VoxCPM2, 22.05 kHz for Gepard and IndexTTS-2, and 24 kHz for many others).
For full request-shape documentation (all parameters, response formats, error codes), see the Speech API reference.
Gepard-1.0¶
Single-stage native AR TTS at 22.05 kHz mono. Zero-shot only: omit voice or pass "default". Voice cloning from reference audio is not available yet.
Prerequisites¶
Same NeMo NanoCodec install as the offline Gepard section. On a host whose CUDA toolkit cannot build kernels — no nvcc/ninja, or a consumer Blackwell (sm_120) card — also export VLLM_USE_FLASHINFER_SAMPLER=0 before launch (same note as offline). That env is not a deploy-YAML field.
Launch¶
vllm-omni serve nineninesix/gepard-1.0 --omni --port 8091 --trust-remote-code \
--stage-init-timeout 900 \
--deploy-config vllm_omni/deploy/gepard.yaml
# or:
./gepard/run_server.sh
--stage-init-timeout 900 matches the online e2e fixture and the offline example. Serve defaults to 300s, which is often too short for a cold download of the talker plus NanoCodec.
The packaged vllm_omni/deploy/gepard.yaml must be passed with --deploy-config. The checkpoint self-identifies as qwen3_5_text, so omitting the YAML launches a diffusion fallback instead of the Gepard pipeline. The YAML sets async_chunk: false, max_num_seqs: 4, and currently pins seed: 42, so serving is deterministic by default until that YAML seed is removed. Pass an explicit per-request seed in tests and clients rather than depending on either default.
Sending requests¶
python examples/online_serving/text_to_speech/gepard/speech_client.py \
--text "Hello, this is Gepard speaking."
python examples/online_serving/text_to_speech/gepard/speech_client.py \
--text "Hello, this is Gepard speaking." --seed 7 --stream --output output.pcm
Notes¶
- Output: 22.05 kHz mono.
max_new_tokensis a frame budget (1 token = 1 frame = 1024 samples ≈ 46.4 ms at 21.5 fps; adapter bounds 1..4096). - Supported request fields:
input(required),voice(defaultonly),response_format(wavdefault;wav/pcm/flac/mp3non-streaming;opus400 because 22.05 kHz is not an Opus sample rate; streamingpcm/wavonly),stream/stream_format,max_new_tokens,seed. - Unsupported:
speed,extra_params(includingtemperature/top_p/top_k),ref_audio,ref_text,speaker_embedding,task_type,instructions,language, andword_timestamps. - Concurrent requests at
max_num_seqs: 4are supported. Native-AR recompute preemption is a known limitation of this architecture (a request that is preempted mid-generation can resume incorrectly); keep concurrency at or belowmax_num_seqsand treat preemption as out of scope until the platform fix lands. - Optional comparison against the upstream Gepard reference server needs Blackwell/Hopper + CUDA 13 + Postgres and is not part of CI.
GLM-TTS¶
2-stage TTS (AR + DiT flow-matching) at 24 kHz. Every request requires ref_audio + ref_text.
Launch¶
vllm serve zai-org/GLM-TTS --omni --trust-remote-code --port 8091
# or:
bash examples/online_serving/text_to_speech/glm_tts/run_server.sh /path/to/GLM-TTS
Sending requests¶
# Voice cloning (required)
python examples/online_serving/text_to_speech/glm_tts/openai_speech_client.py \
--text "你好,这是语音克隆测试。" \
--ref-audio file:///path/to/ref.wav \
--ref-text "这是参考音频的文本内容。"
# Custom format
python examples/online_serving/text_to_speech/glm_tts/openai_speech_client.py \
--text "Hello, this is a voice cloning test." \
--ref-audio file:///path/to/ref.wav \
--ref-text "Transcript of the reference audio." \
--response-format mp3 -o output.mp3
Gradio demo¶
Notes¶
- Output: 24 kHz mono WAV via HiFT vocoder.
ref_audio+ref_textare required together on every request. Reference audio should be 3-10 seconds.- Voice cloning feature extraction (WhisperVQ, CampPlus, mel) runs on the model side — no external dependency on the serving layer.
IndexTTS-2 and IndexTTS-2.5¶
2-stage TTS at 22.05 kHz. Requests use ref_audio for voice cloning, or an uploaded audio voice from /v1/audio/voices. IndexTTS-2.5 uses a multilingual tokenizer, CAMPPlus speaker projection, EnhancedCodec, and code-only Stage 0→1 transfer by default. Both versions support emotion conditioning via emo_audio, emo_text, or emo_vector passed in extra_params.
Prerequisites¶
IndexTTS-2 uses its existing frontend and does not require the IndexTTS-2.5 text dependencies. Before launching IndexTTS-2.5, install its optional frontend dependencies:
Launch¶
vllm serve IndexTeam/IndexTTS-2 --omni --trust-remote-code --port 8092
# or, to pass the bundled deploy config explicitly:
bash examples/online_serving/text_to_speech/indextts2/run_server.sh
# IndexTTS-2.5 default: use_gpt_latent=false
MODEL_VERSION=2.5 \
MODEL=/path/to/indextts-2.5 \
bash examples/online_serving/text_to_speech/indextts2/run_server.sh
Sending requests¶
# Voice cloning (ref_audio required)
python examples/online_serving/text_to_speech/indextts2/speech_client.py \
--text "你好,世界!" \
--ref-audio /path/to/reference.wav
# With emotion audio
python examples/online_serving/text_to_speech/indextts2/speech_client.py \
--text "今天心情很好!" \
--ref-audio /path/to/ref.wav \
--emo-audio /path/to/happy.wav
# IndexTTS-2.5 Japanese request
python examples/online_serving/text_to_speech/indextts2/speech_client.py \
--model-version 2.5 \
--model /path/to/indextts-2.5 \
--lang ja \
--text "こんにちは、音声合成のテストです。" \
--ref-audio /path/to/ref.wav
# IndexTTS-2.5 native speed control
python examples/online_serving/text_to_speech/indextts2/speech_client.py \
--model-version 2.5 \
--model /path/to/indextts-2.5 \
--speed 2.0 \
--text "这是两倍速度的语音合成测试。" \
--ref-audio /path/to/ref.wav
Notes¶
- Output: 22.05 kHz mono WAV.
- Provide
ref_audioon the documented raw request path, or passvoiceonly when it names an uploaded audio voice; neither version provides a built-in text-only preset voice. - IndexTTS-2.5 request controls
langandtext_normalizationare carried inextra_params; the client does this for--model-version 2.5. - IndexTTS-2.5 accepts the public
speedrequest field in[0.5, 2.0]:2.0generates shorter, faster speech and0.5generates longer, slower speech. This is native Stage 1 duration control, so serving does not apply the generic playback-speed adjustment a second time. IndexTTS-2 does not use this control. - IndexTTS-2.5 accepts language codes such as
zh,en,zhen(mixed Chinese/English),ja, andyue.Mandarinis a vLLM-Omni convenience alias forzh. Japanese (ja) usesfugashitokenization and produces audio, but it does not automatically expand numbers, dates, or percentages; callers should first write those inputs as readable Japanese text. - A request
seedcontrols Stage 0 AR sampling and per-request CFM noise. Different concurrent batch compositions do not guarantee a bit-identical waveform. - IndexTTS-2.5 Stage 0 uses plain vLLM sampling and does not reproduce the official default
num_beams=3beam search. For parity comparisons, run upstream withnum_beams=1; output quality can differ from the official beam-search result. - Emotion params (
emo_audio,emo_text,emo_vector,emo_alpha,use_emo_text,use_random) are passed via theextra_paramsfield. Official precedence isuse_emo_text>emo_vector>emo_audio> same emotion as the speaker reference. - IndexTTS-2.5 uses
vllm_omni/deploy/indextts2_5.yamlwith the official code-only Stage 0 to Stage 1 contract. - For IndexTTS-2.5, set
MODELto the local native bundle; asset discovery accepts its nestedcheckpoints/layout.
Audio8 TTS Preview¶
0.6B DualAR TTS at 44.1 kHz, 11 languages, zero-shot voice cloning.
Prerequisites¶
None beyond the base install: the neural audio codec is implemented in tree and its weights (codec.pth) ship with the checkpoint.
Launch¶
The deploy config auto-loads from vllm_omni/deploy/audio8_tts.yaml (HF model_type is arktts). Do not pass --trust-remote-code: vllm-omni registers its own arktts config, and transformers would otherwise prefer the checkpoint's remote code and bypass it.
Text-only synthesis¶
curl -X POST http://localhost:8092/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Welcome to Audio8 TTS.",
"response_format": "wav"
}' --output output.wav
Voice cloning¶
curl -X POST http://localhost:8092/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Welcome to Audio8 TTS.",
"ref_audio": "https://example.com/reference.wav",
"ref_text": "The exact transcript of the reference recording."
}' --output cloned.wav
ref_audio also accepts a data:audio/wav;base64,... URL. ref_text is mandatory whenever ref_audio is present and must match the recording. Uploading a voice via POST /v1/audio/voices lets the server reuse the encoded reference codes across requests (voice: "<name>").
Streaming¶
python audio8_tts/speech_client.py --text "Welcome to Audio8 TTS." --stream --output out.pcm
ffplay -f s16le -ar 44100 -ac 1 out.pcm
Gradio demo¶
./audio8_tts/run_gradio_demo.sh # server + demo
python audio8_tts/gradio_demo.py --api-base http://localhost:8092 # demo only
Benchmark¶
export VLLM_OMNI_BENCH_AUDIO_SAMPLE_RATE=44100 # required: harness defaults to 24 kHz
vllm bench serve --omni --host 127.0.0.1 --port 8092 \
--model Audio8/Audio8-TTS-Preview-0.6b \
--backend openai-audio-speech --endpoint /v1/audio/speech \
--dataset-name seed-tts-text --dataset-path benchmarks/build_dataset/seed_tts_smoke \
--seed-tts-locale en --num-prompts 20 --num-warmups 2 \
--max-concurrency 1 --request-rate inf \
--percentile-metrics ttft,e2el,audio_rtf,audio_ttfp,audio_duration,audio_underrun
Measured on 1x H20 shared with an unrelated job holding ~80% SM, so treat these as a lower bound. 20 prompts from seed_tts_smoke/en, mean audio 4.5 s, deploy defaults:
| HF reference (batch=1) | vllm-omni c=1 | c=4 | c=8 | |
|---|---|---|---|---|
| RTF (mean) | 1.008 | 0.19 | 0.30 | 0.54 |
| Throughput (req/s) | 0.23 | 1.18 | 2.84 | 2.93 |
| E2E latency (mean, ms) | 4375 | 850 | 1349 | 2361 |
| Time to first packet (mean, ms) | n/a (no streaming) | 63 | 105 | 1124 |
| Underrun p99 / continuity OK | n/a | 0 s / 100% | 0 s / 100% | 0 s / 100% |
~5.3x lower per-request RTF than the upstream reference implementation at concurrency 1, ~12.8x aggregate throughput at concurrency 8.
Stage-0 max_num_seqs: first-packet vs throughput¶
This subsection is operator tuning guidance, not a claim this PR optimizes anything: the shipped deploy default is unchanged (max_num_seqs: 4). The numbers below are a single run on the same contended H20 as above (one unrelated job holding ~80% SM, no repeats), so read them as directional, not as measured speedups. Same setup, varying stage 0's max_num_seqs (stage 1 stays at 1):
stage-0 max_num_seqs | client concurrency | req/s | RTF | TTFP mean / p99 (ms) | underrun p99 (s) |
|---|---|---|---|---|---|
| 4 (default) | 1 | 1.18 | 0.19 | 63 / 67 | 0.00 |
| 4 | 4 | 2.84 | 0.30 | 105 / 148 | 0.00 |
| 4 | 8 | 2.93 | 0.54 | 1124 / 1642 | 0.00 |
| 8 | 1 | 1.16 | 0.19 | 64 / 70 | 0.00 |
| 8 | 4 | 2.88 | 0.30 | 107 / 148 | 0.00 |
| 8 | 8 | 3.72 | 0.40 | 247 / 733 | 0.22 |
| 8 | 16 | 3.91 | 0.65 | 1366 / 2530 | 0.29 |
Reading of this:
- Below the limit (
concurrency <= max_num_seqs) the setting is irrelevant, as expected — c=1 and c=4 are identical for both. - At c=8, raising the limit to 8 admits every request instead of queueing half of them: in this single run, roughly +27% throughput and ~4.5x lower first-packet latency.
- Throughput saturates around 3.7-3.9 req/s (~17 s of audio per second). Past c=8 you only buy queueing delay.
- The cost of admitting more streams is a 0.22-0.29 s worst-case buffer deficit in 1 of 20 requests (
Streaming continuity OK rate95%), because stage 1 runsmax_num_seqs: 1and the codec window round-robins across streams. Client-side prebuffering (the bundled Gradio demo holds 2 chunks) absorbs that; raising stage 1'smax_num_seqsis the untested knob if you need both.
The shipped default stays at 4 because every measured point there had zero underrun. Raise it to 8 if first-packet latency under load matters more than a rare sub-300 ms gap.
Notes¶
- Output: 44.1 kHz mono; ~21.5 codec frames per second.
- No built-in speaker presets. Omit
voicefor a random timbre, or clone one. max_new_tokenscaps generated codec frames (1 frame ~= 46 ms).- Context is 2048 packed text+audio positions (
max_model_len: 2048). - Do not send
temperature: 0: greedy decoding is degenerate for this checkpoint and truncates the utterance. The deploy defaults (0.7 / top_k 50 / top_p 0.9) mirrorgeneration_config.json.
Fish Speech S2 Pro¶
4B dual-AR TTS at 44.1 kHz. Server uses the DAC codec, which is vendored in vLLM-Omni (no extra packages required).
Kvcache attention fast path¶
Fish Speech S2 Pro uses a Triton decode-only kvcache attention fast path by default on CUDA builds. Set VLLM_OMNI_FISH_KVCACHE_ATTN=0 to disable it, or VLLM_OMNI_FISH_KVCACHE_ATTN=required to fail fast if the fast path cannot be installed.
# Verify fast path availability.
python - <<'PY'
from vllm_omni.attention import fish_kvcache_attn
print(fish_kvcache_attn.is_available())
print(fish_kvcache_attn.load_error())
PY
# Optional: disable the runtime fast path.
export VLLM_OMNI_FISH_KVCACHE_ATTN=0
Launch¶
The deploy config auto-loads from vllm_omni/deploy/fish_qwen3_omni.yaml (the HF model_type on the fishaudio checkpoint is fish_qwen3_omni).
Voice cloning¶
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Hello, this is a cloned voice.",
"voice": "default",
"ref_audio": "https://example.com/reference.wav",
"ref_text": "Transcript of the reference audio."
}' --output cloned.wav
CLI client¶
cd examples/online_serving/text_to_speech/fish_speech
python speech_client.py --text "Hello, how are you?"
python speech_client.py --text "Hello world" --stream --output output.pcm
Gradio demo¶
./fish_speech/run_gradio_demo.sh # launches server + Gradio
python fish_speech/gradio_demo.py --api-base http://localhost:8091 # if server already running
Notes¶
- Output: 44.1 kHz mono.
- Streaming PCM player command must use
-r 44100.
Ming-omni-tts¶
Dense 0.5B two-stage TTS served through /v1/audio/speech. Ming uses the standard speech endpoint plus structured controls in instructions, voice, language, ref_audio, ref_text, and speaker_embedding.
Launch¶
Equivalent manual command:
vllm-omni serve inclusionAI/Ming-omni-tts-0.5B \
--deploy-config vllm_omni/deploy/ming_tts.yaml \
--host 0.0.0.0 --port 8091 \
--enforce-eager --omni
Sending requests¶
python examples/online_serving/text_to_speech/ming_tts/openai_speech_client.py \
--text "你好,这是 Ming 在线语音合成测试。"
Structured dialect control:
python examples/online_serving/text_to_speech/ming_tts/openai_speech_client.py \
--text "我觉得社会企业同个人都有责任" \
--instruction-json '{"方言":"广粤话"}' \
--ref-audio /path/to/yue_prompt.wav
Zero-shot cloning:
python examples/online_serving/text_to_speech/ming_tts/openai_speech_client.py \
--text "我们的愿景是构建未来服务业的数字化基础设施,为世界带来更多微小而美好的改变。" \
--ref-audio /path/to/10002287-00000094.wav \
--ref-text "在此奉劝大家别乱打美白针。"
Notes¶
run_curl.shkeeps a small sanity subset; use the Ming README for the broader request cookbook.- Online serving is speech-shaped today; music-only
bgmand text-to-audiottaremain offline examples. - Full request details live in
ming_tts/README.md.
Ming-flash-omni-TTS¶
Standalone talker-only deployment of Ming-flash-omni-2.0. Voice is controlled through caption text passed via instructions.
Launch¶
Equivalent manual command:
vllm serve Jonathan1909/Ming-flash-omni-2.0 \
--deploy-config vllm_omni/deploy/ming_flash_omni_tts.yaml \
--host 0.0.0.0 --port 8091 \
--trust-remote-code --omni
Sending requests¶
python examples/online_serving/text_to_speech/ming_flash_omni_tts/speech_client.py \
--text "我们当迎着阳光辛勤耕作,去摘取,去制作,去品尝,去馈赠。" \
--output ming_online.wav
ASMR-style caption via instructions:
python examples/online_serving/text_to_speech/ming_flash_omni_tts/speech_client.py \
--text "我会一直在这里陪着你,直到你慢慢、慢慢地沉入那个最温柔的梦里……好吗?" \
--instructions "这是一种ASMR耳语,属于一种旨在引发特殊感官体验的创意风格。这个女性使用轻柔的普通话进行耳语,声音气音成分重。" \
--output ming_online_asmr.wav
Notes¶
- Server uses
use_zero_spk_emb=Trueand the cookbook decode defaults (max_decode_steps=200,cfg=2.0,sigma=0.25,temperature=0.0). For other caption fields (语速,基频,IP, BGM, etc.) or overriding decode args, use the offline example whereadditional_informationis set explicitly. - This is the online counterpart of
examples/offline_inference/text_to_speech/ming_flash_omni_tts/. - For multimodal Ming-flash-omni online serving, see
examples/online_serving/ming_flash_omni/.
MOSS-TTS Local Transformer v1.5¶
For a single H200, the optional moss_tts_local_h200.yaml deployment places both the talker and codec on logical GPU 0:
CUDA_VISIBLE_DEVICES=0 VLLM_OMNI_EVENT_DRIVEN_ORCH=1 \
vllm serve OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5 \
--omni --trust-remote-code \
--deploy-config vllm_omni/deploy/moss_tts_local_h200.yaml \
--disable-log-stats
Run from the repository root and select an available physical GPU with CUDA_VISIBLE_DEVICES. The command enables event-driven orchestration and disables detailed per-request statistics logging to reduce CPU overhead at high concurrency. Remove --disable-log-stats when those statistics are needed; keep these settings identical when comparing performance. This preset configures 256 request slots per stage, a 32 GiB talker KV cache, talker CUDA Graph buckets through 512 scheduled tokens, and codec batch buckets through 256. It requires more than 80 GiB of GPU memory; request capacity also depends on input and generated lengths. The default moss_tts_local.yaml remains available for smaller deployments.
The preset enables the optional codec backend with hf_overrides.codec_attention_backend: triton on stage 1. It preserves the streaming ring-cache mask and uses BF16 attention with 64-dimensional heads; other attention shapes use PyTorch SDPA. Set the backend to sdpa to use the default attention implementation. Codec terminal tails share execution only when they already map to the same padded CUDA Graph; returned audio retains each request's actual length.
Voice cloning requests use ref_audio and ref_text. For streaming output, set stream: true, stream_format: "audio", and response_format: "pcm"; the native PCM format is 48 kHz, stereo, signed 16-bit little-endian. Compute audio throughput as PCM bytes / (48000 * 2 * 2) / elapsed seconds, and compare the same requests, concurrency, and reference-cache state.
MOSS-TTS-Nano¶
Single-stage 0.1B AR LM + MOSS-Audio-Tokenizer-Nano codec at 48 kHz mono. Every request must include ref_audio; there are no built-in speaker presets.
The OpenAI-schema
voiceandref_textfields are accepted but ignored —voice_clonedoes not consume a transcript, and upstream'scontinuationmode (the only path that acceptsprompt_text) emits near-silent output, so it is not exposed here. Sample reference clips ship in the upstream repo underassets/audio/.
Launch¶
The deploy config at vllm_omni/deploy/moss_tts_nano.yaml auto-loads; no --deploy-config, --trust-remote-code, or --enforce-eager flags are needed.
Sending requests¶
# One-off fetch of a sample reference clip; cache under XDG_CACHE_HOME.
REF_DIR="${XDG_CACHE_HOME:-$HOME/.cache}/moss-tts-nano"
mkdir -p "$REF_DIR"
REF_WAV="$REF_DIR/zh_1.wav"
[ -s "$REF_WAV" ] || curl -L -o "$REF_WAV" https://raw.githubusercontent.com/OpenMOSS/MOSS-TTS-Nano/main/assets/audio/zh_1.wav
REF_AUDIO=$(base64 -w 0 "$REF_WAV")
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d "{
\"input\": \"你好,这是语音合成测试。\",
\"ref_audio\": \"data:audio/wav;base64,${REF_AUDIO}\",
\"response_format\": \"wav\"
}" --output output.wav
Streaming PCM¶
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d "{
\"input\": \"Hello, streaming output from MOSS-TTS-Nano.\",
\"ref_audio\": \"data:audio/wav;base64,${REF_AUDIO}\",
\"stream\": true,
\"stream_format\": \"audio\",
\"response_format\": \"pcm\"
}" --no-buffer | play -t raw -r 48000 -e signed -b 16 -c 1 -
Gradio demo¶
# Option 1: launch server + Gradio together
./moss_tts_nano/run_gradio_demo.sh
# Option 2: server already running
python moss_tts_nano/gradio_demo.py --api-base http://localhost:8091
Then open http://localhost:7860 in your browser.
Notes¶
- Output is 48 kHz mono PCM (the upstream tokenizer is internally stereo at 48 kHz; the wrapper averages to mono before reaching the engine).
- Standard
/v1/audio/speechrequest shape:input,ref_audio(base64 data URL),response_format,stream,max_new_tokens. Thevoiceandref_textfields from the OpenAI schema are accepted but ignored.
OmniVoice¶
Zero-shot multilingual TTS (600+ languages). Online serving currently exposes auto voice only; voice cloning and voice design are available offline.
Prerequisites¶
Voice cloning (offline) needs transformers>=5.3.0; auto voice works with transformers>=4.57.0.
Launch¶
CLI client¶
cd examples/online_serving/text_to_speech/omnivoice
# Text-only (auto voice)
python speech_client.py --text "Hello, how are you?"
# Language hint
python speech_client.py --text "Bonjour, comment allez-vous?" --language French
# Voice cloning (reference audio + optional ref_text)
python speech_client.py \
--text "Bonjour, comment allez-vous?" \
--ref-audio /path/to/ref_audio.wav \
--ref-text "Bonjour, comment allez-vous?"
# Style instruction (voice design-style control)
python speech_client.py \
--text "Bonjour, comment allez-vous?" \
--language French \
--instructions "loud voice"
# Deterministic output with seed parameter
python speech_client.py --text "Hello, how are you?" --seed 42
The client supports --api-base, --model, --text, --response-format, --language, --voice, --ref-audio, --ref-text, --instructions, --seed, and --output.
Qwen3-TTS¶
Three model variants exposed via separate checkpoints:
| Variant | HF repo | Use |
|---|---|---|
| CustomVoice | Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice | Predefined speakers (vivian, ryan, …) with optional style instructions |
| VoiceDesign | Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign | Natural-language voice style description |
| Base | Qwen/Qwen3-TTS-12Hz-1.7B-Base | Voice cloning from a reference audio |
Each variant ships smaller 0.6B companions where available.
Launch¶
vllm serve Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --omni --port 8091
# or:
./qwen3_tts/run_server.sh # default: CustomVoice
./qwen3_tts/run_server.sh VoiceDesign
./qwen3_tts/run_server.sh Base
For a local deployment with multiple API frontend processes sharing one set of TTS stage engines, add --api-server-count:
This mode supports local EngineCore stages only; headless or remote stage deployments are not supported. Runtime voice upload and deletion are disabled with multiple API frontends because their registries are process-local. Built-in voices, inline ref_audio, and voices restored from custom_voice_dir at startup are supported.
Executor backend¶
Single-GPU serves now default to the uniproc executor (lower IPC overhead, the Base cloning use case from #2603 / #2604). vllm_omni/deploy/qwen3_tts.yaml is the only Qwen3-TTS deploy config; pass --deploy-config <path> to override.
To opt out of chunked streaming, pass --no-async-chunk — the pipeline auto-dispatches to the end-to-end codec processor.
Sending requests¶
# CustomVoice with a predefined speaker
python qwen3_tts/openai_speech_client.py \
--model Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
--text "今天天气真好" \
--speaker ryan \
--instructions "用开心的语气说"
# VoiceDesign with a style description
python qwen3_tts/openai_speech_client.py \
--model Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign \
--task-type VoiceDesign \
--text "哥哥,你回来啦" \
--instructions "体现撒娇稚嫩的萝莉女声,音调偏高"
# Base voice cloning
python qwen3_tts/openai_speech_client.py \
--model Qwen/Qwen3-TTS-12Hz-1.7B-Base \
--task-type Base \
--text "Hello, this is a cloned voice" \
--ref-audio /path/to/reference.wav \
--ref-text "Original transcript of the reference audio"
Voices endpoint¶
List available voices, or upload a custom one for Base cloning:
# List
curl http://localhost:8091/v1/audio/voices
# Upload
curl -X POST http://localhost:8091/v1/audio/voices \
-F "audio_sample=@/path/to/voice_sample.wav" \
-F "consent=user_consent_id" \
-F "name=custom_voice_1" \
-F "ref_text=The exact transcript of the audio sample." \
-F "speaker_description=warm narrator"
For Qwen3-TTS, uploaded voices are Base voice-cloning inputs and require a Base checkpoint. When a request names an uploaded voice, the server infers task_type="Base". Built-in presets such as vivian and ryan remain CustomVoice speakers and require a CustomVoice checkpoint.
The runtime upload and delete routes require a single API frontend; with --api-server-count > 1, use inline ref_audio or restore precomputed voices from custom_voice_dir at startup.
Precomputed custom voices¶
For reused Base voice-cloning speakers, precompute the reference artifacts once and load them at server startup. Precomputed voices use the same Base task and checkpoint-matching rules as uploaded voices:
python qwen3_tts/precompute_custom_voice.py \
--model Qwen/Qwen3-TTS-12Hz-1.7B-Base \
--voice-name alice \
--ref-audio /path/to/reference.wav \
--ref-text "Original transcript of the reference audio" \
--mode icl \
--output-dir /path/to/custom_voices
--mode icl stores both speaker_embedding and ref_code; --mode xvec stores only the speaker embedding. Add the output directory to a deploy config:
Then start the server with that config and call the Speech API with only the voice name:
vllm serve Qwen/Qwen3-TTS-12Hz-1.7B-Base --omni --deploy-config /path/to/qwen3_tts_custom_voice.yaml
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input":"Hello from a precomputed voice.","voice":"alice","task_type":"Base"}' \
--output alice.wav
Streaming PCM¶
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Hello, how are you?",
"voice": "vivian",
"language": "English",
"stream": true,
"stream_format": "audio",
"response_format": "pcm"
}' --no-buffer | play -t raw -r 24000 -e signed -b 16 -c 1 -
Raw PCM streaming requires stream_format="audio", response_format="pcm", and async_chunk: true in the deploy config (default in qwen3_tts.yaml). speed is not supported when streaming.
Streaming WebSocket¶
The /v1/audio/speech/stream endpoint accepts text incrementally and, by default, synthesizes the buffered text as one continuous request on input.done:
python qwen3_tts/streaming_speech_client.py --text "Hello world. How are you? I am fine."
python qwen3_tts/streaming_speech_client.py --text "..." --simulate-stt --stt-delay 0.1
For per-sentence audio (including Indic danda ।) before input.done:
python qwen3_tts/streaming_speech_client.py \
--text "नमस्ते। कैसे हो?" \
--split-granularity sentence
input.done flushes without closing, so repeating --text synthesizes several utterances over one connection:
python qwen3_tts/streaming_speech_client.py \
--text "First utterance." \
--text "Second utterance, same connection."
To receive word-level timestamps, launch the server with a forced aligner:
vllm-omni serve Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
--omni \
--deploy-config vllm_omni/deploy/qwen3_tts.yaml \
--trust-remote-code \
--forced-aligner Qwen/Qwen3-ForcedAligner-0.6B
Then request PCM JSON sidecar chunks:
python qwen3_tts/streaming_speech_client.py \
--text "Hello world. How are you?" \
--stream-audio \
--response-format pcm \
--word-timestamps
The client writes one PCM file per sentence and a matching sentence_XXX_timestamps.json sidecar.
Non-streaming requests can also ask for timestamps: pass "word_timestamps": true to POST /v1/audio/speech and read the X-Word-Timestamps response header (JSON list of {word, start_ms, end_ms}, ASCII-escaped). Past 4 KB the header is replaced by X-Word-Timestamps-Omitted: oversize; bytes=<n>; limit=4096 and the audio still returns — use the WebSocket path for long transcripts.
To see the alignment instead of reading a JSON sidecar, run the word-timestamp Gradio demo (server must be launched with --forced-aligner):
Each sentence's audio plays in an <audio> element while its text is rendered as inline word spans; the current word highlights as audio.currentTime crosses each start_ms. The Stop (barge-in) button cuts playback and reports the last-spoken word, useful for the voice-agent barge-in case.
Gradio demos¶
./qwen3_tts/run_gradio_demo.sh # CustomVoice (default)
./qwen3_tts/run_gradio_demo.sh --task-type VoiceDesign
./qwen3_tts/run_gradio_demo.sh --task-type Base
Speaker embedding interpolation¶
qwen3_tts/speaker_embedding_interpolation.py blends two predefined speakers' embeddings to produce intermediate voices. See the script for usage.
Batch client¶
qwen3_tts/batch_speech_client.py issues many concurrent requests for throughput measurement.
Notes¶
- Base voice cloning has uniproc-vs-mp tradeoffs depending on per-request reference audio cost; see the executor-backend section above.
- With async chunking, Qwen3-TTS Base voice cloning sends the full reference context in the first Code2Wav packet, then caches that prefix on the Code2Wav stage for follow-up chunks in the same request.
vllm_omni/deploy/qwen3_tts.yamlis the default deploy config (loaded by HFmodel_type); per-stage runtime overrides are available via--stage-N-<field> <value>.
VoxCPM2¶
Single-stage native AR TTS at 48 kHz.
Launch¶
Deploy config auto-loads from vllm_omni/deploy/voxcpm2.yaml. Pass --deploy-config <path> to override or --stage-N-<field> <value> for per-stage runtime tweaks.
Sending requests¶
# Zero-shot synthesis
python voxcpm2/openai_speech_client.py --text "Hello, this is VoxCPM2."
# Voice cloning
python voxcpm2/openai_speech_client.py \
--text "This should sound like the reference speaker." \
--ref-audio /path/to/reference.wav
The ref_audio field accepts local file paths (auto-base64), HTTP URLs, or data:audio/wav;base64,... data URIs.
Precomputed custom voices¶
For repeated VoxCPM2 speakers, precompute the prompt cache and load it through custom_voice_dir:
python voxcpm2/precompute_custom_voice.py \
--model openbmb/VoxCPM2 \
--voice-name alice \
--ref-audio /path/to/reference.wav \
--mode ref_continuation \
--prompt-text "Original transcript of the reference audio" \
--output-dir /path/to/custom_voices
Add the output directory to the deploy config:
After startup, /v1/audio/voices lists alice, and /v1/audio/speech can use voice="alice" without sending ref_audio.
Gradio demo (gapless streaming via AudioWorklet)¶
Uses an AudioWorklet-based player adapted from the Qwen3-TTS demo for gap-free playback. Raw PCM audio is streamed from the OpenAI Speech endpoint with stream=true and stream_format="audio".
Voxtral TTS¶
Voxtral-4B-TTS (Mistral). Uses the mistral_common SpeechRequest protocol; voice presets are model-specific.
Prerequisites¶
Latest mistral_common with SpeechRequest support:
Launch¶
Deploy config auto-loads from vllm_omni/deploy/voxtral_tts.yaml.
Gradio demo¶
The demo handles voice-preset selection and reference-audio upload. voxtral_tts/text_preprocess.py provides the text-normalization helpers used by the demo (also available for other clients).
Notes¶
- Voice presets are listed on the HF model card (
mistralai/Voxtral-4B-TTS-2603). - Voice cloning is gated upstream and may require a recent
mistral_common. - A standalone CLI client is not yet shipped; the gradio demo is the canonical reference for now.
Breeze-TTS-2¶
Two-stage AR TTS (T5Gemma2 + Qwen3 talker with a depth decoder → bundled Qwen3-TTS codec) at 24 kHz mono, from BreezeBlue. Four modes are selected automatically from the request fields: plain (input), voice design (instructions), clone (ref_audio+ref_text), and voice direction (reference trio + instructions).
Launch¶
The deploy config at vllm_omni/deploy/breeze_tts_2.yaml auto-loads (async-chunk streaming with 1, 2, 4, then 5-frame inter-stage chunks).
Sending requests¶
# Plain synthesis (speaker tags S0..S9, default S0)
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Hello, this is Breeze TTS 2 on vLLM-Omni.",
"voice": "S0",
"response_format": "wav",
"sample_rate": 24000
}' --output breeze_plain.wav
# Voice direction (clone a reference, then steer the delivery)
curl -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "We need to discuss what happened last night.",
"ref_audio": "file:///path/to/reference.wav",
"ref_text": "The exact transcript of the reference audio.",
"instructions": "Speak slowly with a restrained, serious tone.",
"response_format": "wav"
}' --output breeze_direction.wav
Streaming PCM¶
curl -N -X POST http://localhost:8091/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Streaming output from Breeze TTS 2.",
"stream": true,
"stream_format": "audio",
"response_format": "pcm",
"sample_rate": 24000
}' --output breeze_stream.pcm
Notes¶
- Only
sample_rate=24000is accepted;speedmust be1.0. guidance_scale/cfg_scalemust be1.0;negative_promptis rejected (CFG companion support is a follow-up).instructionswithoutref_audiois voice design; withref_audio+ref_textit is voice direction.- See
recipes/BreezeBlue/Breeze-TTS-2.mdfor the full recipe and the offline example underexamples/offline_inference/text_to_speech/breeze_tts_2/.
Example materials¶
audio8_tts/gradio_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/audio8_tts/gradio_demo.py.
audio8_tts/run_gradio_demo.sh
#!/bin/bash
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
# Launch the Audio8 TTS Preview server + Gradio demo together.
#
# Usage:
# ./run_gradio_demo.sh
# CUDA_VISIBLE_DEVICES=0 PORT=8092 GRADIO_PORT=7861 ./run_gradio_demo.sh
set -e
MODEL="${MODEL:-Audio8/Audio8-TTS-Preview-0.6b}"
PORT="${PORT:-8092}"
GRADIO_PORT="${GRADIO_PORT:-7861}"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
echo "Starting Audio8 TTS Preview server (port $PORT)..."
vllm serve "$MODEL" \
--omni \
--host 0.0.0.0 \
--port "$PORT" &
SERVER_PID=$!
cleanup() {
echo "Stopping server (PID $SERVER_PID)..."
kill $SERVER_PID 2>/dev/null
wait $SERVER_PID 2>/dev/null
}
trap cleanup EXIT
echo "Waiting for server to start..."
for _ in $(seq 1 120); do
if curl -s "http://localhost:$PORT/health" > /dev/null 2>&1; then
echo "Server ready."
break
fi
sleep 2
done
echo "Starting Gradio demo (port $GRADIO_PORT)..."
python "$SCRIPT_DIR/gradio_demo.py" \
--api-base "http://localhost:$PORT" \
--port "$GRADIO_PORT"
audio8_tts/run_server.sh
#!/bin/bash
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
# Launch a vLLM-Omni server for Audio8 TTS Preview 0.6B.
#
# Usage:
# ./run_server.sh
# CUDA_VISIBLE_DEVICES=0 PORT=8092 ./run_server.sh
set -e
MODEL="${MODEL:-Audio8/Audio8-TTS-Preview-0.6b}"
PORT="${PORT:-8092}"
echo "Starting Audio8 TTS Preview server with model: $MODEL"
# --trust-remote-code is deliberately NOT passed: vllm-omni registers its own
# `arktts` config, and transformers would otherwise prefer the checkpoint's
# remote code and bypass it.
vllm serve "$MODEL" \
--omni \
--host 0.0.0.0 \
--port "$PORT"
audio8_tts/speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
"""Client for Audio8 TTS Preview via the /v1/audio/speech endpoint.
Examples:
# Basic TTS
python speech_client.py --text "Welcome to Audio8 TTS."
# Zero-shot voice cloning (the transcript must match the reference audio)
python speech_client.py --text "Welcome to Audio8 TTS." \
--ref-audio ref.wav --ref-text "The exact transcript of the reference recording."
# Streaming PCM output
python speech_client.py --text "Welcome to Audio8 TTS." --stream --output output.pcm
"""
import argparse
import base64
import os
import httpx
DEFAULT_API_BASE = "http://localhost:8092"
DEFAULT_API_KEY = "EMPTY"
DEFAULT_MODEL = "Audio8/Audio8-TTS-Preview-0.6b"
#: The codec runs at 44.1 kHz, so raw PCM must be played back at that rate.
PCM_SAMPLE_RATE = 44100
def encode_audio_to_base64(audio_path: str) -> str:
"""Encode a local audio file as a base64 data URL."""
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
ext = audio_path.lower().rsplit(".", 1)[-1]
mime_type = {"wav": "audio/wav", "mp3": "audio/mpeg", "flac": "audio/flac", "ogg": "audio/ogg"}.get(
ext, "audio/wav"
)
with open(audio_path, "rb") as handle:
audio_b64 = base64.b64encode(handle.read()).decode("utf-8")
return f"data:{mime_type};base64,{audio_b64}"
def build_payload(args) -> dict:
payload = {
"model": args.model,
"input": args.text,
"response_format": args.response_format,
}
if args.ref_audio:
if args.ref_audio.startswith(("http://", "https://")):
payload["ref_audio"] = args.ref_audio
else:
payload["ref_audio"] = encode_audio_to_base64(args.ref_audio)
if args.ref_text:
payload["ref_text"] = args.ref_text
if args.max_new_tokens is not None:
payload["max_new_tokens"] = args.max_new_tokens
if args.stream:
payload["stream"] = True
payload["stream_format"] = "audio"
payload["response_format"] = "pcm"
return payload
def run_tts(args) -> None:
payload = build_payload(args)
api_url = f"{args.api_base}/v1/audio/speech"
headers = {"Content-Type": "application/json", "Authorization": f"Bearer {args.api_key}"}
print(f"Model: {args.model}")
print(f"Text: {args.text}")
if args.ref_audio:
print(f"Voice cloning: ref_audio={args.ref_audio!r} ref_text={args.ref_text!r}")
if args.stream:
output_path = args.output or "output.pcm"
with (
httpx.Client(timeout=300.0) as client,
client.stream("POST", api_url, json=payload, headers=headers) as resp,
):
if resp.status_code != 200:
print(f"Error {resp.status_code}: {resp.read().decode()}")
return
total_bytes = 0
with open(output_path, "wb") as handle:
for chunk in resp.iter_bytes():
handle.write(chunk)
total_bytes += len(chunk)
seconds = total_bytes / 2 / PCM_SAMPLE_RATE
print(f"Streamed {total_bytes} bytes (~{seconds:.2f}s of s16le @ {PCM_SAMPLE_RATE} Hz) to {output_path}")
print(f"Play with: ffplay -f s16le -ar {PCM_SAMPLE_RATE} -ac 1 {output_path}")
return
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error {response.status_code}: {response.text}")
return
if response.headers.get("content-type", "").startswith("application/json"):
print(f"Error: {response.text}")
return
output_path = args.output or f"output.{args.response_format}"
with open(output_path, "wb") as handle:
handle.write(response.content)
print(f"Audio saved to: {output_path} ({len(response.content)} bytes)")
def main() -> None:
parser = argparse.ArgumentParser(description="Audio8 TTS Preview client")
parser.add_argument("--api-base", default=DEFAULT_API_BASE, help="API base URL")
parser.add_argument("--api-key", default=DEFAULT_API_KEY, help="API key")
parser.add_argument("--model", "-m", default=DEFAULT_MODEL, help="Model name")
parser.add_argument("--text", required=True, help="Text to synthesize")
parser.add_argument("--ref-audio", default=None, help="Reference audio for voice cloning (path or URL)")
parser.add_argument("--ref-text", default=None, help="Transcript of the reference audio")
parser.add_argument("--max-new-tokens", type=int, default=None, help="Cap the number of generated codec frames")
parser.add_argument("--stream", action="store_true", help="Stream PCM instead of returning a whole file")
parser.add_argument(
"--response-format",
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio format (default: wav)",
)
parser.add_argument("--output", "-o", default=None, help="Output file path")
run_tts(parser.parse_args())
if __name__ == "__main__":
main()
breeze_tts_2/run_server.sh
#!/usr/bin/env bash
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
# Launch Breeze-TTS-2 online serving with the async-chunk deploy config.
set -euo pipefail
MODEL="${1:-BreezeBlue/Breeze-TTS-2}"
PORT="${PORT:-8091}"
exec vllm-omni serve "${MODEL}" \
--deploy-config vllm_omni/deploy/breeze_tts_2.yaml \
--omni --port "${PORT}"
cosyvoice3/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for CosyVoice3 TTS
#
# Usage:
# ./run_server.sh
# CUDA_VISIBLE_DEVICES=0 ./run_server.sh
#
# Streaming (async-chunk) is on by default via vllm_omni/deploy/cosyvoice3.yaml.
# Set NO_ASYNC_CHUNK=1 to use the legacy synchronous path.
set -e
MODEL="${MODEL:-FunAudioLLM/Fun-CosyVoice3-0.5B-2512}"
PORT="${PORT:-8091}"
EXTRA_ARGS=()
if [[ -n "${NO_ASYNC_CHUNK:-}" ]]; then
EXTRA_ARGS+=(--no-async-chunk)
fi
echo "Starting CosyVoice3 server with model: $MODEL"
vllm serve "$MODEL" \
--host 0.0.0.0 \
--port "$PORT" \
--trust-remote-code \
--omni \
"${EXTRA_ARGS[@]}"
cosyvoice3/speech_client.py
"""Client for CosyVoice3 TTS via /v1/audio/speech endpoint.
CosyVoice3 has no built-in voice presets: every request is voice cloning
driven by ``ref_audio`` + ``ref_text``. The defaults below point at the
official upstream zero-shot prompt so the script runs out of the box.
Examples:
# Voice cloning with the default upstream prompt
python speech_client.py --text "收到好友从远方寄来的生日礼物。"
# Custom reference clip + transcript
python speech_client.py --text "Hello, this is a cloned voice." \
--ref-audio /path/to/reference.wav \
--ref-text "Transcript of the reference audio."
# Streaming PCM output
python speech_client.py --text "Hello world" --stream --output output.pcm
"""
import argparse
import base64
import os
import httpx
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
DEFAULT_MODEL = "FunAudioLLM/Fun-CosyVoice3-0.5B-2512"
# Official CosyVoice zero-shot prompt and its transcript.
DEFAULT_REF_AUDIO = "https://raw.githubusercontent.com/FunAudioLLM/CosyVoice/main/asset/zero_shot_prompt.wav"
DEFAULT_REF_TEXT = "希望你以后能够做的比我还好呦。"
def encode_audio_to_base64(audio_path: str) -> str:
"""Encode a local audio file to a base64 data URL."""
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
ext = audio_path.lower().rsplit(".", 1)[-1]
mime_map = {"wav": "audio/wav", "mp3": "audio/mpeg", "flac": "audio/flac", "ogg": "audio/ogg"}
mime_type = mime_map.get(ext, "audio/wav")
with open(audio_path, "rb") as f:
audio_b64 = base64.b64encode(f.read()).decode("utf-8")
return f"data:{mime_type};base64,{audio_b64}"
def run_tts(args) -> None:
"""Generate speech via the /v1/audio/speech API."""
payload = {
"model": args.model,
"input": args.text,
"response_format": args.response_format,
}
if args.ref_audio.startswith(("http://", "https://")):
payload["ref_audio"] = args.ref_audio
else:
payload["ref_audio"] = encode_audio_to_base64(args.ref_audio)
payload["ref_text"] = args.ref_text
if args.stream:
payload["stream"] = True
payload["stream_format"] = "audio"
payload["response_format"] = "pcm"
print(f"Model: {args.model}")
print(f"Text: {args.text}")
print(f"Voice cloning: ref_audio={args.ref_audio}, ref_text={args.ref_text}")
print("Generating audio...")
api_url = f"{args.api_base}/v1/audio/speech"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
if args.stream:
output_path = args.output or "output.pcm"
with httpx.Client(timeout=300.0) as client:
with client.stream("POST", api_url, json=payload, headers=headers) as resp:
if resp.status_code != 200:
print(f"Error: {resp.status_code}")
print(resp.read().decode())
return
total_bytes = 0
with open(output_path, "wb") as f:
for chunk in resp.iter_bytes():
f.write(chunk)
total_bytes += len(chunk)
print(f"Streamed {total_bytes} bytes to: {output_path}")
else:
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return
try:
text = response.content.decode("utf-8")
if text.startswith('{"error"'):
print(f"Error: {text}")
return
except UnicodeDecodeError:
pass
output_path = args.output or "output.wav"
with open(output_path, "wb") as f:
f.write(response.content)
print(f"Audio saved to: {output_path}")
def main():
parser = argparse.ArgumentParser(description="CosyVoice3 TTS client")
parser.add_argument("--api-base", default=DEFAULT_API_BASE, help="API base URL")
parser.add_argument("--api-key", default=DEFAULT_API_KEY, help="API key")
parser.add_argument("--model", "-m", default=DEFAULT_MODEL, help="Model name")
parser.add_argument("--text", required=True, help="Text to synthesize")
parser.add_argument(
"--ref-audio",
default=DEFAULT_REF_AUDIO,
help="Reference audio for voice cloning (path or URL)",
)
parser.add_argument(
"--ref-text",
default=DEFAULT_REF_TEXT,
help="Transcript of the reference audio",
)
parser.add_argument("--stream", action="store_true", help="Enable streaming (PCM output)")
parser.add_argument(
"--response-format",
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio format (default: wav)",
)
parser.add_argument("--output", "-o", default=None, help="Output file path")
args = parser.parse_args()
run_tts(args)
if __name__ == "__main__":
main()
fish_speech/gradio_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/fish_speech/gradio_demo.py.
fish_speech/run_gradio_demo.sh
#!/bin/bash
# Launch Fish Speech S2 Pro server + Gradio demo together.
#
# Usage:
# ./run_gradio_demo.sh
# CUDA_VISIBLE_DEVICES=0 PORT=8091 GRADIO_PORT=7860 ./run_gradio_demo.sh
set -e
MODEL="${MODEL:-fishaudio/s2-pro}"
PORT="${PORT:-8091}"
GRADIO_PORT="${GRADIO_PORT:-7860}"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
echo "Starting Fish Speech S2 Pro server (port $PORT)..."
FLASHINFER_DISABLE_VERSION_CHECK=1 \
vllm serve "$MODEL" \
--omni \
--host 0.0.0.0 \
--port "$PORT" &
SERVER_PID=$!
cleanup() {
echo "Stopping server (PID $SERVER_PID)..."
kill $SERVER_PID 2>/dev/null
wait $SERVER_PID 2>/dev/null
}
trap cleanup EXIT
# Wait for server to be ready.
echo "Waiting for server to start..."
for i in $(seq 1 120); do
if curl -s "http://localhost:$PORT/health" > /dev/null 2>&1; then
echo "Server ready."
break
fi
sleep 2
done
echo "Starting Gradio demo (port $GRADIO_PORT)..."
python "$SCRIPT_DIR/gradio_demo.py" \
--api-base "http://localhost:$PORT" \
--port "$GRADIO_PORT"
fish_speech/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for Fish Speech S2 Pro
#
# Usage:
# ./run_server.sh
# CUDA_VISIBLE_DEVICES=0 ./run_server.sh
set -e
MODEL="${MODEL:-fishaudio/s2-pro}"
PORT="${PORT:-8091}"
echo "Starting Fish Speech S2 Pro server with model: $MODEL"
FLASHINFER_DISABLE_VERSION_CHECK=1 \
vllm serve "$MODEL" \
--omni \
--host 0.0.0.0 \
--port "$PORT"
fish_speech/speech_client.py
"""Client for Fish Speech S2 Pro via /v1/audio/speech endpoint.
Examples:
# Basic TTS
python speech_client.py --text "Hello, how are you?"
# Voice cloning
python speech_client.py --text "Hello, how are you?" \
--ref-audio ref.wav --ref-text "This is the reference transcript."
# Streaming PCM output
python speech_client.py --text "Hello world" --stream --output output.pcm
"""
import argparse
import base64
import os
import httpx
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
def encode_audio_to_base64(audio_path: str) -> str:
"""Encode a local audio file to base64 data URL."""
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
ext = audio_path.lower().rsplit(".", 1)[-1]
mime_map = {"wav": "audio/wav", "mp3": "audio/mpeg", "flac": "audio/flac", "ogg": "audio/ogg"}
mime_type = mime_map.get(ext, "audio/wav")
with open(audio_path, "rb") as f:
audio_b64 = base64.b64encode(f.read()).decode("utf-8")
return f"data:{mime_type};base64,{audio_b64}"
def run_tts(args) -> None:
"""Generate speech via /v1/audio/speech API."""
payload = {
"model": args.model,
"input": args.text,
"response_format": args.response_format,
}
# Voice cloning parameters.
if args.ref_audio:
if args.ref_audio.startswith(("http://", "https://")):
payload["ref_audio"] = args.ref_audio
else:
payload["ref_audio"] = encode_audio_to_base64(args.ref_audio)
if args.ref_text:
payload["ref_text"] = args.ref_text
if args.stream:
payload["stream"] = True
payload["stream_format"] = "audio"
payload["response_format"] = "pcm"
print(f"Model: {args.model}")
print(f"Text: {args.text}")
if args.ref_audio:
print(f"Voice cloning: ref_audio={args.ref_audio}, ref_text={args.ref_text}")
print("Generating audio...")
api_url = f"{args.api_base}/v1/audio/speech"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
if args.stream:
output_path = args.output or "output.pcm"
with httpx.Client(timeout=300.0) as client:
with client.stream("POST", api_url, json=payload, headers=headers) as resp:
if resp.status_code != 200:
print(f"Error: {resp.status_code}")
print(resp.read().decode())
return
total_bytes = 0
with open(output_path, "wb") as f:
for chunk in resp.iter_bytes():
f.write(chunk)
total_bytes += len(chunk)
print(f"Streamed {total_bytes} bytes to: {output_path}")
else:
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return
try:
text = response.content.decode("utf-8")
if text.startswith('{"error"'):
print(f"Error: {text}")
return
except UnicodeDecodeError:
pass
output_path = args.output or "output.wav"
with open(output_path, "wb") as f:
f.write(response.content)
print(f"Audio saved to: {output_path}")
def main():
parser = argparse.ArgumentParser(description="Fish Speech S2 Pro TTS client")
parser.add_argument("--api-base", default=DEFAULT_API_BASE, help="API base URL")
parser.add_argument("--api-key", default=DEFAULT_API_KEY, help="API key")
parser.add_argument("--model", "-m", default="fishaudio/s2-pro", help="Model name")
parser.add_argument("--text", required=True, help="Text to synthesize")
parser.add_argument("--ref-audio", default=None, help="Reference audio for voice cloning (path or URL)")
parser.add_argument("--ref-text", default=None, help="Transcript of reference audio")
parser.add_argument("--stream", action="store_true", help="Enable streaming (PCM output)")
parser.add_argument(
"--response-format",
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio format (default: wav)",
)
parser.add_argument("--output", "-o", default=None, help="Output file path")
args = parser.parse_args()
run_tts(args)
if __name__ == "__main__":
main()
gepard/run_server.sh
#!/bin/bash
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
# Launch vLLM-Omni server for Gepard-1.0 TTS (zero-shot default voice).
#
# Usage:
# ./run_server.sh
# CUDA_VISIBLE_DEVICES=0 ./run_server.sh
#
# Packaged deploy config: vllm_omni/deploy/gepard.yaml (async_chunk=false,
# 22.05 kHz mono, default seed 42 until the YAML seed is removed).
# Cold start downloads the talker and NeMo NanoCodec; the serve default of
# --stage-init-timeout 300 is too tight (offline end2end.py uses 900).
#
# On a host whose CUDA toolkit cannot JIT FlashInfer (no nvcc/ninja, or a
# consumer Blackwell sm_120 card), also:
# export VLLM_USE_FLASHINFER_SAMPLER=0
set -e
MODEL="${MODEL:-nineninesix/gepard-1.0}"
PORT="${PORT:-8091}"
echo "Starting Gepard-1.0 server with model: $MODEL"
vllm-omni serve "$MODEL" \
--host 0.0.0.0 \
--port "$PORT" \
--trust-remote-code \
--omni \
--stage-init-timeout 900 \
--deploy-config vllm_omni/deploy/gepard.yaml
gepard/speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
"""Client for Gepard-1.0 TTS via /v1/audio/speech.
Gepard is zero-shot: omit ``voice`` or pass ``"default"``. Output is 22.05 kHz
mono. ``seed`` is optional; the packaged deploy YAML currently pins ``seed: 42``
until that default is removed, so requests without an explicit seed are still
deterministic.
Examples:
python speech_client.py --text "Hello, this is Gepard speaking."
python speech_client.py --text "Hello, this is Gepard speaking." --seed 7
python speech_client.py --text "Hello, this is Gepard speaking." --stream --output output.pcm
"""
from __future__ import annotations
import argparse
import httpx
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
DEFAULT_MODEL = "nineninesix/gepard-1.0"
def run_tts(args) -> None:
payload = {
"model": args.model,
"input": args.text,
"voice": args.voice,
"response_format": args.response_format,
}
if args.seed is not None:
payload["seed"] = args.seed
if args.max_new_tokens is not None:
payload["max_new_tokens"] = args.max_new_tokens
if args.stream:
payload["stream"] = True
payload["stream_format"] = "audio"
if args.response_format not in ("pcm", "wav"):
print(f"Note: streaming requires pcm/wav; overriding --response-format {args.response_format} to 'pcm'")
payload["response_format"] = "pcm"
print(f"Model: {args.model}")
print(f"Text: {args.text}")
print(f"Voice: {args.voice}")
if args.seed is not None:
print(f"Seed: {args.seed}")
print("Generating audio...")
api_url = f"{args.api_base}/v1/audio/speech"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
if args.stream:
output_path = args.output or ("output.wav" if payload["response_format"] == "wav" else "output.pcm")
with httpx.Client(timeout=300.0) as client:
with client.stream("POST", api_url, json=payload, headers=headers) as resp:
if resp.status_code != 200:
print(f"Error: {resp.status_code}")
print(resp.read().decode())
return
total_bytes = 0
with open(output_path, "wb") as f:
for chunk in resp.iter_bytes():
f.write(chunk)
total_bytes += len(chunk)
print(f"Streamed {total_bytes} bytes to: {output_path}")
return
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return
output_path = args.output or "output.wav"
with open(output_path, "wb") as f:
f.write(response.content)
print(f"Audio saved to: {output_path}")
def main() -> None:
parser = argparse.ArgumentParser(description="Gepard-1.0 TTS client")
parser.add_argument("--api-base", default=DEFAULT_API_BASE, help="API base URL")
parser.add_argument("--api-key", default=DEFAULT_API_KEY, help="API key")
parser.add_argument("--model", "-m", default=DEFAULT_MODEL, help="Model name")
parser.add_argument("--text", required=True, help="Text to synthesize")
parser.add_argument("--voice", default="default", help="Voice name (only 'default' is supported)")
parser.add_argument("--seed", type=int, default=None, help="Sampling seed")
parser.add_argument(
"--max-new-tokens",
type=int,
default=None,
help="Frame budget (1 token = 1 frame ≈ 46.4 ms at 21.5 fps)",
)
parser.add_argument("--stream", action="store_true", help="Enable streaming (PCM output)")
parser.add_argument(
"--response-format",
default="wav",
choices=["wav", "mp3", "flac", "pcm"],
help="Audio format (default: wav). Streaming is pcm/wav only. Opus is not supported at 22.05 kHz.",
)
parser.add_argument("--output", "-o", default=None, help="Output file path")
args = parser.parse_args()
run_tts(args)
if __name__ == "__main__":
main()
glm_tts/gradio_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/glm_tts/gradio_demo.py.
glm_tts/openai_speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""OpenAI-compatible client for GLM-TTS via /v1/audio/speech endpoint.
GLM-TTS is a two-stage TTS system (AR + DiT) that generates audio from text
conditioned on reference speech. Each request requires ref_audio + ref_text.
Usage:
# Voice cloning
python openai_speech_client.py --text "你好" --ref-audio file:///path/to/ref.wav --ref-text "参考文本"
# Streaming response, for async_chunk server mode
python openai_speech_client.py --text "你好" --stream --ref-audio file:///path/to/ref.wav --ref-text "参考文本"
# Specify output format
python openai_speech_client.py --text "你好" --ref-audio file:///path/to/ref.wav \
--ref-text "参考文本" --response-format mp3 -o output.mp3
"""
import argparse
import httpx
# Default server configuration
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
def run_tts_generation(args) -> None:
"""Run TTS generation via OpenAI-compatible /v1/audio/speech API."""
if not args.ref_audio or not args.ref_text:
raise ValueError("GLM-TTS requires --ref-audio and --ref-text for voice cloning.")
payload = {
"model": args.model,
"voice": "default",
"input": args.text,
"response_format": args.response_format,
"stream": bool(args.stream),
"ref_audio": args.ref_audio,
"ref_text": args.ref_text,
}
if args.stream:
payload["stream_format"] = "audio"
payload["response_format"] = "pcm"
if args.max_new_tokens:
payload["max_new_tokens"] = args.max_new_tokens
print(f"Model: {args.model}")
print(f"Text: {args.text}")
print(f"Voice cloning: ref_audio={args.ref_audio}, ref_text={args.ref_text}")
print(f"Stream: {args.stream}")
print("Generating audio...")
api_url = f"{args.api_base}/v1/audio/speech"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
if args.stream:
output_path = args.output or "tts_output.pcm"
with httpx.Client(timeout=300.0) as client, open(output_path, "wb") as f:
with client.stream("POST", api_url, json=payload, headers=headers) as response:
if response.status_code != 200:
print(f"Error: {response.status_code}")
response.read()
print(response.text)
return
for chunk in response.iter_bytes():
f.write(chunk)
print(f"Streaming audio saved to: {output_path}")
else:
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return
try:
text = response.content.decode("utf-8")
if text.startswith('{"error"'):
print(f"Error: {text}")
return
except UnicodeDecodeError:
pass
output_path = args.output or f"tts_output.{args.response_format}"
with open(output_path, "wb") as f:
f.write(response.content)
print(f"Audio saved to: {output_path}")
def parse_args():
"""Parse command line arguments."""
parser = argparse.ArgumentParser(
description="OpenAI-compatible client for GLM-TTS via /v1/audio/speech",
)
# Server configuration
parser.add_argument(
"--api-base",
type=str,
default=DEFAULT_API_BASE,
help=f"API base URL (default: {DEFAULT_API_BASE})",
)
parser.add_argument(
"--api-key",
type=str,
default=DEFAULT_API_KEY,
help="API key (default: EMPTY)",
)
parser.add_argument(
"--model",
"-m",
type=str,
default="glm-tts",
help="Model name/path",
)
# Input text
parser.add_argument(
"--text",
type=str,
required=True,
help="Text to synthesize",
)
# Generation parameters
parser.add_argument(
"--max-new-tokens",
type=int,
default=None,
help="Maximum new tokens to generate (default: model default)",
)
# Output
parser.add_argument(
"--response-format",
type=str,
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio output format (default: wav)",
)
parser.add_argument(
"--stream",
action="store_true",
help="Request a streaming audio response (use with async_chunk server mode).",
)
parser.add_argument(
"--output",
"-o",
type=str,
default=None,
help="Output audio file path (default: tts_output.<format>)",
)
# Voice cloning parameters
parser.add_argument(
"--ref-audio",
type=str,
default=None,
help="Reference audio URL, file:// URI, or base64 data URL for voice cloning",
)
parser.add_argument(
"--ref-text",
type=str,
default=None,
help="Transcript of the reference audio (required with --ref-audio)",
)
return parser.parse_args()
if __name__ == "__main__":
args = parse_args()
run_tts_generation(args)
glm_tts/run_gradio_demo.sh
#!/bin/bash
# Launch GLM-TTS server + Gradio demo together.
#
# Usage:
# ./run_gradio_demo.sh
# CUDA_VISIBLE_DEVICES=0 PORT=8091 GRADIO_PORT=7860 ./run_gradio_demo.sh
set -e
MODEL="${MODEL:-zai-org/GLM-TTS}"
PORT="${PORT:-8091}"
GRADIO_PORT="${GRADIO_PORT:-7860}"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/../../../.." && pwd)"
echo "Starting GLM-TTS server (port $PORT)..."
FLASHINFER_DISABLE_VERSION_CHECK=1 \
vllm-omni serve "$MODEL" \
--deploy-config "$REPO_ROOT/vllm_omni/deploy/glm_tts.yaml" \
--host 0.0.0.0 \
--port "$PORT" \
--gpu-memory-utilization 0.9 \
--trust-remote-code \
--enforce-eager \
--omni &
SERVER_PID=$!
cleanup() {
echo "Stopping server (PID $SERVER_PID)..."
kill $SERVER_PID 2>/dev/null
wait $SERVER_PID 2>/dev/null
}
trap cleanup EXIT
# Wait for server to be ready.
echo "Waiting for server to start..."
for i in $(seq 1 120); do
if curl -s "http://localhost:$PORT/health" > /dev/null 2>&1; then
echo "Server ready."
break
fi
sleep 2
done
echo "Starting Gradio demo (port $GRADIO_PORT)..."
python "$SCRIPT_DIR/gradio_demo.py" \
--api-base "http://localhost:$PORT" \
--port "$GRADIO_PORT"
glm_tts/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for GLM-TTS models
#
# Usage:
# ./run_server.sh # Default model path, async_chunk mode
# ./run_server.sh /path/to/GLM-TTS # Custom model path, async_chunk mode
# ./run_server.sh /path/to/GLM-TTS sync # Sync two-stage mode
#
# NOTE: The model path should point to the repo ROOT (not llm/ subdirectory).
# model_subdir/tokenizer_subdir in the pipeline config resolve subdirectories.
set -e
MODEL="${1:-zai-org/GLM-TTS}"
MODE="${2:-async}"
EXTRA_ARGS=()
case "$MODE" in
async|async_chunk)
;;
sync|no_async_chunk)
EXTRA_ARGS+=("--no-async-chunk")
;;
*)
echo "Unknown mode: $MODE (expected async or sync)" >&2
exit 1
;;
esac
echo "Starting GLM-TTS server with model: $MODEL (mode: $MODE)"
vllm-omni serve "$MODEL" \
--deploy-config vllm_omni/deploy/glm_tts.yaml \
--host 0.0.0.0 \
--port 8091 \
--trust-remote-code \
--omni \
"${EXTRA_ARGS[@]}"
higgs_audio_v2/README.md
higgs-audio v2 online example¶
This directory contains the online-serving entry points for boson-ai's higgs-audio v2 as integrated by vllm-omni: a 2-stage TTS pipeline (Llama-3.2-3B talker with DualFFN audio expert + HiggsAudio codec decoder) emitting 24 kHz mono speech.
Prerequisites¶
Voice clone uses HF's HiggsAudioV2TokenizerModel loaded from k2-fsa/OmniVoice/audio_tokenizer/ (the boson-ai standalone tokenizer Hub repo's model.safetensors is the 3B talker LM, not the codec). Only that ~806 MB subdir is downloaded.
Files¶
run_server.sh— launch the vllm-omni server with the bundledvllm_omni/deploy/higgs_audio_v2.yamldeploy config.batch_speech_client.py— send a list of prompts to/v1/audio/speechand save the returned WAV / PCM bytes to a directory; optionally passes--ref-audio+--ref-textfor shallow voice clone.
Launching the server¶
Environment overrides:
MODEL— HF id of the talker (defaultbosonai/higgs-audio-v2-generation-3B-base).PORT— server port (default8094).GPUS—CUDA_VISIBLE_DEVICESvalue (default6,7).GPU_UTIL—--gpu-memory-utilization(default0.4).
The script also exports VLLM_USE_DEEP_GEMM=0 / VLLM_MOE_USE_DEEP_GEMM=0 so the example works on images without the optional deep_gemm backend.
The deploy YAML ships with async_chunk: false and codec_streaming: true, i.e. Stage 0 finishes its codec frames before Stage 1 starts decoding, and Stage 1 streams WAV/PCM bytes to the client chunk-by-chunk.
Driving the server¶
Plain TTS:
python examples/online_serving/text_to_speech/higgs_audio_v2/batch_speech_client.py \
--base-url http://localhost:8094 \
--model bosonai/higgs-audio-v2-generation-3B-base \
--output-dir /tmp/higgs_audio_v2_batch \
--prompts "Hello world." \
"The quick brown fox jumps over the lazy dog."
Voice clone — pass a reference clip and its transcript (both required together):
python examples/online_serving/text_to_speech/higgs_audio_v2/batch_speech_client.py \
--base-url http://localhost:8094 \
--model bosonai/higgs-audio-v2-generation-3B-base \
--output-dir /tmp/higgs_audio_v2_clone \
--ref-audio /path/to/reference.wav \
--ref-text "Exact transcript spoken in reference.wav." \
--prompts "Hello, this is a cloned voice."
Notes¶
--ref-textmust be the real transcript of--ref-audio; mismatched text degrades cloned-voice quality.- Out of scope (rejected with explicit 4xx by the request validator): multi-speaker
[SPEAKERn]tags insideinput,profile:text-only speaker descriptions, theref_audio_in_system_messagesystem-block variant, chunked long-form generation, and per-requestvoice/instructions/task_type/language/speed != 1.0/x_vector_only_mode/speaker_embedding.
higgs_audio_v2/batch_speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""Batch client for the higgs-audio v2 online server.
Sends a fixed list of prompts to ``/v1/audio/speech`` and saves the returned
WAV files (or raw PCM bytes when ``--format pcm``) into ``--output-dir``.
Usage (plain text -> speech):
python examples/online_serving/text_to_speech/higgs_audio_v2/batch_speech_client.py \
--base-url http://localhost:8094 \
--output-dir /tmp/higgs_audio_v2_batch \
--prompts "Hello world." "The quick brown fox jumps over the lazy dog."
Usage (shallow voice clone — pass a reference clip + its transcript):
python examples/online_serving/text_to_speech/higgs_audio_v2/batch_speech_client.py \
--base-url http://localhost:8094 \
--output-dir /tmp/higgs_audio_v2_clone \
--ref-audio path/to/reference.wav \
--ref-text "the transcript of the reference clip" \
--prompts "Hello world."
"""
from __future__ import annotations
import argparse
import base64
import sys
from pathlib import Path
DEFAULT_PROMPTS = (
"Hello world.",
"The quick brown fox jumps over the lazy dog.",
"It was the night before my birthday.",
"Innovation distinguishes between a leader and a follower.",
)
def _slug(text: str) -> str:
import re
s = re.sub(r"\s+", "_", text.strip().lower())
return re.sub(r"[^a-z0-9_]+", "", s)[:32] or "prompt"
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
parser.add_argument("--base-url", default="http://localhost:8094")
parser.add_argument("--model", default="higgs_audio_v2")
parser.add_argument("--prompts", nargs="+", default=list(DEFAULT_PROMPTS))
parser.add_argument("--output-dir", type=Path, default=Path("/tmp/higgs_audio_v2_batch"))
parser.add_argument("--format", choices=("wav", "pcm"), default="wav")
parser.add_argument("--max-new-tokens", type=int, default=300)
parser.add_argument("--seed", type=int, default=42)
parser.add_argument("--timeout-s", type=float, default=120.0)
parser.add_argument(
"--ref-audio",
type=Path,
default=None,
help="Reference clip for voice clone (path to a WAV file). Must be paired with --ref-text.",
)
parser.add_argument(
"--ref-text",
type=str,
default=None,
help="Transcript of the reference clip. Required when --ref-audio is set.",
)
args = parser.parse_args()
if (args.ref_audio is None) != (args.ref_text is None):
print("--ref-audio and --ref-text must be supplied together", file=sys.stderr)
return 2
ref_audio_data_url: str | None = None
if args.ref_audio is not None:
if not args.ref_audio.exists():
print(f"ref-audio file not found: {args.ref_audio}", file=sys.stderr)
return 2
mime = "audio/wav" if args.ref_audio.suffix.lower() == ".wav" else "audio/mpeg"
ref_b64 = base64.b64encode(args.ref_audio.read_bytes()).decode("ascii")
ref_audio_data_url = f"data:{mime};base64,{ref_b64}"
try:
import httpx
except ImportError:
print(
"this client needs `httpx`. Install with `pip install httpx`.",
file=sys.stderr,
)
return 2
args.output_dir.mkdir(parents=True, exist_ok=True)
url = args.base_url.rstrip("/") + "/v1/audio/speech"
failures = 0
with httpx.Client(timeout=args.timeout_s) as client:
for prompt in args.prompts:
payload = {
"model": args.model,
"input": prompt,
"response_format": args.format,
"max_new_tokens": args.max_new_tokens,
"seed": args.seed,
}
if ref_audio_data_url is not None:
payload["ref_audio"] = ref_audio_data_url
payload["ref_text"] = args.ref_text
resp = client.post(url, json=payload)
if resp.status_code != 200:
print(f"[FAIL] {prompt!r} -> {resp.status_code}: {resp.text[:200]}", file=sys.stderr)
failures += 1
continue
suffix = ".wav" if args.format == "wav" else ".pcm"
out = args.output_dir / f"{_slug(prompt)}{suffix}"
out.write_bytes(resp.content)
print(f"[ ok ] {prompt!r} -> {out} ({len(resp.content)} bytes)")
return 1 if failures else 0
if __name__ == "__main__":
sys.exit(main())
higgs_audio_v2/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for higgs-audio v2.
#
# v1 scope: plain text -> 24 kHz speech only. Voice cloning, multi-speaker,
# ChatML rich content, and language overrides are rejected by the validator
# with explicit 4xx (see vllm_omni/entrypoints/openai/serving_speech.py).
#
# Usage:
# ./run_server.sh # default port 8094, GPUs 6 and 7
# PORT=8095 GPUS=6,7 ./run_server.sh
# MODEL=bosonai/higgs-audio-v2-generation-3B-base ./run_server.sh
set -e
MODEL="${MODEL:-bosonai/higgs-audio-v2-generation-3B-base}"
PORT="${PORT:-8094}"
GPUS="${GPUS:-6,7}"
GPU_UTIL="${GPU_UTIL:-0.4}"
echo "Starting higgs-audio v2 server"
echo " MODEL=$MODEL"
echo " PORT=$PORT"
echo " CUDA_VISIBLE_DEVICES=$GPUS"
# DeepGEMM FP8 kernels are optional and trip warmup on builds without
# the deep_gemm backend; disable them so the example works out of the box.
# Users with deep_gemm installed can re-enable via the same env vars.
CUDA_VISIBLE_DEVICES="$GPUS" \
VLLM_USE_DEEP_GEMM=0 \
VLLM_MOE_USE_DEEP_GEMM=0 \
vllm-omni serve "$MODEL" \
--deploy-config vllm_omni/deploy/higgs_audio_v2.yaml \
--host 0.0.0.0 \
--port "$PORT" \
--gpu-memory-utilization "$GPU_UTIL" \
--trust-remote-code \
--omni
higgs_audio_v3/README.md
Higgs-Audio V3 Online Serving¶
Start the server¶
# Default: GPU 0, port 8095
./examples/online_serving/text_to_speech/higgs_audio_v3/run_server.sh
# Custom GPU / port
PORT=8096 GPUS=0,1 ./examples/online_serving/text_to_speech/higgs_audio_v3/run_server.sh
Plain text TTS¶
python examples/online_serving/text_to_speech/higgs_audio_v3/batch_speech_client.py \
--base-url http://localhost:8095 \
--output-dir /tmp/higgs_v3_batch \
--prompts "Hello world." "The quick brown fox jumps over the lazy dog."
Voice clone¶
python examples/online_serving/text_to_speech/higgs_audio_v3/batch_speech_client.py \
--base-url http://localhost:8095 \
--output-dir /tmp/higgs_v3_clone \
--ref-audio path/to/reference.wav \
--ref-text "transcript of the reference clip" \
--prompts "Text to synthesize in the cloned voice."
curl example¶
higgs_audio_v3/batch_speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""Batch client for the higgs-audio v3 online server.
Sends prompts to ``/v1/audio/speech`` and saves the returned WAV files.
Usage (plain text -> speech):
python examples/online_serving/text_to_speech/higgs_audio_v3/batch_speech_client.py \
--base-url http://localhost:8095 \
--output-dir /tmp/higgs_v3_batch \
--prompts "Hello world." "The quick brown fox jumps over the lazy dog."
Usage (voice clone):
python examples/online_serving/text_to_speech/higgs_audio_v3/batch_speech_client.py \
--base-url http://localhost:8095 \
--output-dir /tmp/higgs_v3_clone \
--ref-audio path/to/reference.wav \
--ref-text "the transcript of the reference clip" \
--prompts "Hello world."
"""
from __future__ import annotations
import argparse
import base64
import sys
from pathlib import Path
DEFAULT_PROMPTS = (
"Hello world.",
"The quick brown fox jumps over the lazy dog.",
"Today is a beautiful day for a walk in the park.",
"Innovation distinguishes between a leader and a follower.",
)
def _slug(text: str) -> str:
import re
s = re.sub(r"\s+", "_", text.strip().lower())
return re.sub(r"[^a-z0-9_]+", "", s)[:32] or "prompt"
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
parser.add_argument("--base-url", default="http://localhost:8095")
parser.add_argument("--model", default="higgs_audio_v3")
parser.add_argument("--prompts", nargs="+", default=list(DEFAULT_PROMPTS))
parser.add_argument("--output-dir", type=Path, default=Path("/tmp/higgs_v3_batch"))
parser.add_argument("--format", choices=("wav", "pcm"), default="wav")
parser.add_argument("--max-new-tokens", type=int, default=2048)
parser.add_argument("--seed", type=int, default=42)
parser.add_argument("--timeout-s", type=float, default=120.0)
parser.add_argument(
"--ref-audio",
type=Path,
default=None,
help="Reference clip for voice clone (WAV/FLAC/MP3 path). Pair with --ref-text.",
)
parser.add_argument(
"--ref-text",
type=str,
default=None,
help="Transcript of the reference clip. Optional but improves fidelity.",
)
args = parser.parse_args()
ref_audio_data_url: str | None = None
if args.ref_audio is not None:
if not args.ref_audio.exists():
print(f"ref-audio file not found: {args.ref_audio}", file=sys.stderr)
return 2
mime = "audio/wav" if args.ref_audio.suffix.lower() == ".wav" else "audio/mpeg"
ref_b64 = base64.b64encode(args.ref_audio.read_bytes()).decode("ascii")
ref_audio_data_url = f"data:{mime};base64,{ref_b64}"
try:
import httpx
except ImportError:
print("this client needs `httpx`. Install with `pip install httpx`.", file=sys.stderr)
return 2
args.output_dir.mkdir(parents=True, exist_ok=True)
url = args.base_url.rstrip("/") + "/v1/audio/speech"
failures = 0
with httpx.Client(timeout=args.timeout_s) as client:
for prompt in args.prompts:
payload = {
"model": args.model,
"input": prompt,
"response_format": args.format,
"max_new_tokens": args.max_new_tokens,
"seed": args.seed,
}
if ref_audio_data_url is not None:
payload["ref_audio"] = ref_audio_data_url
if args.ref_text:
payload["ref_text"] = args.ref_text
resp = client.post(url, json=payload)
if resp.status_code != 200:
print(f"[FAIL] {prompt!r} -> {resp.status_code}: {resp.text[:200]}", file=sys.stderr)
failures += 1
continue
suffix = ".wav" if args.format == "wav" else ".pcm"
out = args.output_dir / f"{_slug(prompt)}{suffix}"
out.write_bytes(resp.content)
print(f"[ ok ] {prompt!r} -> {out} ({len(resp.content)} bytes)")
return 1 if failures else 0
if __name__ == "__main__":
sys.exit(main())
higgs_audio_v3/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for higgs-audio v3.
#
# Supports plain text TTS and voice cloning via /v1/audio/speech.
#
# Usage:
# ./run_server.sh # default port 8095, GPU 0
# PORT=8096 GPUS=0,1 ./run_server.sh
# MODEL=/path/to/local/checkpoint ./run_server.sh
set -e
MODEL="${MODEL:-bosonai/higgs-audio-v3-tts-4b}"
PORT="${PORT:-8095}"
GPUS="${GPUS:-0}"
GPU_UTIL="${GPU_UTIL:-0.6}"
echo "Starting higgs-audio v3 server"
echo " MODEL=$MODEL"
echo " PORT=$PORT"
echo " CUDA_VISIBLE_DEVICES=$GPUS"
CUDA_VISIBLE_DEVICES="$GPUS" \
VLLM_USE_DEEP_GEMM=0 \
VLLM_MOE_USE_DEEP_GEMM=0 \
vllm-omni serve "$MODEL" \
--deploy-config vllm_omni/deploy/higgs_multimodal_qwen3.yaml \
--host 0.0.0.0 \
--port "$PORT" \
--gpu-memory-utilization "$GPU_UTIL" \
--trust-remote-code \
--omni
indextts2/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for IndexTTS 2.0 or 2.5.
#
# Usage from repository root:
# examples/online_serving/text_to_speech/indextts2/run_server.sh
# MODEL_VERSION=2.5 MODEL=/path/to/native/bundle \
# examples/online_serving/text_to_speech/indextts2/run_server.sh
set -e
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
ROOT_DIR="$(cd -- "$SCRIPT_DIR/../../../.." && pwd)"
MODEL_VERSION="${MODEL_VERSION:-2.0}"
PORT="${PORT:-8092}"
if [[ "$MODEL_VERSION" == "2.5" ]]; then
if [[ -z "${MODEL:-}" ]]; then
echo "Usage: MODEL_VERSION=2.5 MODEL=/path/to/native/bundle bash $0" >&2
exit 2
fi
DEFAULT_DEPLOY_CONFIG="$ROOT_DIR/vllm_omni/deploy/indextts2_5.yaml"
elif [[ "$MODEL_VERSION" == "2.0" ]]; then
MODEL="${MODEL:-IndexTeam/IndexTTS-2}"
DEFAULT_DEPLOY_CONFIG="$ROOT_DIR/vllm_omni/deploy/indextts2.yaml"
else
echo "MODEL_VERSION must be 2.0 or 2.5" >&2
exit 2
fi
DEPLOY_CONFIG="${DEPLOY_CONFIG:-$DEFAULT_DEPLOY_CONFIG}"
echo "Starting IndexTTS $MODEL_VERSION server with model: $MODEL"
echo "Deploy config: $DEPLOY_CONFIG"
FLASHINFER_DISABLE_VERSION_CHECK=1 \
vllm serve "$MODEL" \
--host 0.0.0.0 \
--port "$PORT" \
--omni \
--trust-remote-code \
--deploy-config "$DEPLOY_CONFIG"
indextts2/speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""OpenAI-compatible client for IndexTTS 2.0/2.5 via /v1/audio/speech.
Examples:
# With reference audio for voice cloning
python speech_client.py --text "你好,世界!" \
--ref-audio /path/to/reference.wav
# With emotion audio
python speech_client.py --text "今天心情很好!" \
--ref-audio /path/to/ref.wav \
--emo-audio /path/to/happy.wav
Server setup:
vllm serve IndexTeam/IndexTTS-2 --omni --host 0.0.0.0 --port 8092
"""
from __future__ import annotations
import argparse
import base64
import os
import httpx
DEFAULT_API_BASE = "http://localhost:8092"
DEFAULT_API_KEY = "sk-empty"
def encode_audio_to_base64(audio_path: str) -> str:
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
ext = audio_path.lower().rsplit(".", 1)[-1]
mime = {"wav": "audio/wav", "mp3": "audio/mpeg", "flac": "audio/flac"}.get(ext, "audio/wav")
with open(audio_path, "rb") as f:
b64 = base64.b64encode(f.read()).decode("utf-8")
return f"data:{mime};base64,{b64}"
def main() -> None:
parser = argparse.ArgumentParser(description="IndexTTS2 OpenAI speech client")
parser.add_argument("--text", type=str, required=True)
parser.add_argument("--ref-audio", type=str, default=None, help="Reference audio for voice cloning")
parser.add_argument("--emo-audio", type=str, default=None, help="Emotion reference audio")
parser.add_argument("--emo-text", type=str, default=None, help="Emotion description text")
parser.add_argument(
"--emo-vector",
type=float,
nargs=8,
default=None,
help="8-dim emotion vector: happy angry sad afraid disgusted melancholic surprised calm",
)
parser.add_argument("--emo-alpha", type=float, default=None, help="Emotion weight in [0, 1]")
parser.add_argument("--use-emo-text", action="store_true", help="Infer emotion vector from emo-text or text")
parser.add_argument("--use-random", action="store_true", help="Use random emotion prototypes")
parser.add_argument(
"--model-version",
choices=("2.0", "2.5"),
default="2.0",
)
parser.add_argument("--model", type=str, default=None)
parser.add_argument(
"--lang",
default="zh",
help=(
"IndexTTS 2.5 language code, for example zh/en/zhen/ja/yue; "
"zhen is mixed Chinese/English and Mandarin is a vLLM-Omni alias for zh"
),
)
parser.add_argument(
"--no-text-normalization",
action="store_false",
dest="text_normalization",
help="Disable IndexTTS 2.5 text normalization",
)
parser.add_argument(
"--speed",
type=float,
default=1.0,
help="IndexTTS 2.5 native synthesis speed in [0.5, 2.0]; 2.0 is faster",
)
parser.add_argument("--voice", type=str, default=None, help="Uploaded voice name to use instead of --ref-audio")
parser.add_argument("--output", type=str, default="output.wav")
parser.add_argument("--api-base", type=str, default=DEFAULT_API_BASE)
parser.add_argument("--api-key", type=str, default=DEFAULT_API_KEY)
parser.add_argument("--response-format", type=str, default="wav")
args = parser.parse_args()
if args.model is None:
if args.model_version == "2.5":
parser.error("--model is required for IndexTTS 2.5")
args.model = "IndexTeam/IndexTTS-2"
if args.model_version == "2.5" and not 0.5 <= args.speed <= 2.0:
parser.error("IndexTTS 2.5 --speed must be between 0.5 and 2.0")
if not args.ref_audio and not args.voice:
parser.error("IndexTTS2 requires --ref-audio or --voice for voice cloning")
payload: dict = {
"model": args.model,
"input": args.text,
"response_format": args.response_format,
}
if args.voice:
payload["voice"] = args.voice
if args.ref_audio:
ref = args.ref_audio
if ref.startswith(("http://", "https://", "data:")):
payload["ref_audio"] = ref
else:
payload["ref_audio"] = encode_audio_to_base64(ref)
extra_params = {}
if args.emo_audio:
emo = args.emo_audio
if emo.startswith(("http://", "https://", "data:")):
extra_params["emo_audio"] = emo
else:
extra_params["emo_audio"] = encode_audio_to_base64(emo)
if args.emo_text:
extra_params["emo_text"] = args.emo_text
if args.emo_vector is not None:
extra_params["emo_vector"] = args.emo_vector
if args.emo_alpha is not None:
extra_params["emo_alpha"] = args.emo_alpha
if args.use_emo_text:
extra_params["use_emo_text"] = True
if args.use_random:
extra_params["use_random"] = True
if args.model_version == "2.5":
payload["speed"] = args.speed
extra_params["lang"] = args.lang
extra_params["text_normalization"] = args.text_normalization
if extra_params:
payload["extra_params"] = extra_params
url = f"{args.api_base}/v1/audio/speech"
print(f"POST {url}")
print(f" text: {args.text}")
if args.ref_audio:
print(f" ref_audio: {args.ref_audio[:80]}...")
with httpx.Client(timeout=300) as client:
resp = client.post(
url,
json=payload,
headers={"Authorization": f"Bearer {args.api_key}"},
)
if resp.status_code != 200:
print(f"Error {resp.status_code}: {resp.text[:500]}")
return
with open(args.output, "wb") as f:
f.write(resp.content)
print(f"Saved: {args.output} ({len(resp.content):,} bytes)")
if __name__ == "__main__":
main()
ming_flash_omni_tts/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for Ming-flash-omni-2.0 standalone talker (TTS).
#
# Usage:
# ./run_server.sh
# MODEL=/path/to/local/model ./run_server.sh
# PORT=8091 ./run_server.sh
# HOST=127.0.0.1 ./run_server.sh # bind only to loopback
set -e
MODEL="${MODEL:-Jonathan1909/Ming-flash-omni-2.0}"
HOST="${HOST:-0.0.0.0}"
PORT="${PORT:-8091}"
DEPLOY_CONFIG="${DEPLOY_CONFIG:-vllm_omni/deploy/ming_flash_omni_tts.yaml}"
echo "Starting Ming standalone TTS server with model: $MODEL"
echo "Deploy config: $DEPLOY_CONFIG"
vllm serve "$MODEL" \
--deploy-config "$DEPLOY_CONFIG" \
--host "$HOST" \
--port "$PORT" \
--trust-remote-code \
--omni
ming_flash_omni_tts/speech_client.py
"""Client for Ming standalone TTS via /v1/audio/speech endpoint."""
import argparse
import json
import sys
import httpx
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
DEFAULT_MODEL = "Jonathan1909/Ming-flash-omni-2.0"
def run_tts(args) -> None:
payload = {
"model": args.model,
"input": args.text,
"response_format": args.response_format,
}
instructions = args.instructions
if args.instruction_json:
if instructions:
sys.exit("--instructions and --instruction-json are mutually exclusive")
try:
parsed = json.loads(args.instruction_json)
except json.JSONDecodeError as exc:
sys.exit(f"--instruction-json must be valid JSON: {exc}")
if not isinstance(parsed, dict):
sys.exit("--instruction-json must decode to a JSON object")
# Re-encode with ensure_ascii=False so UTF-8 Chinese keys/values
# arrive at the server intact rather than as \\uXXXX escapes.
instructions = json.dumps(parsed, ensure_ascii=False)
if instructions:
payload["instructions"] = instructions
print(f"Model: {args.model}")
print(f"Text: {args.text}")
print("Generating audio...")
api_url = f"{args.api_base}/v1/audio/speech"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return
output_path = args.output or "ming_tts_output.wav"
with open(output_path, "wb") as f:
f.write(response.content)
print(f"Audio saved to: {output_path}")
def main():
parser = argparse.ArgumentParser(description="Ming standalone TTS speech client")
parser.add_argument("--api-base", default=DEFAULT_API_BASE, help="API base URL")
parser.add_argument("--api-key", default=DEFAULT_API_KEY, help="API key")
parser.add_argument("--model", "-m", default=DEFAULT_MODEL, help="Model name or local path")
parser.add_argument("--text", required=True, help="Text to synthesize")
parser.add_argument(
"--response-format",
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio format (default: wav)",
)
parser.add_argument("--output", "-o", default=None, help="Output file path")
parser.add_argument(
"--instructions",
default=None,
help="Free-form style description (mapped to caption 风格 on the server).",
)
parser.add_argument(
"--instruction-json",
default=None,
help=(
"Structured caption JSON forwarded as `instructions`. Accepts Ming "
"caption keys: 方言, 风格, 语速, 基频, 音量, 情感, IP, 说话人, BGM. "
),
)
args = parser.parse_args()
run_tts(args)
if __name__ == "__main__":
main()
ming_tts/README.md
Ming-omni-tts Online Serving¶
Serve the dense inclusionAI/Ming-omni-tts-0.5B two-stage TTS model through the OpenAI-compatible /v1/audio/speech endpoint.
Start Server¶
vllm-omni serve inclusionAI/Ming-omni-tts-0.5B \
--deploy-config vllm_omni/deploy/ming_tts.yaml \
--omni \
--port 8091 \
--enforce-eager
Or:
The tested ROCm environment is summarized in the Ming recipe.
Send Requests¶
The Python client targets http://localhost:8091/v1 with api_key=EMPTY; it does not call OpenAI's hosted API.
Style or dialect controls can be plain text or Ming JSON. The upstream dialect example also uses yue_prompt.wav for speaker conditioning:
python openai_speech_client.py \
--text "我觉得社会企业同个人都有责任" \
--instruction-json '{"方言":"广粤话"}' \
--ref-audio /path/to/yue_prompt.wav \
--max-new-tokens 200
When --ref-audio is supplied without --ref-text, the server extracts the Ming speaker embedding, matching upstream use_spk_emb=True, without using the audio as a zero-shot prompt.
Reference-audio cloning:
python openai_speech_client.py \
--text "我们的愿景是构建未来服务业的数字化基础设施,为世界带来更多微小而美好的改变。" \
--ref-audio /path/to/10002287-00000094.wav \
--ref-text "在此奉劝大家别乱打美白针。" \
--max-new-tokens 200
Podcast-style multi-speaker prompt:
python openai_speech_client.py \
--text " speaker_1:你可以说一下,就大概说一下,可能虽然我也不知道,我看过那部电影没有。
speaker_2:就是那个叫什么,变相一节课的嘛。
speaker_1:嗯。
speaker_2:一部搞笑的电影。
speaker_1:一部搞笑的。" \
--ref-audio /path/to/CTS-CN-F2F-2019-11-11-423-012-A.wav \
--ref-audio /path/to/CTS-CN-F2F-2019-11-11-423-012-B.wav \
--ref-text " speaker_1:并且我们还要进行每个月还要考核 笔试的话还要进行笔试,做个,当服务员还要去笔试了
speaker_2:对啊,这真的很奇怪,就是 单纯的因,单纯自己工资不高,只是因为可能人家那个店比较出名一点,就对你苛刻要求"
Streaming PCM:
run_curl.sh keeps small smoke checks:
./run_curl.sh basic
REF_AUDIO=/path/to/reference.wav REF_TEXT="在此奉劝大家别乱打美白针。" ./run_curl.sh zero_shot
./run_curl.sh stream
Request Fields¶
| Field | Ming meaning |
|---|---|
input | target text |
instructions | plain style text, or JSON object for structured Ming controls |
voice | Ming IP voice label unless it resolves to an uploaded speaker |
language | Ming 方言 control |
ref_audio | speaker reference; with ref_text, also supplies the prompt waveform |
ref_text | transcript enabling zero-shot or podcast prompt-latent conditioning |
speaker_embedding | 192-d Ming speaker embedding |
max_new_tokens | Ming max_decode_steps |
Notes¶
ref_audioaccepts local paths through the client, remote URLs,file://, ordata:URLs.- Non-streaming responses return WAV bytes; streaming responses return PCM.
- Music-only
bgmgeneration is offline-only until the API exposes Ming prompt-mode selection.
ming_tts/openai_speech_client.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/ming_tts/openai_speech_client.py.
ming_tts/run_curl.sh
#!/bin/bash
set -euo pipefail
MODE="${1:-basic}"
HOST="${HOST:-localhost}"
PORT="${PORT:-8091}"
MODEL="${MODEL:-inclusionAI/Ming-omni-tts-0.5B}"
API_URL="http://${HOST}:${PORT}/v1/audio/speech"
TEXT="${TEXT:-你好,这是 Ming 在线语音合成测试。}"
OUTPUT="${OUTPUT:-ming_output.wav}"
STREAM_OUTPUT="${STREAM_OUTPUT:-ming_output.pcm}"
REF_AUDIO="${REF_AUDIO:-}"
REF_TEXT="${REF_TEXT:-}"
post_json() {
local payload="$1"
local output_path="$2"
curl -X POST "$API_URL" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer EMPTY" \
-d "$payload" \
--output "$output_path"
}
case "$MODE" in
basic)
post_json "{
\"model\": \"${MODEL}\",
\"input\": \"${TEXT}\",
\"response_format\": \"wav\"
}" "$OUTPUT"
;;
zero_shot)
if [ -z "$REF_AUDIO" ] || [ -z "$REF_TEXT" ]; then
echo "zero_shot requires REF_AUDIO and REF_TEXT" >&2
exit 1
fi
python - <<'PY' > /tmp/ming_zero_shot_payload.json
import base64
import json
import mimetypes
import os
from pathlib import Path
path = Path(os.environ["REF_AUDIO"])
mime_type = mimetypes.guess_type(path.name)[0] or "audio/wav"
payload = {
"model": os.environ["MODEL"],
"input": os.environ["TEXT"],
"ref_audio": f"data:{mime_type};base64,{base64.b64encode(path.read_bytes()).decode('utf-8')}",
"ref_text": os.environ["REF_TEXT"],
"response_format": "wav",
}
print(json.dumps(payload, ensure_ascii=False))
PY
curl -X POST "$API_URL" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer EMPTY" \
--data-binary @/tmp/ming_zero_shot_payload.json \
--output "$OUTPUT"
rm -f /tmp/ming_zero_shot_payload.json
;;
stream)
post_json "{
\"model\": \"${MODEL}\",
\"input\": \"${TEXT}\",
\"stream\": true,
\"stream_format\": \"audio\",
\"response_format\": \"pcm\"
}" "$STREAM_OUTPUT"
;;
*)
echo "Unknown mode: $MODE" >&2
echo "Supported sanity checks: basic, zero_shot, stream" >&2
exit 1
;;
esac
ming_tts/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for Ming-omni-tts.
#
# Usage:
# ./run_server.sh
# PORT=8000 ./run_server.sh
set -e
DIR="$(cd "$(dirname "$0")" && pwd)"
ROOT="$(cd "$DIR/../../../.." && pwd)"
MODEL="${MODEL:-inclusionAI/Ming-omni-tts-0.5B}"
PORT="${PORT:-8091}"
DEPLOY_CONFIG="${DEPLOY_CONFIG:-$ROOT/vllm_omni/deploy/ming_tts.yaml}"
echo "Starting Ming-omni-tts server with model: $MODEL"
echo "Deploy config: $DEPLOY_CONFIG"
vllm-omni serve "$MODEL" \
--deploy-config "$DEPLOY_CONFIG" \
--host 0.0.0.0 \
--port "$PORT" \
--enforce-eager \
--omni
moss_tts_nano/gradio_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/moss_tts_nano/gradio_demo.py.
moss_tts_nano/run_gradio_demo.sh
#!/bin/bash
# Launch MOSS-TTS-Nano server + Gradio demo together.
#
# Usage:
# ./run_gradio_demo.sh
# CUDA_VISIBLE_DEVICES=0 PORT=8091 GRADIO_PORT=7860 ./run_gradio_demo.sh
set -e
MODEL="${MODEL:-OpenMOSS-Team/MOSS-TTS-Nano}"
PORT="${PORT:-8091}"
GRADIO_PORT="${GRADIO_PORT:-7860}"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
echo "Starting MOSS-TTS-Nano server (port $PORT)..."
FLASHINFER_DISABLE_VERSION_CHECK=1 \
vllm serve "$MODEL" \
--host 0.0.0.0 \
--port "$PORT" \
--omni &
SERVER_PID=$!
cleanup() {
echo "Stopping server (PID $SERVER_PID)..."
kill $SERVER_PID 2>/dev/null
wait $SERVER_PID 2>/dev/null
}
trap cleanup EXIT
# Wait for server to be ready.
echo "Waiting for server to start..."
for i in $(seq 1 120); do
if curl -s "http://localhost:$PORT/health" > /dev/null 2>&1; then
echo "Server ready."
break
fi
sleep 2
done
echo "Starting Gradio demo (port $GRADIO_PORT)..."
python "$SCRIPT_DIR/gradio_demo.py" \
--api-base "http://localhost:$PORT" \
--port "$GRADIO_PORT"
moss_tts_nano/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for MOSS-TTS-Nano
#
# Usage:
# ./run_server.sh
# CUDA_VISIBLE_DEVICES=0 PORT=8091 ./run_server.sh
set -e
MODEL="${MODEL:-OpenMOSS-Team/MOSS-TTS-Nano}"
PORT="${PORT:-8091}"
echo "Starting MOSS-TTS-Nano server with model: $MODEL"
FLASHINFER_DISABLE_VERSION_CHECK=1 \
vllm serve "$MODEL" \
--host 0.0.0.0 \
--port "$PORT" \
--omni
omnivoice/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for OmniVoice TTS
#
# Usage:
# ./run_server.sh
# CUDA_VISIBLE_DEVICES=0 ./run_server.sh
set -e
MODEL="${MODEL:-k2-fsa/OmniVoice}"
PORT="${PORT:-8091}"
echo "Starting OmniVoice server with model: $MODEL"
vllm serve "$MODEL" \
--host 0.0.0.0 \
--port "$PORT" \
--trust-remote-code \
--omni
omnivoice/speech_client.py
"""Client for OmniVoice TTS via /v1/audio/speech endpoint.
Examples:
# Basic TTS (auto voice)
python speech_client.py --text "Hello, how are you?"
# Specify language
python speech_client.py --text "Bonjour, comment allez-vous?" --language French
# Use a specific uploaded/supported voice
python speech_client.py --text "Hello" --voice my_uploaded_voice
"""
import argparse
import base64
import os
import httpx
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
def encode_audio_to_base64(audio_path: str) -> str:
"""Encode a local audio file to a base64 data URL."""
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
ext = audio_path.lower().rsplit(".", 1)[-1]
mime = {
"wav": "audio/wav",
"mp3": "audio/mpeg",
"flac": "audio/flac",
"ogg": "audio/ogg",
}.get(ext, "audio/wav")
with open(audio_path, "rb") as f:
b64 = base64.b64encode(f.read()).decode("utf-8")
return f"data:{mime};base64,{b64}"
def run_tts(args) -> None:
"""Generate speech via /v1/audio/speech API."""
payload = {
"model": args.model,
"input": args.text,
"response_format": args.response_format,
}
if args.seed is not None:
payload["extra_params"] = {}
payload["extra_params"]["seed"] = args.seed
if args.voice:
payload["voice"] = args.voice
if args.language:
payload["language"] = args.language
if args.ref_audio:
ref = args.ref_audio
if ref.startswith(("http://", "https://", "data:")):
payload["ref_audio"] = ref
else:
payload["ref_audio"] = encode_audio_to_base64(ref)
if args.ref_text:
payload["ref_text"] = args.ref_text
if args.instructions:
payload["instructions"] = args.instructions
print(f"Model: {args.model}")
print(f"Text: {args.text}")
if args.seed:
print(f"Seed: {args.seed}")
if args.voice:
print(f"Voice: {args.voice}")
if args.language:
print(f"Language: {args.language}")
print("Generating audio...")
api_url = f"{args.api_base}/v1/audio/speech"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return
try:
text = response.content.decode("utf-8")
if text.startswith('{"error"'):
print(f"Error: {text}")
return
except UnicodeDecodeError:
pass
output_path = args.output or "omnivoice_output.wav"
with open(output_path, "wb") as f:
f.write(response.content)
print(f"Audio saved to: {output_path}")
def main():
parser = argparse.ArgumentParser(description="OmniVoice TTS client")
parser.add_argument("--api-base", default=DEFAULT_API_BASE, help="API base URL")
parser.add_argument("--api-key", default=DEFAULT_API_KEY, help="API key")
parser.add_argument("--model", "-m", default="k2-fsa/OmniVoice", help="Model name")
parser.add_argument("--text", required=True, help="Text to synthesize")
parser.add_argument(
"--voice",
default=None,
help="Voice name (omit for auto voice; must match a supported or uploaded speaker if set)",
)
parser.add_argument("--language", default=None, help="Language hint (e.g., English, Chinese, French)")
parser.add_argument(
"--response-format",
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio format (default: wav)",
)
parser.add_argument(
"--ref-audio",
type=str,
default=None,
help="Reference audio for voice cloning (local path, URL, or data: URI)",
)
parser.add_argument(
"--ref-text",
type=str,
default=None,
help="Reference text for voice cloning",
)
parser.add_argument(
"--instructions",
type=str,
default=None,
help="Voice style/emotion instructions",
)
parser.add_argument(
"--seed",
type=int,
default=None,
help="Random seed for generation, default: None for stochastic output)",
)
parser.add_argument("--output", "-o", default=None, help="Output file path")
args = parser.parse_args()
run_tts(args)
if __name__ == "__main__":
main()
qwen3_tts/batch_speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
"""Batch speech client for Qwen3-TTS via /v1/audio/speech/batch endpoint.
This script demonstrates how to synthesize multiple texts in a single request.
A particularly useful scenario is voice cloning: set ref_audio once at the
batch level and generate many utterances in the cloned voice without repeating
the reference for each item.
Start the server (with batch-optimized stage settings for best throughput):
vllm serve Qwen/Qwen3-TTS-12Hz-1.7B-Base \
--omni \
--trust-remote-code \
--stage-overrides '{"0":{"max_num_seqs":4,"gpu_memory_utilization":0.2},
"1":{"max_num_seqs":4,"gpu_memory_utilization":0.2}}'
Examples:
# Batch with a predefined voice
python batch_speech_client.py \
--texts "Hello, how are you?" "Goodbye, see you later!"
# Voice cloning: one ref_audio, many outputs
python batch_speech_client.py \
--task-type Base \
--ref-audio /path/to/reference.wav \
--ref-text "Transcript of the reference audio" \
--texts "First cloned sentence." "Second cloned sentence." \
"Third cloned sentence."
"""
import argparse
import base64
import os
import httpx
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
def encode_audio_to_base64(audio_path: str) -> str:
"""Encode a local audio file to a base64 data URL."""
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
ext = os.path.splitext(audio_path)[1].lower()
mime_map = {".wav": "audio/wav", ".mp3": "audio/mpeg", ".flac": "audio/flac", ".ogg": "audio/ogg"}
mime_type = mime_map.get(ext, "audio/wav")
with open(audio_path, "rb") as f:
audio_b64 = base64.b64encode(f.read()).decode("utf-8")
return f"data:{mime_type};base64,{audio_b64}"
def run_batch(args) -> None:
"""Send a batch TTS request and save each result to a file."""
items = [{"input": text} for text in args.texts]
payload: dict = {
"items": items,
"response_format": args.response_format,
}
if args.voice:
payload["voice"] = args.voice
if args.language:
payload["language"] = args.language
if args.task_type:
payload["task_type"] = args.task_type
if args.instructions:
payload["instructions"] = args.instructions
if args.max_new_tokens:
payload["max_new_tokens"] = args.max_new_tokens
# Voice cloning parameters (shared across all items)
if args.ref_audio:
if args.ref_audio.startswith(("http://", "https://")):
payload["ref_audio"] = args.ref_audio
else:
payload["ref_audio"] = encode_audio_to_base64(args.ref_audio)
if args.ref_text:
payload["ref_text"] = args.ref_text
print(f"Sending batch of {len(items)} item(s) to {args.api_base}")
if args.ref_audio:
print("Voice cloning mode — ref_audio applied to all items")
url = f"{args.api_base}/v1/audio/speech/batch"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
with httpx.Client(timeout=300.0) as client:
response = client.post(url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error {response.status_code}: {response.text}")
return
data = response.json()
print(f"Total: {data['total']} Succeeded: {data['succeeded']} Failed: {data['failed']}")
os.makedirs(args.output_dir, exist_ok=True)
for result in data["results"]:
idx = result["index"]
if result["status"] == "success":
audio_bytes = base64.b64decode(result["audio_data"])
out_path = os.path.join(args.output_dir, f"batch_{idx}.{args.response_format}")
with open(out_path, "wb") as f:
f.write(audio_bytes)
print(f" [{idx}] saved {len(audio_bytes)} bytes -> {out_path}")
else:
print(f" [{idx}] FAILED: {result['error']}")
def parse_args():
parser = argparse.ArgumentParser(
description="Batch speech client for /v1/audio/speech/batch",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("--api-base", default=DEFAULT_API_BASE, help="API base URL")
parser.add_argument("--api-key", default=DEFAULT_API_KEY, help="API key")
# Texts to synthesize
parser.add_argument(
"--texts",
nargs="+",
required=True,
help="One or more texts to synthesize",
)
# Shared voice settings
parser.add_argument("--voice", default="vivian", help="Speaker name (default: vivian)")
parser.add_argument("--language", default=None, help="Language: Auto, Chinese, English, etc.")
parser.add_argument("--instructions", default=None, help="Voice style/emotion instructions")
parser.add_argument(
"--task-type",
default=None,
choices=["CustomVoice", "VoiceDesign", "Base"],
help="TTS task type (default: CustomVoice)",
)
# Voice cloning (Base task)
parser.add_argument("--ref-audio", default=None, help="Reference audio path or URL for voice cloning")
parser.add_argument("--ref-text", default=None, help="Reference audio transcript for voice cloning")
# Generation
parser.add_argument("--max-new-tokens", type=int, default=None, help="Max new tokens per item")
parser.add_argument(
"--response-format",
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio format (default: wav)",
)
parser.add_argument("--output-dir", "-o", default="batch_output", help="Output directory (default: batch_output)")
return parser.parse_args()
if __name__ == "__main__":
args = parse_args()
run_batch(args)
qwen3_tts/gradio_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/qwen3_tts/gradio_demo.py.
qwen3_tts/openai_speech_client.py
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
"""OpenAI-compatible client for Qwen3-TTS via /v1/audio/speech endpoint.
This script demonstrates how to use the OpenAI-compatible speech API
to generate audio from text using Qwen3-TTS models.
Examples:
# CustomVoice task (predefined speaker)
python openai_speech_client.py --text "Hello, how are you?" --speaker vivian
# CustomVoice with emotion instruction
python openai_speech_client.py --text "I'm so happy!" --speaker vivian \
--instructions "Speak with excitement"
# VoiceDesign task (voice from description)
python openai_speech_client.py --text "Hello world" \
--model Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign \
--task-type VoiceDesign \
--instructions "A warm, friendly female voice"
# Base task (voice cloning)
python openai_speech_client.py --text "Hello world" \
--model Qwen/Qwen3-TTS-12Hz-1.7B-Base \
--task-type Base \
--ref-audio "https://example.com/reference.wav" \
--ref-text "This is the reference transcript"
# Base task with pre-computed speaker embedding
python openai_speech_client.py --text "Hello world" \
--model Qwen/Qwen3-TTS-12Hz-1.7B-Base \
--task-type Base \
--speaker-embedding embedding.json
"""
import argparse
import base64
import json
import os
import httpx
# Default server configuration
DEFAULT_API_BASE = "http://localhost:8091"
DEFAULT_API_KEY = "EMPTY"
def encode_audio_to_base64(audio_path: str) -> str:
"""Encode a local audio file to base64 data URL."""
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
# Detect MIME type from extension
audio_path_lower = audio_path.lower()
if audio_path_lower.endswith(".wav"):
mime_type = "audio/wav"
elif audio_path_lower.endswith((".mp3", ".mpeg")):
mime_type = "audio/mpeg"
elif audio_path_lower.endswith(".flac"):
mime_type = "audio/flac"
elif audio_path_lower.endswith(".ogg"):
mime_type = "audio/ogg"
else:
mime_type = "audio/wav" # Default
with open(audio_path, "rb") as f:
audio_bytes = f.read()
audio_b64 = base64.b64encode(audio_bytes).decode("utf-8")
return f"data:{mime_type};base64,{audio_b64}"
def run_tts_generation(args) -> None:
"""Run TTS generation via OpenAI-compatible /v1/audio/speech API."""
# Build request payload
payload = {
"model": args.model,
"input": args.text,
"voice": args.speaker,
"response_format": args.response_format,
}
if args.sample_rate is not None:
payload["sample_rate"] = args.sample_rate
# Add optional parameters
if args.instructions:
payload["instructions"] = args.instructions
if args.task_type:
payload["task_type"] = args.task_type
if args.language:
payload["language"] = args.language
if args.max_new_tokens:
payload["max_new_tokens"] = args.max_new_tokens
# Voice clone parameters (Base task)
if args.ref_audio:
if args.ref_audio.startswith(("http://", "https://")):
payload["ref_audio"] = args.ref_audio
elif args.ref_audio.startswith("data:"):
payload["ref_audio"] = args.ref_audio
else:
payload["ref_audio"] = encode_audio_to_base64(args.ref_audio)
if args.ref_text:
payload["ref_text"] = args.ref_text
if args.x_vector_only:
payload["x_vector_only_mode"] = True
if args.speaker_embedding:
with open(args.speaker_embedding) as f:
payload["speaker_embedding"] = json.load(f)
print(f"Model: {args.model}")
print(f"Task type: {args.task_type or 'CustomVoice'}")
print(f"Text: {args.text}")
print(f"Speaker: {args.speaker}")
print("Generating audio...")
# Make the API call
api_url = f"{args.api_base}/v1/audio/speech"
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {args.api_key}",
}
with httpx.Client(timeout=300.0) as client:
response = client.post(api_url, json=payload, headers=headers)
if response.status_code != 200:
print(f"Error: {response.status_code}")
print(response.text)
return
# Check for JSON error response (only if content is valid UTF-8 text)
try:
text = response.content.decode("utf-8")
if text.startswith('{"error"'):
print(f"Error: {text}")
return
except UnicodeDecodeError:
pass # Binary audio data, not an error
# Save audio response
output_path = args.output or "tts_output.wav"
with open(output_path, "wb") as f:
f.write(response.content)
print(f"Audio saved to: {output_path}")
def parse_args():
"""Parse command line arguments."""
parser = argparse.ArgumentParser(
description="OpenAI-compatible client for Qwen3-TTS via /v1/audio/speech",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
# Server configuration
parser.add_argument(
"--api-base",
type=str,
default=DEFAULT_API_BASE,
help=f"API base URL (default: {DEFAULT_API_BASE})",
)
parser.add_argument(
"--api-key",
type=str,
default=DEFAULT_API_KEY,
help="API key (default: EMPTY)",
)
parser.add_argument(
"--model",
"-m",
type=str,
default="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
help="Model name/path",
)
# Task configuration
parser.add_argument(
"--task-type",
"-t",
type=str,
default=None,
choices=["CustomVoice", "VoiceDesign", "Base"],
help="TTS task type (default: CustomVoice)",
)
# Input text
parser.add_argument(
"--text",
type=str,
required=True,
help="Text to synthesize",
)
# Voice/speaker
parser.add_argument(
"--speaker",
type=str,
default="vivian",
help="Speaker name (default: vivian). Options: vivian, ryan, aiden, etc.",
)
parser.add_argument(
"--language",
type=str,
default=None,
help="Language: Auto, Chinese, English, etc.",
)
parser.add_argument(
"--instructions",
type=str,
default=None,
help="Voice style/emotion instructions",
)
# Base (voice clone) parameters
parser.add_argument(
"--ref-audio",
type=str,
default=None,
help="Reference audio file path, URL, or base64 for voice cloning (Base task)",
)
parser.add_argument(
"--ref-text",
type=str,
default=None,
help="Reference audio transcript for voice cloning (Base task)",
)
parser.add_argument(
"--x-vector-only",
action="store_true",
help="Use x-vector only mode for voice cloning (no ICL)",
)
parser.add_argument(
"--speaker-embedding",
type=str,
default=None,
help="Path to JSON file containing a pre-computed speaker embedding vector (1024-dim for 0.6B, 2048-dim for 1.7B)",
)
# Generation parameters
parser.add_argument(
"--max-new-tokens",
type=int,
default=None,
help="Maximum new tokens to generate",
)
# Output
parser.add_argument(
"--response-format",
type=str,
default="wav",
choices=["wav", "mp3", "flac", "pcm", "aac", "opus"],
help="Audio output format (default: wav)",
)
parser.add_argument(
"--sample-rate",
type=int,
default=None,
choices=[8000, 24000],
help="Output sample rate in Hz (default: model native)",
)
parser.add_argument(
"--output",
"-o",
type=str,
default=None,
help="Output audio file path (default: tts_output.wav)",
)
return parser.parse_args()
if __name__ == "__main__":
args = parse_args()
run_tts_generation(args)
qwen3_tts/precompute_custom_voice.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/qwen3_tts/precompute_custom_voice.py.
qwen3_tts/run_gradio_demo.sh
#!/bin/bash
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project
# Launch both vLLM server and Gradio demo for Qwen3-TTS
#
# Usage:
# ./run_gradio_demo.sh # Default: CustomVoice
# ./run_gradio_demo.sh --task-type VoiceDesign # VoiceDesign model
# ./run_gradio_demo.sh --task-type Base --gradio-port 7861
#
# Options:
# --task-type TYPE Task type: CustomVoice, VoiceDesign, Base (default: CustomVoice)
# --server-port PORT Port for vLLM server (default: 8000)
# --gradio-port PORT Port for Gradio demo (default: 7860)
# --server-host HOST Host for vLLM server (default: 0.0.0.0)
# --gradio-ip IP IP for Gradio demo (default: 127.0.0.1)
# --share Share Gradio demo publicly
set -e
# Default values
TASK_TYPE="CustomVoice"
SERVER_PORT=8000
GRADIO_PORT=7860
SERVER_HOST="0.0.0.0"
GRADIO_IP="127.0.0.1"
GRADIO_SHARE=false
# Parse command line arguments
while [[ $# -gt 0 ]]; do
case $1 in
--task-type)
TASK_TYPE="$2"
shift 2
;;
--server-port)
SERVER_PORT="$2"
shift 2
;;
--gradio-port)
GRADIO_PORT="$2"
shift 2
;;
--server-host)
SERVER_HOST="$2"
shift 2
;;
--gradio-ip)
GRADIO_IP="$2"
shift 2
;;
--share)
GRADIO_SHARE=true
shift
;;
--help)
echo "Usage: $0 [OPTIONS]"
echo ""
echo "Options:"
echo " --task-type TYPE Task type: CustomVoice, VoiceDesign, Base (default: CustomVoice)"
echo " --server-port PORT Port for vLLM server (default: 8000)"
echo " --gradio-port PORT Port for Gradio demo (default: 7860)"
echo " --server-host HOST Host for vLLM server (default: 0.0.0.0)"
echo " --gradio-ip IP IP for Gradio demo (default: 127.0.0.1)"
echo " --share Share Gradio demo publicly"
echo ""
exit 0
;;
*)
echo "Unknown option: $1"
echo "Use --help for usage information"
exit 1
;;
esac
done
# Map task type to model
case "$TASK_TYPE" in
CustomVoice)
MODEL="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice"
;;
VoiceDesign)
MODEL="Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign"
;;
Base)
MODEL="Qwen/Qwen3-TTS-12Hz-1.7B-Base"
;;
*)
echo "Unknown task type: $TASK_TYPE"
echo "Supported: CustomVoice, VoiceDesign, Base"
exit 1
;;
esac
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
API_BASE="http://localhost:${SERVER_PORT}"
echo "=========================================="
echo "Qwen3-TTS Gradio Demo"
echo "=========================================="
echo "Task Type : $TASK_TYPE"
echo "Model : $MODEL"
echo "Server : http://${SERVER_HOST}:${SERVER_PORT}"
echo "Gradio : http://${GRADIO_IP}:${GRADIO_PORT}"
echo "=========================================="
# Cleanup on exit
cleanup() {
echo ""
echo "Shutting down..."
if [ -n "$SERVER_PID" ]; then
echo "Stopping vLLM server (PID: $SERVER_PID)..."
kill "$SERVER_PID" 2>/dev/null || true
wait "$SERVER_PID" 2>/dev/null || true
fi
if [ -n "$GRADIO_PID" ]; then
echo "Stopping Gradio demo (PID: $GRADIO_PID)..."
kill "$GRADIO_PID" 2>/dev/null || true
wait "$GRADIO_PID" 2>/dev/null || true
fi
echo "Cleanup complete"
exit 0
}
trap cleanup SIGINT SIGTERM
# Start vLLM server
echo ""
echo "Starting vLLM server..."
LOG_FILE="/tmp/vllm_tts_server_${SERVER_PORT}.log"
vllm-omni serve "$MODEL" \
--deploy-config vllm_omni/deploy/qwen3_tts.yaml \
--host "$SERVER_HOST" \
--port "$SERVER_PORT" \
--gpu-memory-utilization 0.9 \
--trust-remote-code \
--omni 2>&1 | tee "$LOG_FILE" &
SERVER_PID=$!
# Wait for server startup
echo ""
echo "Waiting for vLLM server to be ready..."
STARTUP_FLAG="/tmp/vllm_tts_startup_flag_${SERVER_PORT}.tmp"
rm -f "$STARTUP_FLAG"
(
tail -f "$LOG_FILE" 2>/dev/null | grep -m 1 "Application startup complete" > /dev/null && touch "$STARTUP_FLAG"
) &
TAIL_PID=$!
MAX_WAIT=300
ELAPSED=0
while [ $ELAPSED -lt $MAX_WAIT ]; do
if [ -f "$STARTUP_FLAG" ]; then
kill "$TAIL_PID" 2>/dev/null || true
wait "$TAIL_PID" 2>/dev/null || true
echo ""
echo "vLLM server is ready!"
break
fi
if ! kill -0 "$SERVER_PID" 2>/dev/null; then
kill "$TAIL_PID" 2>/dev/null || true
echo ""
echo "Error: vLLM server failed to start"
exit 1
fi
sleep 1
ELAPSED=$((ELAPSED + 1))
done
rm -f "$STARTUP_FLAG"
if [ $ELAPSED -ge $MAX_WAIT ]; then
kill "$TAIL_PID" 2>/dev/null || true
echo "Error: Server startup timed out after ${MAX_WAIT}s"
kill "$SERVER_PID" 2>/dev/null || true
exit 1
fi
# Start Gradio demo
echo ""
echo "Starting Gradio demo..."
cd "$SCRIPT_DIR"
GRADIO_CMD=("python" "gradio_demo.py" "--api-base" "$API_BASE" "--host" "$GRADIO_IP" "--port" "$GRADIO_PORT" "--task-type" "$TASK_TYPE")
if [ "$GRADIO_SHARE" = true ]; then
GRADIO_CMD+=("--share")
fi
"${GRADIO_CMD[@]}" &
GRADIO_PID=$!
echo ""
echo "=========================================="
echo "Both services are running!"
echo "=========================================="
echo "vLLM Server : http://${SERVER_HOST}:${SERVER_PORT}"
echo "Gradio Demo : http://${GRADIO_IP}:${GRADIO_PORT}"
echo ""
echo "Press Ctrl+C to stop both services"
echo "=========================================="
echo ""
wait $SERVER_PID $GRADIO_PID || true
cleanup
qwen3_tts/run_server.sh
#!/bin/bash
# Launch vLLM-Omni server for Qwen3-TTS models
#
# Usage:
# ./run_server.sh # Default: CustomVoice model
# ./run_server.sh CustomVoice # CustomVoice model
# ./run_server.sh VoiceDesign # VoiceDesign model
# ./run_server.sh Base # Base (voice clone) model
set -e
TASK_TYPE="${1:-CustomVoice}"
case "$TASK_TYPE" in
CustomVoice)
MODEL="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice"
;;
VoiceDesign)
MODEL="Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign"
;;
Base)
MODEL="Qwen/Qwen3-TTS-12Hz-1.7B-Base"
;;
*)
echo "Unknown task type: $TASK_TYPE"
echo "Supported: CustomVoice, VoiceDesign, Base"
exit 1
;;
esac
echo "Starting Qwen3-TTS server with model: $MODEL"
vllm-omni serve "$MODEL" \
--deploy-config vllm_omni/deploy/qwen3_tts.yaml \
--host 0.0.0.0 \
--port 8091 \
--gpu-memory-utilization 0.9 \
--trust-remote-code \
--omni
qwen3_tts/speaker_embedding_interpolation.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/qwen3_tts/speaker_embedding_interpolation.py.
qwen3_tts/streaming_speech_client.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/qwen3_tts/streaming_speech_client.py.
qwen3_tts/tts_common.py
"""Shared constants, helpers, and payload building for Qwen3-TTS Gradio demos."""
import base64
import io
try:
import gradio as gr
except ImportError:
raise ImportError("gradio is required to run this demo. Install it with: pip install 'vllm-omni[demo]'") from None
import httpx
import numpy as np
import soundfile as sf
SUPPORTED_LANGUAGES = [
"Auto",
"Chinese",
"English",
"Japanese",
"Korean",
"German",
"French",
"Russian",
"Portuguese",
"Spanish",
"Italian",
]
TASK_TYPES = ["CustomVoice", "VoiceDesign", "Base"]
PCM_SAMPLE_RATE = 24000
DEFAULT_API_BASE = "http://localhost:8000"
def fetch_voices(api_base: str) -> list[str]:
"""Fetch available voices from the server."""
try:
with httpx.Client(timeout=10.0) as client:
resp = client.get(
f"{api_base}/v1/audio/voices",
headers={"Authorization": "Bearer EMPTY"},
)
if resp.status_code == 200:
data = resp.json()
voices = data.get("voices") or []
if voices:
return voices
except Exception:
pass
return ["Vivian", "Ryan"]
def encode_audio_to_base64(audio_data: tuple) -> str:
"""Encode Gradio audio input (sample_rate, numpy_array) to base64 data URL."""
sample_rate, audio_np = audio_data
if audio_np.dtype != np.int16:
if audio_np.dtype in (np.float32, np.float64):
audio_np = np.clip(audio_np, -1.0, 1.0)
audio_np = (audio_np * 32767).astype(np.int16)
else:
audio_np = audio_np.astype(np.int16)
buf = io.BytesIO()
sf.write(buf, audio_np, sample_rate, format="WAV")
wav_b64 = base64.b64encode(buf.getvalue()).decode("utf-8")
return f"data:audio/wav;base64,{wav_b64}"
def build_payload(
text: str,
task_type: str,
voice: str,
language: str,
instructions: str,
ref_audio: tuple | None,
ref_audio_url: str,
ref_text: str,
x_vector_only: bool,
response_format: str = "pcm",
speed: float = 1.0,
stream: bool = True,
) -> dict:
"""Build the /v1/audio/speech request payload.
Raises gr.Error for invalid input so callers don't need to validate.
"""
if not text or not text.strip():
raise gr.Error("Please enter text to synthesize.")
payload: dict = {
"input": text.strip(),
"response_format": "pcm" if stream else response_format,
"stream": stream,
}
if stream:
payload["stream_format"] = "audio"
if not stream:
payload["speed"] = speed
if task_type:
payload["task_type"] = task_type
if language:
payload["language"] = language
if task_type == "CustomVoice":
if voice:
payload["voice"] = voice
if instructions and instructions.strip():
payload["instructions"] = instructions.strip()
elif task_type == "VoiceDesign":
if not instructions or not instructions.strip():
raise gr.Error("VoiceDesign task requires voice style instructions.")
payload["instructions"] = instructions.strip()
elif task_type == "Base":
ref_audio_url_stripped = ref_audio_url.strip() if ref_audio_url else ""
if ref_audio_url_stripped:
payload["ref_audio"] = ref_audio_url_stripped
elif ref_audio is not None:
payload["ref_audio"] = encode_audio_to_base64(ref_audio)
else:
raise gr.Error("Base (voice clone) task requires reference audio. Upload a file or provide a URL.")
if ref_text and ref_text.strip():
payload["ref_text"] = ref_text.strip()
if x_vector_only:
payload["x_vector_only_mode"] = True
return payload
def on_task_type_change(task_type: str):
"""Update UI visibility based on selected task type."""
if task_type == "CustomVoice":
return (
gr.update(visible=True), # voice dropdown
gr.update(visible=True, info="Optional style/emotion instructions"),
gr.update(visible=False), # ref_audio
gr.update(visible=False), # ref_audio_url
gr.update(visible=False), # ref_text
gr.update(visible=False), # x_vector_only
)
elif task_type == "VoiceDesign":
return (
gr.update(visible=False),
gr.update(visible=True, info="Required: describe the voice style"),
gr.update(visible=False),
gr.update(visible=False),
gr.update(visible=False),
gr.update(visible=False),
)
elif task_type == "Base":
return (
gr.update(visible=False),
gr.update(visible=False),
gr.update(visible=True),
gr.update(visible=True),
gr.update(visible=True),
gr.update(visible=True),
)
return (
gr.update(visible=True),
gr.update(visible=True),
gr.update(visible=False),
gr.update(visible=False),
gr.update(visible=False),
gr.update(visible=False),
)
def stream_pcm_chunks(api_base: str, payload: dict):
"""Stream raw PCM bytes from the server, yielding int16 numpy arrays.
Handles odd-byte boundaries between network chunks.
"""
leftover = b""
with httpx.Client(timeout=300.0) as client:
with client.stream(
"POST",
f"{api_base}/v1/audio/speech",
json=payload,
headers={
"Content-Type": "application/json",
"Authorization": "Bearer EMPTY",
},
) as resp:
if resp.status_code != 200:
resp.read()
raise gr.Error(f"Server error ({resp.status_code}): {resp.text}")
for chunk in resp.iter_bytes():
if not chunk:
continue
raw = leftover + chunk
usable = len(raw) - (len(raw) % 2)
leftover = raw[usable:]
if usable == 0:
continue
yield np.frombuffer(raw[:usable], dtype=np.int16).copy()
def add_common_args(parser):
"""Add CLI arguments shared by both demos."""
parser.add_argument(
"--api-base",
default=DEFAULT_API_BASE,
help=f"Base URL for the vLLM API server (default: {DEFAULT_API_BASE}).",
)
parser.add_argument(
"--host",
default="0.0.0.0",
help="Host/IP for Gradio server (default: 0.0.0.0).",
)
parser.add_argument(
"--port",
type=int,
default=7860,
help="Port for Gradio server (default: 7860).",
)
parser.add_argument(
"--share",
action="store_true",
help="Share the Gradio demo publicly.",
)
return parser
qwen3_tts/word_timestamps_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/qwen3_tts/word_timestamps_demo.py.
voxcpm2/gradio_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/voxcpm2/gradio_demo.py.
voxcpm2/openai_speech_client.py
"""OpenAI-compatible client for VoxCPM2 TTS via /v1/audio/speech endpoint.
Examples:
# Zero-shot synthesis
python openai_speech_client.py --text "Hello, this is VoxCPM2."
# Voice cloning with a local reference audio file
python openai_speech_client.py --text "Hello world" \
--ref-audio /path/to/reference.wav
# Voice cloning with a URL
python openai_speech_client.py --text "Hello world" \
--ref-audio "https://example.com/reference.wav"
Server setup:
vllm serve openbmb/VoxCPM2 --omni --host 0.0.0.0 --port 8000
"""
from __future__ import annotations
import argparse
import base64
import os
import httpx
DEFAULT_API_BASE = "http://localhost:8000"
DEFAULT_API_KEY = "sk-empty"
def encode_audio_to_base64(audio_path: str) -> str:
"""Encode a local audio file to a base64 data URL."""
if not os.path.exists(audio_path):
raise FileNotFoundError(f"Audio file not found: {audio_path}")
ext = audio_path.lower().rsplit(".", 1)[-1]
mime = {
"wav": "audio/wav",
"mp3": "audio/mpeg",
"flac": "audio/flac",
"ogg": "audio/ogg",
}.get(ext, "audio/wav")
with open(audio_path, "rb") as f:
b64 = base64.b64encode(f.read()).decode("utf-8")
return f"data:{mime};base64,{b64}"
def main() -> None:
parser = argparse.ArgumentParser(description="VoxCPM2 OpenAI speech client")
parser.add_argument("--text", type=str, required=True, help="Text to synthesize")
parser.add_argument(
"--ref-audio",
type=str,
default=None,
help="Reference audio for voice cloning (local path, URL, or data: URI)",
)
parser.add_argument("--model", type=str, default="voxcpm2")
parser.add_argument("--output", type=str, default="output.wav")
parser.add_argument("--api-base", type=str, default=DEFAULT_API_BASE)
parser.add_argument("--api-key", type=str, default=DEFAULT_API_KEY)
parser.add_argument("--response-format", type=str, default="wav")
args = parser.parse_args()
# VoxCPM2 has no predefined voices. The "voice" field is required by
# the OpenAI API schema but ignored by VoxCPM2 — use any placeholder.
# For voice cloning, pass --ref-audio instead.
payload: dict = {
"model": args.model,
"input": args.text,
"voice": "default",
"response_format": args.response_format,
}
if args.ref_audio:
ref = args.ref_audio
if ref.startswith(("http://", "https://", "data:")):
payload["ref_audio"] = ref
else:
payload["ref_audio"] = encode_audio_to_base64(ref)
url = f"{args.api_base}/v1/audio/speech"
print(f"POST {url}")
print(f" text: {args.text}")
if args.ref_audio:
print(f" ref_audio: {args.ref_audio[:80]}...")
with httpx.Client(timeout=300) as client:
resp = client.post(
url,
json=payload,
headers={"Authorization": f"Bearer {args.api_key}"},
)
if resp.status_code != 200:
print(f"Error {resp.status_code}: {resp.text[:500]}")
return
with open(args.output, "wb") as f:
f.write(resp.content)
print(f"Saved: {args.output} ({len(resp.content):,} bytes)")
if __name__ == "__main__":
main()
voxcpm2/precompute_custom_voice.py
"""Pre-compute VoxCPM2 custom voice profiles.
The generated directory can be passed to the server via
``custom_voice_dir`` in ``vllm_omni/deploy/voxcpm2.yaml``. Requests can then
use ``/v1/audio/speech`` with ``voice="<name>"`` and no per-request ref_audio.
"""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
from typing import Any
import torch
from safetensors.torch import save_file
REPO_ROOT = Path(__file__).resolve().parents[4]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from vllm_omni.utils.custom_voice_io import safe_voice_stem # noqa: E402
MANIFEST_NAME = "custom_voice_manifest.json"
def _load_tts(model: str, device: torch.device):
from vllm_omni.model_executor.models.voxcpm2.voxcpm2_import_utils import import_voxcpm2_core
VoxCPM = import_voxcpm2_core()
native = VoxCPM.from_pretrained(model, load_denoiser=False, optimize=False)
return native.tts_model.to(device).eval()
def _load_manifest(output_dir: Path, model: str) -> dict[str, Any]:
path = output_dir / MANIFEST_NAME
if path.exists():
return json.loads(path.read_text(encoding="utf-8"))
return {
"schema_version": 1,
"model_type": "voxcpm2",
"model": model,
"voices": {},
}
def _write_voice(
*,
model: str,
output_dir: Path,
voice_name: str,
ref_audio: str,
prompt_text: str | None,
mode: str,
speaker_description: str | None,
device: torch.device,
) -> None:
if mode in ("continuation", "ref_continuation") and not prompt_text:
raise ValueError("--prompt-text is required for continuation/ref_continuation modes")
tts = _load_tts(model, device)
tensors: dict[str, torch.Tensor] = {}
with torch.inference_mode():
if mode in ("reference", "ref_continuation"):
tensors["ref_audio_feat"] = tts._encode_wav(ref_audio, padding_mode="right").float().cpu().contiguous()
if mode in ("continuation", "ref_continuation"):
tensors["audio_feat"] = tts._encode_wav(ref_audio, padding_mode="left").float().cpu().contiguous()
output_dir.mkdir(parents=True, exist_ok=True)
filename = f"{safe_voice_stem(voice_name)}.safetensors"
save_file(tensors, str(output_dir / filename))
manifest = _load_manifest(output_dir, model)
entry: dict[str, Any] = {
"name": voice_name,
"file": filename,
"mode": mode,
}
if "ref_audio_feat" in tensors:
entry["ref_audio_feat_len"] = int(tensors["ref_audio_feat"].shape[0])
if "audio_feat" in tensors:
entry["audio_feat_len"] = int(tensors["audio_feat"].shape[0])
if prompt_text:
entry["prompt_text"] = prompt_text
if speaker_description:
entry["speaker_description"] = speaker_description
manifest.setdefault("voices", {})[voice_name] = entry
(output_dir / MANIFEST_NAME).write_text(json.dumps(manifest, indent=2, ensure_ascii=False), encoding="utf-8")
print(f"Wrote {output_dir / filename}")
print(f"Updated {output_dir / MANIFEST_NAME}")
def main() -> None:
parser = argparse.ArgumentParser(description="Pre-compute VoxCPM2 custom voice profile")
parser.add_argument("--model", default="openbmb/VoxCPM2", help="VoxCPM2 model path or Hugging Face ID")
parser.add_argument("--voice-name", required=True)
parser.add_argument("--ref-audio", required=True)
parser.add_argument(
"--prompt-text",
default=None,
help="Transcript of ref audio for continuation/ref_continuation modes",
)
parser.add_argument(
"--mode",
choices=["reference", "continuation", "ref_continuation"],
default="reference",
)
parser.add_argument("--speaker-description", default=None)
parser.add_argument("--output-dir", required=True)
parser.add_argument("--device", default="cuda" if torch.cuda.is_available() else "cpu")
args = parser.parse_args()
_write_voice(
model=args.model,
output_dir=Path(args.output_dir),
voice_name=args.voice_name,
ref_audio=args.ref_audio,
prompt_text=args.prompt_text,
mode=args.mode,
speaker_description=args.speaker_description,
device=torch.device(args.device),
)
if __name__ == "__main__":
main()
voxtral_tts/gradio_demo.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/voxtral_tts/gradio_demo.py.
voxtral_tts/text_preprocess.py
Large file omitted from the rendered docs. View it on GitHub: https://github.com/vllm-project/vllm-omni/blob/main/examples/online_serving/text_to_speech/voxtral_tts/text_preprocess.py.