vllm_omni.transformers_utils.configs.audio8_tts ¶
Audio8 TTS config registration with transformers AutoConfig.
Registers model_type = "arktts" (and the two sub-config types) so AutoConfig.from_pretrained("Audio8/Audio8-TTS-Preview-0.6b") resolves to vllm-omni's Qwen2-shaped config instead of the checkpoint's remote code.
This only wins when trust_remote_code is off -- transformers prefers a checkpoint's auto_map over registered classes otherwise -- which is why deploy/audio8_tts.yaml sets trust_remote_code: false.
Audio8TTSConfig ¶
Bases: PretrainedConfig
Top-level Audio8 TTS config (model_type = "arktts").
Accepts the flat HF checkpoint fields and derives text_config (Slow AR) and fast_ar_config (Fast AR). get_text_config() returns the Slow AR config, which is what Qwen2Model reads.
codec_post_intermediate_size instance-attribute ¶
codec_post_intermediate_size = int(
codec_post_intermediate_size
)
codec_post_n_local_heads instance-attribute ¶
codec_post_n_local_heads = int(codec_post_n_local_heads)
fast_ar_config instance-attribute ¶
fast_ar_config = fast_ar_config or Audio8TTSFastARConfig(
codebook_size=codebook_size,
num_codebooks=num_codebooks,
fast_dim=fast_dim,
fast_n_head=fast_n_head,
fast_n_local_heads=fast_n_local_heads,
fast_head_dim=fast_head_dim,
n_fast_layer=n_fast_layer,
fast_intermediate_size=fast_intermediate_size,
fast_attention_qkv_bias=fast_attention_qkv_bias,
fast_attention_qk_norm=fast_attention_qk_norm,
rope_base=rope_base,
norm_eps=norm_eps,
)
sub_configs class-attribute instance-attribute ¶
sub_configs = {
"text_config": Audio8TTSSlowARConfig,
"fast_ar_config": Audio8TTSFastARConfig,
}
text_config instance-attribute ¶
text_config = text_config or Audio8TTSSlowARConfig(
vocab_size=vocab_size,
dim=dim,
n_head=n_head,
n_local_heads=n_local_heads,
head_dim=head_dim,
n_layer=n_layer,
intermediate_size=intermediate_size,
max_seq_len=max_seq_len,
rope_base=rope_base,
norm_eps=norm_eps,
attention_qkv_bias=attention_qkv_bias,
attention_qk_norm=attention_qk_norm,
tie_word_embeddings=tie_word_embeddings,
codebook_size=codebook_size,
num_codebooks=num_codebooks,
semantic_begin_id=semantic_begin_id,
semantic_end_id=semantic_end_id,
eos_token_id=eos_token_id,
pad_token_id=pad_token_id,
)
Audio8TTSFastARConfig ¶
Bases: PretrainedConfig
Fast AR config: the n_fast_layer residual-codebook predictor.