Qwen Audio 3.0 Tts Flash

qwen-audio-3.0-tts-flash · Qwen

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.

API Pricing

Input$14.2 / 1M tokens
Output$14.2 / 1M tokens

Specifications

Modalitiestext

Frequently asked questions

What is Qwen Audio 3.0 Tts Flash?

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.

How much does Qwen Audio 3.0 Tts Flash cost?

On AIHubMix, Qwen Audio 3.0 Tts Flash costs $14.2 per million input tokens and $14.2 per million output tokens.

What modalities does Qwen Audio 3.0 Tts Flash support?

Qwen Audio 3.0 Tts Flash accepts text input.

How do I call Qwen Audio 3.0 Tts Flash via API?

Qwen Audio 3.0 Tts Flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to qwen-audio-3.0-tts-flash — no other code changes needed.

Who created Qwen Audio 3.0 Tts Flash?

Qwen Audio 3.0 Tts Flash is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More models from Qwen

See all Qwen models →

Qwen3.8 Flash

by Qwen

Qwen3.8 Flash is Alibaba Cloud Qwen’s flagship native vision-language model for coding…

$0.113/1M in · $0.38/1M out
1,000,000 tokens context

Wan3.0 Video

by Qwen

Wan 3.0 (Tongyi Wanxiang 3.0) is an integrated video generation and editing model…

$2/1M in · $2/1M out

Wan3.0 Video Prime

by Qwen

Wan3.0 Video Prime is Alibaba Cloud’s preview high-speed edition of its All-in-One video…

$2/1M in · $2/1M out

Qwen3.8 2.4t A95B

by Qwen

Qwen3.8-2.4T-A95B is Alibaba’s most powerful Qwen model to date. It is a…

$2/1M in · $6/1M out
262,000 tokens context

Qwen Image 3.0

by Qwen

Qwen Image 3.0(qwen-image-3.0) is an image generation and editing model developed by…

$2/1M in

Qwen Image 3.0 Pro

by Qwen

Qwen Image 3.0 Pro (qwen-image-3.0-pro) is Alibaba Cloud Qwen’s flagship image generation…

$2/1M in

Use Qwen Audio 3.0 Tts Flash via the AIHubMix unified API — one interface for every major LLM.