Alibaba’s Qwen-Audio-3.0-TTS: Flash and Plus models covering 16 languages
Decision Brief
Qwen-Audio-3.0-TTS offers two variants: Flash for real-time interaction (~300ms first-packet latency) and Plus for high-quality generation. Plus achieves first place on the Artificial Analysis Voice Arena with Elo ~1236, priced at ~$27.59 per million characters—roughly one-third of ElevenLabs and MiniMax equivalents. However, Plus throughput is ~16 chars/s, lower than Simba 3.2 (30.2) and Sonic 3.5 (120). The model supports 16 languages and 20 Chinese dialects, achieving the lowest character error rate in 10/16 languages (Flash 3.87%, Plus 3.96%), with Plus scoring 82.75 in speaker similarity (first). It introduces 86 fine-grained inline tags for non-speech events like laughter, breathing, and coughing, but only in unidirectional streaming mode. Built on a 12.5Hz low-frame-rate speech tokenizer and a five-stage progressive training paradigm, it supports up to 3-minute single synthesis and 48kHz output. For developers needing low-latency real-time TTS, Flash is ideal; teams prioritizing audio quality and naturalness will find Plus highly competitive in cost and quality, but should evaluate throughput constraints.
Sources
- MarkTechPost
Fast research-paper and ML tooling summaries, useful for infra and agent updates.
- MarkTechPost
留言
登入后即可留言,和其他 builder 交换实测心得。
还没有留言,抢头香。