Qwen-Audio-3.1-TTS: Alibaba's Voice Model Takes Direction
Most text-to-speech systems have been judged on one thing: whether the voice sounds human. That matters, but it is less useful to a studio producing an audiobook or a game character, where the voice also has to laugh on cue and slow down when the scene calls for it.
Alibaba's new Qwen-Audio-3.1-TTS is built with that second group in mind. The model turns written text into speech across several languages and Chinese dialects, and it lets users shape the delivery with plain-language instructions. Tone, tempo and speaking style do not require technical parameters. A short description is enough.
Borrowing a voice from a short recording
The feature most likely to draw attention is voice cloning. Along with the text, users can supply a brief audio clip, and the generated speech then takes on the sound of that speaker. According to the project page, the model should cope with recordings that contain background noise and reverb. That is a practical detail, because real-world samples are rarely captured in a clean studio.
German-language outlet All-AI tested this with a clip and a German sentence warning that slippery roads, pavements and especially steps are the most common cause of accidents in winter. Its verdict on the German output was very good. It rated the model's Chinese performance as the strongest.
The project page lists 16 supported languages and 20 Chinese dialect regions. It also states that the model can produce up to three minutes of speech in a single pass, which reduces the need to stitch shorter segments together.
Markers for laughter, sighs and mood changes
For anyone voicing dialogue, Qwen-Audio-3.1-TTS offers a finer level of control. Individual passages in the script can be tagged so that the model inserts a laugh, a sigh or a shift in emotional tone at exactly that point. Instead of hoping the model reads the mood correctly, the writer decides where it changes.
That kind of control points to clear use cases: audiobooks, podcasts and game characters that need a voice. In each case, the difference between a flat reading and a convincing performance often comes down to small, well-timed reactions.
What the benchmark charts show
Alibaba has published charts based on the CV3-Eval benchmark. On speaker similarity, meaning how closely the output matches the reference voice, Qwen-Audio-3.1-TTS comes out ahead of MiniMax Speech 2.8 HD and ElevenLabs v3 in many of the languages shown. That comparison carries weight given how prominent ElevenLabs has become in AI voice.
On text accuracy, the results are less clear. In some languages, competing models score slightly higher, so the lead is not uniform across every metric.
There is also a caveat on the wider evidence. The linked research paper and the independent leaderboard both refer to version 3.0, not the new release. On that leaderboard, the predecessor already sits in second place. All-AI takes this as a sign that version 3.1 is probably the best text-to-speech model available right now. That is a reasonable inference, but it is an inference rather than a published ranking of the new model.
A Flash version and lower prices for developers
Alongside the main model, Alibaba offers Qwen-Audio-3.1-TTS-Flash through Alibaba Cloud Model Studio. This version is aimed at applications that need continuous speech output as it is generated, rather than waiting for a complete audio file.
Alibaba has also reduced the prices of its audio models. Together, the Flash variant and the lower pricing target developers who want to build voices into assistants and other products where fast responses matter. The move fits a broader pattern of voice becoming a main interface for AI assistants, where delays in speech output are immediately noticeable to users.
Why it matters
The release shows where text-to-speech competition is heading. Natural-sounding output is becoming the baseline. The next distinction is control: cloning a specific voice from a noisy sample, steering delivery with everyday instructions and placing emotional cues exactly where a script needs them.
For Alibaba, the model extends the Qwen family further into consumer-facing experiences, following efforts such as its Qwen Intelligence agent platform for smartphones. Whether version 3.1 holds the top spot will depend on independent evaluations of the new release itself. Based on the published charts and the standing of its predecessor, it enters the market as a serious contender.
Sponsored Recommended for you – discover more →
