The Decoder→ original

Qwen Audio 3.0 TTS Plus от Alibaba возглавила рейтинг синтеза речи Speech Arena

Alibaba вывела модель синтеза речи Qwen Audio 3.0 TTS Plus на первое место рейтинга Speech Arena от Artificial Analysis. Модель озвучивает текст на 16 языках и даёт управлять стилем речи обычными словами или тегами вроде [angry]. Слабое место — скорость: всего 16 символов в секунду, заметно медленнее конкурентов Sonic 3.5 и Simba 3.2.

AI-processed from The Decoder; edited by Hamidun News
Qwen Audio 3.0 TTS Plus от Alibaba возглавила рейтинг синтеза речи Speech Arena
Source: The Decoder. Collage: Hamidun News.
◐ Listen to article

Alibaba in July 2026 pushed its speech synthesis model Qwen Audio 3.0 TTS Plus to first place on the SpeechArena leaderboard from Artificial Analysis — it voices text in 16 languages and lets you set speech style with ordinary words or tags like [angry].

First Place on SpeechArena

Qwen Audio 3.0 TTS Plus topped SpeechArena — the public leaderboard for text-to-speech systems maintained by the analytics company Artificial Analysis. The leaderboard compares voice models by synthesis quality, and Alibaba's new model overtook its rivals in the overall ranking.

  • Developer — Alibaba, model Qwen Audio 3.0 TTS Plus
  • First place on the SpeechArena leaderboard from Artificial Analysis
  • Support for 16 languages
  • Speech style control: natural-language description or tags like [angry]
  • Generation speed — 16 characters per second, slower than Sonic 3.5 and Simba 3.2

How to Set the Voice's Emotion

Speech style in Qwen Audio 3.0 TTS Plus can be controlled in two ways: with a natural-language description or with service tags. For example, the [angry] tag makes the model sound irritated. This kind of control over intonation sets modern TTS models apart from the robotic synthesis of the previous generation, where the voice sounded flat and unemotional.

Support for 16 languages broadens its use: the same model can be used for voicing content, voice assistants, and dubbing across different markets without switching engines.

Why the Model Is Slower Than Competitors

Qwen Audio 3.0 TTS Plus generates speech at a rate of 16 characters per second — noticeably slower than its direct rivals Sonic 3.5 and Simba 3.2. For offline tasks like preparing an audiobook or pre-recorded voiceover, this isn't critical: the file is rendered just once anyway. But in real-time scenarios — voice agents and live conversations — that speed means a noticeable delay between text input and sound.

The result is a characteristic trade-off: Alibaba wins on synthesis quality, which is exactly what SpeechArena evaluates, but loses on speed. For some tasks, voice naturalness matters more; for others, minimal latency does.

First place on

SpeechArena at a speed of just 16 characters per second — that's how the Artificial Analysis leaderboard describes Qwen Audio 3.0 TTS Plus's result, as reported by The Decoder.

What This Means

Qwen Audio 3.0 TTS Plus's lead shows that the race among TTS models has shifted from simply "reading text intelligibly" to controlling emotion and style across many languages. But there's no universal winner yet: quality and speed remain a mutual trade-off, and the choice of model depends on what matters most for the task.

Frequently Asked Questions

How many languages does Qwen Audio 3.0 TTS Plus support?

The model voices text in 16 languages. It also lets you control speech style and emotion through a natural-language description or service tags like [angry].

Where does

Qwen Audio 3.0 TTS Plus fall short of competitors?

In synthesis speed. The model generates 16 characters per second — according to Artificial Analysis, that's noticeably slower than competitors Sonic 3.5 and Simba 3.2, which is critical for real-time scenarios.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…