Qwen-Audio 3.1 Release: Alibaba Cuts Voice AI Prices 95%
Alibaba's Qwen team shipped five upgraded voice AI models and slashed API prices by up to 95 percent, intensifying competition in speech recognition, TTS and real-time voice.

Alibaba's Qwen team released Qwen-Audio 3.1 on September 23, 2026, a five-model voice AI lineup that upgrades its speech recognition (ASR), text-to-speech (TTS) and real-time voice models while adding two new ones, ASR-Next and TTS-Next. Alongside the release, Alibaba Cloud cut API prices across the audio lineup by as much as 95 percent, with TTS down roughly 70 percent, Realtime down about 85 percent, and ASR down up to 95 percent, effective from 00:00 Beijing time on September 22, 2026.
- Released: September 23, 2026, by Alibaba's Qwen team, following an Apsara Conference preview on September 22
- Lineup: five models — ASR, ASR-Next, TTS, TTS-Next, Realtime
- Price cuts: TTS ~70% cheaper, Realtime ~85% cheaper, ASR up to 95% cheaper
- Effective date for new pricing: 00:00 Beijing time, September 22, 2026
- Where to get it: Alibaba Cloud Model Studio API and the Qwen Cloud console (qwencloud.com)
- Realtime context window: up to 262,144 tokens on Qwen-Audio-3.1-Realtime-Plus
What is Qwen-Audio 3.1?
Qwen-Audio 3.1 is Alibaba's latest voice AI product line, built around what the Qwen team calls "one complete audio stack: understanding, generation, interaction and creation." It replaces the standalone Qwen-Audio 3.0 models with five API-accessible models covering four jobs: turning speech into text (ASR), turning text into speech (TTS), holding a live spoken conversation (Realtime), and generating full audio scenes such as dialogue mixed with sound effects and ambience (TTS-Next).
The release was folded into Alibaba's broader Apsara Conference announcements on full-stack AI strategy, where the company also introduced Qwen3.8-LiveTranslate, a simultaneous-interpretation model, and previewed its Qwen 4.5 and Qwen 5 roadmap. But the audio stack got its own dedicated rollout a day later, when Qwen's official account posted the five-model lineup along with the pricing news that has driven most of the search interest since.
The five Qwen-Audio 3.1 models, explained
Each model in the lineup targets a different part of the voice pipeline. According to Qwen's own release notes and Alibaba Cloud Model Studio documentation, the breakdown is:
| Model | Job | Status | What's new |
|---|---|---|---|
| Qwen-Audio-3.1-ASR | Speech recognition | Upgraded | Stronger multilingual and Chinese-dialect recognition; automatically strips filler words and repetitions from transcripts |
| Qwen-Audio-3.1-ASR-Next | Audio understanding | New | Multi-speaker diarization with timestamps; detects emotion, ambient sound and machine noise for captioning and audio QA |
| Qwen-Audio-3.1-TTS | Text-to-speech | Upgraded | Multilingual and dialect synthesis with cross-lingual voice transfer; emotion, speed and style controlled via natural-language instructions |
| Qwen-Audio-3.1-TTS-Next | Audio creation ("AudioGen") | New | Unified language-model-plus-diffusion pipeline that generates voice, sound effects and background audio in a single pass, for podcasts, audiobooks and games |
| Qwen-Audio-3.1-Realtime | Live voice interaction | Upgraded | Full-duplex speaking and listening with instant interruption; tool use for agentic tasks; a Realtime-Plus variant with a 262K-token context window |
The ASR upgrade focuses on cleanup and coverage: Alibaba Cloud's model documentation describes Qwen-Audio-3.1-ASR-Flash as tuned for "high-quality, multilingual, and cultural content transcription scenarios," including Chinese dialects and even classical-poetry rhythm recognition, with punctuation prediction and text normalization built in. ASR-Next goes further, adding speaker labels and timestamps for multi-person recordings and tagging emotional tone and background sound events — closer to an audio-understanding model than a plain transcriber.
On the generation side, Qwen-Audio-3.1-TTS is described in the team's own technical write-up as a "production-oriented speech synthesis system" built on a 12.5 Hz low-frame-rate speech tokenizer and a five-stage training pipeline, supporting 16 languages and 20 Chinese dialect regions, one-pass synthesis up to three minutes long, and 86 fine-grained inline tags for non-verbal sounds like laughter or breathing. Qwen's team says the model ranks first on the independent Artificial Analysis Text-to-Speech Leaderboard. TTS-Next is the more novel addition: rather than just reading text aloud, it composes speech, sound effects and ambience together from a script, aimed at podcast production, audiobooks and game or film soundscapes.
Qwen-Audio-3.1-Realtime is the conversational model, built for products like voice assistants and customer-service bots that need to listen and speak at once and handle mid-sentence interruptions naturally. A companion research paper on Qwen-Audio-3.1-Realtime, published around the same time, reports the upgrade raised overall task success from 78.4 percent to 82.0 percent on a half-duplex speech-to-text adaptation benchmark, and cut the model's tendency to respond to irrelevant background speech from a 73.0 percent false-trigger rate down to 13.0 percent — a meaningful gain for any product using it as an always-listening agent.
Qwen-Audio 3.1 price cuts: what actually changed
The headline number from Qwen's own announcement is blunt: TTS API pricing is down about 70 percent, Realtime is down about 85 percent, and ASR is down by as much as 95 percent, compared with the equivalent Qwen-Audio 3.0 tiers. Alibaba Cloud's Model Studio pricing notice confirms the change took effect at 00:00 Beijing time on September 22, 2026, across both its Singapore (international) and Beijing sites, and applies automatically with no action required from existing API users.
The clearest documented example is the realtime tier. Alibaba Cloud's own price-change notice lists the before-and-after rates for the Flash variant like this:
| Qwen-Audio Realtime-Flash (Singapore, per million tokens) | Before | After | Change |
|---|---|---|---|
| Input text | $0.45 | $0.23 | ~49% lower |
| Input audio | $4.50 | $0.93 | ~79% lower |
| Output text | $4.50 | $0.70 | ~84% lower |
| Output (text + audio) | $15.00 | $1.87 | ~88% lower |
Those per-line reductions land between roughly 49 percent and 88 percent depending on the input or output type, which is consistent with Qwen's own "~85% off" headline once averaged across a typical mixed conversation. Alibaba Cloud's notice also lists a new rate for Qwen-Audio-3.1-TTS-Flash of $0.23 per million input tokens and $1.87 per million output tokens on the Singapore site, though the prior TTS rate was billed per 10,000 input characters rather than per token, so the two units aren't directly comparable line for line; the percentage Qwen cites (~70%) is the company's own like-for-like comparison rather than one this article can independently recompute from the notice alone.
Current list pricing on Alibaba Cloud Model Studio, for reference, includes qwen-audio-3.1-asr-flash at $0.113 per million input tokens and $0.382 per million output tokens, and qwen-audio-3.1-tts-next (the new AudioGen model) at 6 CNY per million input tokens and 12 CNY per million output tokens on the Beijing site. Qwen-Audio-3.1-Realtime-Plus, the higher-context realtime variant, lists at $0.80 per million text-input tokens, $6.40 per million audio-input tokens, $6.40 per million text-output tokens and $24 per million combined text-and-audio output tokens — a reminder that the "Plus" and "Flash" tiers are priced very differently, and the percentage cuts apply within each tier rather than across the whole family uniformly.
How the new prices compare with before
For anyone shopping around, the practical read is that Alibaba has pushed its audio APIs toward the aggressive end of what other major voice AI vendors charge, particularly on ASR and Flash-tier realtime interaction. A 95 percent cut on ASR pricing, for a model family that already competes on transcription accuracy, makes high-volume transcription — call center audits, meeting notes, subtitling — dramatically cheaper to run at scale than it was under Qwen-Audio 3.0 pricing just weeks earlier. The same logic applies to conversational Realtime-Flash traffic, where the input-audio rate alone fell by roughly 79 percent.
It's worth comparing this against how other frontier labs price comparable capability tiers; Pandromeda's own coverage of recent GPT-6 API pricing and how ChatGPT, Claude and Gemini plans stack up on price shows the same general pattern playing out across modalities: as more vendors ship credible alternatives, list prices for API access keep moving down, not up.
How to access Qwen-Audio 3.1
Qwen-Audio 3.1 is not (as of this release) a downloadable open-weight checkpoint in the way Qwen's smaller Qwen3-ASR and Qwen3-TTS models are on GitHub and Hugging Face. Instead, it ships as a hosted API:
- Alibaba Cloud Model Studio — developers call the models (e.g.
qwen-audio-3.1-asr-flash,qwen-audio-3.1-tts-flash,qwen-audio-3.1-tts-next) over HTTP using a DashScope API key, with per-model documentation covering supported languages, audio formats, context limits and rate limits. - Realtime via WebSocket — Qwen-Audio-3.1-Realtime and Realtime-Plus are accessed through a WebSocket endpoint for full-duplex, low-latency streaming rather than a simple request/response call, which is the standard pattern for conversational voice APIs.
- Qwen Cloud console — Qwen has also started listing individual audio models with their own pages and quick-start links on qwencloud.com, alongside the Model Studio documentation.
For teams building agentic voice products — assistants that need to use tools mid-conversation rather than just transcribe or narrate — the Realtime model's function-calling and tool-use support is the more relevant feature than raw pricing. Readers who want the broader context on how these systems are built can see Pandromeda's explainer on what an AI agent actually is and how it works.
Benchmarks and quality claims
Alongside the price story, Qwen has published performance claims meant to justify treating 3.1 as more than a discount refresh. On text-to-speech, the team's technical write-up reports state-of-the-art or top-tier results on SEED-TTS-Eval, CV3-Eval, instruction-following, long-form synthesis and robustness benchmarks, and a first-place ranking on the independent Artificial Analysis Text-to-Speech Leaderboard. On dialect recognition, Qwen's own published charts show Qwen-Audio-3.1-ASR posting lower character error rates than Doubao-ASR and close to, or better than, Tencent's Hy-ASR-3.0-preview across most tested Chinese dialect and Cantonese (WSYue) benchmarks.
On the Realtime side, the accompanying research paper on agentic voice interaction reports the version bump raised task success from 78.4 percent to 82.0 percent in half-duplex evaluation, and — arguably the more practically important number — cut the rate at which the model incorrectly responds to background speech it overhears from 73.0 percent down to 13.0 percent. That kind of false-trigger reduction matters more for real deployments than headline accuracy, since an assistant that talks over background noise is unusable regardless of how good its transcription is.
Why Alibaba is cutting voice AI prices so aggressively
Price cuts this steep, arriving in the same announcement as a genuine model upgrade, point to a voice AI market that's getting crowded fast. Alibaba's own benchmark comparisons name Doubao (ByteDance) and Tencent's Hy-ASR as direct rivals on transcription quality, and the broader real-time and text-to-speech categories have filled out with competing offerings from Western labs over the past year. Cutting ASR pricing by up to 95 percent is the kind of move a vendor makes when it wants transcription volume to be a non-factor in a buyer's decision, shifting the competition purely to accuracy, language coverage and latency.
It also fits a pattern that's shown up repeatedly across the AI industry this year: once a modality's underlying compute cost drops and more capable competitors ship, list prices compress quickly. A released model rarely holds its launch price for long once a comparable alternative appears — and voice AI, after lagging text generation for a while, is now going through the same cycle.
What's next for Qwen-Audio
Qwen's announcement described "more APIs coming soon," suggesting the five-model lineup announced this week isn't the final shape of the 3.1 generation. Alibaba's broader Apsara Conference roadmap also flagged Qwen3.8-LiveTranslate for simultaneous interpretation and previewed Qwen 4.5 and Qwen 5 as the company's next flagship text and multimodal models, which will likely fold in further audio capability over time. For now, the near-term story is straightforward: Qwen-Audio 3.1 gives developers a cheaper, broader voice AI stack than they had a week ago, and the pricing pressure it puts on rival ASR, TTS and realtime-voice vendors is probably the bigger long-term story than any single benchmark number.
Frequently asked questions
What is Qwen-Audio 3.1?
Qwen-Audio 3.1 is a five-model voice AI lineup from Alibaba's Qwen team, released September 23, 2026. It covers speech recognition (ASR and ASR-Next), text-to-speech (TTS and TTS-Next), and real-time voice conversation (Realtime), all available through Alibaba Cloud's API.
When was Qwen-Audio 3.1 released, and how much cheaper is it?
Qwen-Audio 3.1 launched September 23, 2026, a day after Alibaba previewed it at the Apsara Conference. Alibaba Cloud cut API pricing on the lineup effective 00:00 Beijing time on September 22, 2026, with TTS down about 70 percent, Realtime down about 85 percent, and ASR down by as much as 95 percent versus Qwen-Audio 3.0 pricing.
What are the five Qwen-Audio 3.1 models?
They are Qwen-Audio-3.1-ASR and the new Qwen-Audio-3.1-ASR-Next for speech understanding, Qwen-Audio-3.1-TTS and the new Qwen-Audio-3.1-TTS-Next for speech and audio generation, and Qwen-Audio-3.1-Realtime (with a higher-context Realtime-Plus variant) for live voice interaction.
How do I access Qwen-Audio 3.1?
The models are available as a hosted API through Alibaba Cloud Model Studio, using a DashScope API key for ASR and TTS calls and a WebSocket endpoint for the Realtime models. Individual model pages are also listed on Alibaba's Qwen Cloud console. There is no downloadable open-weight release of Qwen-Audio 3.1 at this time.
What's new in Qwen-Audio-3.1-TTS-Next and ASR-Next?
TTS-Next is an 'AudioGen' model that generates voice, sound effects and background ambience together in one pass from a script, aimed at podcasts, audiobooks and game or film sound design. ASR-Next adds multi-speaker diarization with timestamps and can detect emotion, ambient sound and machine noise, going beyond plain transcription.
How does Qwen-Audio-3.1-Realtime compare with the previous version?
According to Qwen's published research, Qwen-Audio-3.1-Realtime raised task success from 78.4 percent to 82.0 percent on a half-duplex evaluation and cut its false-trigger rate on background speech from 73.0 percent to 13.0 percent, making it noticeably better at ignoring irrelevant audio during a live conversation.
Sources
- Alibaba Unveils Roadmap on Full-Stack AI Strategy (Alibaba Cloud Press Room)alibabacloud.com
- Model Studio Price Reduction Notice for Selected Audio Models (Alibaba Cloud)alibabacloud.com
- Qwen-Audio-3.1 announcement (@Alibaba_Qwen on X)x.com
- qwen-audio-3.1-realtime-plus model page (Qwen Cloud)qwencloud.com
- qwen-audio-3.1-tts-flash model info (Alibaba Cloud Model Studio docs)help.aliyun.com
- Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction (arXiv)arxiv.org
Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.


