Back to feed
News Story
APriority74
阿里云开发者
1 sources

Alibaba Launches Qwen-Audio-3.0 Voice Models on Qianwen AI Platform

Alibaba has officially launched the Qwen-Audio-3.0 series of voice models on its Qianwen AI platform, including ASR, TTS, and real-time interaction variants. The models achieved top rankings in three categories on Artificial Analysis's voice leaderboard, and developers can access them via API or Token Plan subscriptions. They are already integrated into products like the Qianwen App.

SynthePulse Insight · AI deep reading

Qwen-Audio-3.0 Tops Global Speech Rankings, Alibaba's Speech Models Achieve 'Grand Slam'

Version 1 · 1 source

Alibaba has released the Qwen-Audio-3.0 series of speech models, which have secured first place globally in ASR, TTS, and real-time interaction on the Artificial Analysis speech leaderboard, and have been integrated into products such as the Qwen App. This article reviews their capabilities, pricing, and platform strategy.

  • The Qwen-Audio-3.0 series includes three models: ASR, TTS, and Realtime, now available on the Qwen AI platform.
  • The series achieved first place globally in ASR, TTS, and real-time interaction on the Artificial Analysis July speech leaderboard.
  • The Realtime model supports end-to-end speech understanding, listening while speaking, interruption at any time, and tool invocation.
  • The models have been applied in products such as the Qwen App, Qwen Office, and Qoder.
  • The Qwen AI platform has recently launched multiple models covering text, code, speech, image, and video.
  • Token Plan subscription offers limited-time discounts: personal plans start at 39 RMB/month, and team seats are discounted up to 24% off.
Open section navigationRelease and Product Matrix

Release and Product Matrix

On August 17, Alibaba announced that the Qwen-Audio-3.0 series of speech models has officially launched on the Qwen AI platform, and developers can access them via API or Token Plan subscription. The series includes the speech recognition model Qwen-Audio-3.0-ASR, the speech synthesis model Qwen-Audio-3.0-TTS, and the real-time speech interaction dialogue model Qwen-Audio-3.0-Realtime.

According to official introductions, the ASR model supports recognition in complex contexts and professional domains; the TTS model supports multiple languages and dialects, and can control emotion, tone, and rhythm; the Realtime model supports end-to-end speech understanding and dialogue, can listen while speaking, interrupt at any time for follow-up questions, and supports tool invocation.

Global Benchmark Leadership

Officially, on the July speech leaderboard of the globally authoritative AI evaluation platform Artificial Analysis, the Qwen-Audio-3.0 series achieved first place globally in three tracks: speech recognition (ASR), real-time interaction (RealTime), and speech synthesis (TTS), securing a 'Grand Slam'.

It should be noted that this achievement comes from an official press release and is a source claim, not yet independently verified by third parties. The specific scores, test sets, and comparison models have not been disclosed.

Application and Platform Ecosystem

Officially, the series has been applied in products such as the Qwen App, Qwen Office, and Qoder. This indicates that speech capabilities have moved from the model layer to the product layer, but specific application scenarios and effects have not been detailed.

The Qwen AI platform has recently launched multiple flagship models, including the new-generation base model Qwen3.8-Max, the image generation model Qwen-image 3.0 pro, the video generation model Wan3.0, and Kimi K3. The platform is attempting to cover multimodal capabilities such as text, code, speech, image, and video, creating a one-stop model invocation entry.

Commercialization and Pricing

The Token Plan subscription service offers limited-time discounts. The personal edition is divided into Lite (39 RMB/month, 2500 Credits every 7 days), Standard (139 RMB/month, 4 times Lite usage), and Pro (499 RMB/month, 16 times Lite usage). The team edition standard seat is 50 RMB/seat/month (reduced to 7.6折), the premium seat is 550 RMB/seat/month (reduced to 7.9折), and the deluxe seat is 1398 RMB/seat/month.

The pricing strategy shows Alibaba's intention to lower the barrier for developers while covering different scales of demand through multiple tiers. However, the specific consumption rules for Credits have not been announced, making actual costs difficult to assess.

Credibility boundary

The main information in this article comes from the official Alibaba Cloud developer WeChat account, which is a first-party release. However, the benchmark results and product capabilities have not been provided with third-party verification or detailed data, so the relevant conclusions should be regarded as official claims rather than independent confirmation.

Insight takeaway

The Qwen-Audio-3.0 series has, according to official statements, topped global speech benchmarks and has been quickly integrated into products and platforms, but independent evaluation is needed to verify its true performance. Alibaba is strengthening the appeal of the Qwen AI platform to developers through a multimodal model matrix and discounted subscription strategies.

Primary report

阿里云开发者

Primary source