Back to feed
News Story
阿里云开发者
1 sources

Alibaba Releases Qwen-Audio-3.0-Realtime Voice Interaction Model

Alibaba has released Qwen-Audio-3.0-Realtime, an upgraded voice interaction model with emotional perception, voice cloning, duplex dialogue, and dynamic tool calling. It ranks first overall in Artificial Analysis benchmarks, surpassing OpenAI's GPT-Realtime-2. The model aims to make voice interactions more natural and intelligent for applications like emotional companionship, customer service, and education.

SynthePulse Insight · AI deep reading

Qwen-Audio-3.0-Realtime: How Alibaba's Voice Model Achieves Real-Time Interaction That 'Listens and Understands'

Version 1 · 1 source

Alibaba Cloud releases Qwen-Audio-3.0-Realtime, ranking first in multiple Artificial Analysis evaluations, surpassing OpenAI GPT-Realtime-2. Its core breakthroughs lie in high-expressive speech, duplex interaction, and dynamic tool calling, though performance drops on spoken prompts in some benchmarks.

  • Qwen-Audio-3.0-Realtime ranks first overall in Artificial Analysis's Speech Reasoning subcategory, surpassing OpenAI GPT-Realtime-2.
  • The model supports emotion-aware generation, prosody dynamic adjustment, and paralinguistic information processing, enabling human-like voice interaction.
  • Built-in multimodal duplex control model supports voiceprint-level background filtering, maintaining smooth conversation in noisy environments.
  • Dynamic tool calling capability enables tasks like route planning and information retrieval, with multi-tool coordination.
  • Adopts On-Policy Distillation and multi-teacher distillation framework, balancing reasoning ability and response speed.
  • Performs well on benchmarks like VoiceBench and AudioMultiChallenge, but scores drop on spoken prompts.
Open section navigationHigh Expressiveness and Empathetic Dialogue: Breaking the Mechanical Reading Feel

High Expressiveness and Empathetic Dialogue: Breaking the Mechanical Reading Feel

Qwen-Audio-3.0-Realtime dynamically adjusts tone, rhythm, pitch, and emotional expression based on conversation context, achieving human-like voice interaction. It supports voice cloning and achieves SOTA results on the S2S voice instruction following public benchmark VStyle.

The model possesses emotion-aware generation, prosody dynamic adjustment, and paralinguistic information processing capabilities, understanding and generating non-verbal sound signals such as laughter, sighs, and hesitations. In complex contexts like debates, role-playing, and emotional companionship, it adjusts speaking style and emotional expression.

These capabilities enable applications in AI companions, intelligent customer service, educational assistants, and other scenarios, providing warm emotional support.

Smooth Duplex Interaction: Precise Dialogue in Noisy Environments

The model incorporates a multimodal-aware duplex control model that perceives rich information such as voice, environment, background, and speaker, enabling precise duplex rhythm judgment. It remains composed in noisy backgrounds, multi-person conversations, and speaker switching scenarios.

Its multimodal perception and intelligent interruption detection simultaneously analyze audio, semantics, and voiceprint features, distinguishing environmental noise from real interruptions. The Realtime-API reserves an audio_prompt field, allowing developers to implement voiceprint-level anti-interference and precisely target specific users.

Achieves SOTA results in Artificial Analysis's dialogue dynamics subcategory, leading in key metrics such as interruption detection accuracy and response latency.

Agentic and Dynamic Tool Calling: More Than Just Chat

The model possesses strong Agentic capabilities and a dynamic tool calling mechanism, intelligently deciding when to call tools and when to engage in casual conversation based on dialogue context. In TauBench tests (covering real business domains like retail, aviation, and telecom), it accurately decomposes complex spoken instructions and automatically calls external tools for multi-step reasoning.

Core features include dynamic tool routing, multi-tool coordination, unified capability access (based on FunctionCall standard protocol supporting MCP, API, knowledge base), and context awareness. It can perform tasks such as route planning, restaurant recommendations, real-time news queries, stock quotes, schedule management, and data analysis.

Foundation Support: Balancing Speed and Intelligence

To address the IQ degradation issue in voice models, the On-Policy Distillation framework is adopted, distilling the complete reasoning ability of text large models into the voice model. Additionally, a multi-teacher distillation strategy is introduced, including spoken multi-turn preference teacher, general teacher, Agentic teacher, and audio understanding teacher, ensuring the model is well-rounded.

The model provides millisecond-level response for latency-sensitive scenarios like daily Q&A. It ranks first overall in Artificial Analysis's Speech Reasoning subcategory and performs excellently on benchmarks like VoiceBench and AudioMultiChallenge.

However, note that in VoiceBench, the plus version scores 92.5 on standard prompts and 90.5 on spoken prompts (a drop of 2.0); the flash version scores 89.8 on standard prompts and 87.5 on spoken prompts (a drop of 2.4). In AudioMultiChallenge, the plus version scores 44.0 on standard prompts and 37.6 on spoken prompts (a drop of 6.4); the flash version scores 43.6 on standard prompts and 38.1 on spoken prompts (a drop of 5.5). Performance decreases on spoken prompts.

Credibility boundary

This article's information primarily comes from Alibaba Cloud's official public account article, which is a product release promotion. Some evaluation data is based on the Preview version. Artificial Analysis rankings and benchmark results are self-reported by Alibaba Cloud and have not been independently verified by third parties. The performance drop on spoken prompts is from the original text, but test conditions are not specified.

Insight takeaway

Qwen-Audio-3.0-Realtime shows significant improvements in speech expressiveness, duplex interaction, and tool calling, achieving leading rankings in authoritative evaluations. However, the performance drop on spoken prompts suggests it may still have limitations in real-world complex spoken scenarios.

Primary report

阿里云开发者

Primary source