Back to feed
News Story
Artificial Analysis (X)
1 sources

Hugging Face Launches Speech Agent Arena to Evaluate Speech-to-Speech Models

Hugging Face announced the Speech Agent Arena, a new benchmark for evaluating speech-to-speech models in real-world scenarios. It measures conversational preference and task success rate through human comparisons, with initial results showing Google's Gemini 3.1 Flash Live Preview leading the leaderboard.

SynthePulse Insight · AI deep readingMembers

Voice Agent Arena: Why Do Preferences and Task Success Rates Diverge?

Version 1 · 1 source

Artificial Analysis releases Speech Agent Arena to evaluate speech-to-speech models in real-world scenarios. Gemini 3.1 Flash Live Preview leads in preference, but its task success rate is lower than Grok Voice and GPT-Realtime series, revealing the gap between 'sounding good' and 'doing well.'

  • Speech Agent Arena, introduced by Artificial Analysis, evaluates voice models using 15 agentic and 20 non-agentic scenarios, measuring conversational preference and task success rate.
  • Gemini 3.1 Flash Live Preview - Minimal leads the preference leaderboard with 1046 Elo, but its task success rate is only 74.6%, lower than Grok Voice Think Fast 2.0 High's 94.7%.
  • The top task success rate is Grok Voice Think Fast 2.0 High (94.7%), followed by GPT-Realtime-2.1 High (91.5%) and ElevenLabs Agents (90.5%).
Open section navigationPositioning and Design of the New Benchmark

Positioning and Design of the New Benchmark

On August 21, 2026, Artificial Analysis released Speech Agent Arena, aiming to evaluate speech-to-speech models in real-world scenarios. The benchmark covers 15 agentic scenarios (requiring tool calls, such as ordering food) and 20 non-agentic scenarios (no tool calls, such as asking for business hours). It collects preference votes from human participants in live conversations with two hidden models and fits Elo scores.

Task success rate is defined as the proportion of qualifying conversations where the model completes the request with the correct final tool call. To reduce overfitting, scenario prompts, tool schemas, and participant instructions are kept confidential except for one example; participants are paid, screened third-party individuals.

Free for now

Read the full analysis

3 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This report is based on the official release from Artificial Analysis, a primary source. All data comes from their announcement and has not been independently verified. Some conclusions (such as the reasons for the divergence between preference and task success) are inferences based on observations in the release.

Primary report

Artificial Analysis (X)

Primary source