Skip to main content
Back to Blog
AI/MLData Analysis
13 August 20264 min readUpdated 25 August 2026

Speech Agent Arena Evaluates Speech-to-Speech Models in Real-World Conversations

Speech Agent Arena Evaluates Speech to Speech Models in Real World Conversations The new Speech Agent Arena evaluates speech to speech models and cascaded systems on real world...

By AI Engineering Team

Speech Agent Arena Evaluates Speech-to-Speech Models in Real-World Conversations

The new Speech Agent Arena evaluates speech-to-speech models and cascaded systems on real-world scenarios. It measures both conversational preference and task success, providing a view of how models perform when people use them to complete practical requests.

Existing speech-to-speech benchmarks examine reasoning, simulated agentic tasks, and conversational behavior such as turn-taking and interruption handling. The Speech Agent Arena instead has participants interact with models while completing real-world tasks, including scenarios that require tool calls.

How the Speech Agent Arena Works

Participants compare two hidden speech-to-speech models assigned to the same scenario. The evaluation includes 15 agentic scenarios, such as ordering takeout, and 20 non-agentic scenarios, such as asking about opening hours.

After completing separate live conversations with both models, participants select the one they preferred. These pairwise votes are used to calculate a Preference Elo score.

For agentic scenarios, Task Success Rate is the percentage of eligible conversations in which the model completes the requested action through the correct final tool call or calls. Conversations involving participant deviations or unverifiable outcomes are excluded from this calculation.

Except for the New Patient Dental Booking example, scenario prompts, tool schemas, and participant instructions are currently private to reduce overfitting. The evaluation uses a qualified pool of paid, screened third-party participants to conduct and assess the interactions.

Initial Results

Preference Elo

The initial Preference Elo rankings are:

  1. GoogleAI Gemini 3.1 Flash Live Preview - Minimal: 1,046 Elo
  2. Gemini 3.1 Flash Live Preview - High: 1,014 Elo
  3. OpenAI GPT-Realtime-1.5: 1,000 Elo
  4. GPT Realtime (Aug '25): 944 Elo
  5. ElevenLabs Agents: 937 Elo

The ElevenLabs Agents result uses the default cascaded system of Scribe v2 Realtime, GPT-4o Mini, and Eleven v3, with a pre-registered tool schema.

In reviewed conversations, highly preferred models generally responded quickly, sounded more natural, and produced fewer unnatural sounds or audio artifacts.

Task Success Rate

The leading Task Success Rate results are:

  1. SpaceXAI Grok Voice Think Fast 2.0 High: 94.7%
  2. OpenAI GPT-Realtime-2.1 High: 91.5%
  3. ElevenLabs Agents (Default Cascaded System): 90.5%
  4. GPT-Realtime-2 (High): 89.8%
  5. GPT Realtime (Aug '25) and GPT-Realtime-2.1 Minimal: 89.4%, tied

Gemini 3.1 Flash Live Preview - Minimal leads overall preference with 1,046 Elo, but has a Task Success Rate of 74.6%. This difference shows that conversational preference and successful task completion do not always align. A conversation may sound as though the requested action was completed even when the required final tool call was unsuccessful.

Responsiveness and Preference

Overall Preference Elo generally increases as Time to First Audio (TTFA) decreases, suggesting that responsiveness contributes to a preferred conversational experience.

Gemini 3.1 Flash Live Preview - Minimal leads overall preference with 1,046 Elo and a 0.96-second TTFA. GPT-Realtime-2 (High) records 914 Elo and a 1.14-second TTFA, while Qwen Audio 3.0 Realtime Plus records 699 Elo and a 1.54-second TTFA.

Cost, Preference, and Task Success

Models achieve different levels of conversational preference and task success at substantially different prices.

Gemini 3.1 Flash Live Preview - Minimal leads overall preference with 1,046 Elo, costs $1.50 per hour of input audio, and has a 74.6% Task Success Rate. Grok Voice Think Fast 2.0 High leads Task Success Rate at 94.7% and costs $4.80 per hour. GPT-Realtime-2.1 High follows with a 91.5% Task Success Rate and costs $10.75 per hour.

The Speech Agent Arena is continuing to expand its coverage of native and cascaded speech-to-speech systems, along with the number of models, providers, and scenarios included in the evaluation.