last updated 2026-10-01

VAmoS Pro Bench

Voice Agent Simulation Bench for utility billing and payment assistance. Every call chains two to four caller requests, over street, café or television noise, to a calm or angry caller. Every world effect and spoken figure is graded against the trace.

Leaderboard

Voice agentTask completionLatencyTurnsBarge-in/callCostTypeModel
grok voice44.7%2.64s161.07~$0.210speech‑to‑speechGrok Voice Think Fast 2.0
gemini 3.8 live40.1%1.66s120.13~$0.193speech‑to‑speechGemini 3.8 Live
gpt-live 139.2%1.62s171.05~$0.192speech‑to‑speechGPT-Live 1
openai realtime 2.137.9%2.12s140.44~$0.633speech‑to‑speechGPT Realtime 2.1
elevenlabs36.4%2.18s120.04~$0.272cascadeElevenLabs ASR → gpt-4.1-mini → eleven_flash_v2
openai realtime33.6%2.32s140.35~$0.607speech‑to‑speechGPT Realtime 2
livekit29.4%2.96s120.04~$0.077cascadeDeepgram nova-3 → gpt-4.1-mini → ElevenLabs eleven_flash_v2
gemini 3.1 live25.6%2.66s120.10~$0.199speech‑to‑speechGemini 3.1 Flash Live
mistral22.7%3.75s120.08~$0.279cascadeVoxtral realtime → mistral-medium-2604 → Voxtral TTS
gradium21.3%2.42s141.18~$0.120cascadeGradium STT → gpt-4.1-mini → Gradium TTS
vapi20.5%2.08s200.71~$0.271cascadeDeepgram → gpt-4.1-mini → ElevenLabs
deepgram19.8%2.04s130.34~$0.193cascadenova-3 → gpt-4.1-mini → aura-2
hugging face19.0%5.28s129.86~$0.053cascadeWhisper large-v3 → gpt-oss-120b → Kokoro-82M
pipecat17.3%1.96s187.94~$0.109cascadeDeepgram nova-3 → gpt-4.1-mini → ElevenLabs eleven_flash_v2

Cost vs. completed tasks

10% 20% 30% 40% 50% $0.05 $0.10 $0.20 $0.50 Mean cost per call (USD, log scale) · Y: task completion (% of calls) Pareto frontier grok voice: 44.7% completion, ~$0.210 per call 3.8 Livegemini 3.8 live: 40.1% completion, ~$0.193 per call GPT-Livegpt-live 1: 39.2% completion, ~$0.192 per call Realtime 2.1openai realtime 2.1: 37.9% completion, ~$0.633 per call elevenlabs: 36.4% completion, ~$0.272 per call Realtime 2openai realtime: 33.6% completion, ~$0.607 per call livekit: 29.4% completion, ~$0.077 per call 3.1 Livegemini 3.1 live: 25.6% completion, ~$0.199 per call mistral: 22.7% completion, ~$0.279 per call gradium: 21.3% completion, ~$0.120 per call vapi: 20.5% completion, ~$0.271 per call deepgram: 19.8% completion, ~$0.193 per call hugging face: 19.0% completion, ~$0.053 per call pipecat: 17.3% completion, ~$0.109 per call grok voice: 44.7% completion, ~$0.210 per callgrok voice gemini 3.8 live: 40.1% completion, ~$0.193 per callgemini 3.8 live gpt-live 1: 39.2% completion, ~$0.192 per callgpt-live 1 openai realtime 2.1: 37.9% completion, ~$0.633 per callopenai realtime 2.1 elevenlabs: 36.4% completion, ~$0.272 per callelevenlabs openai realtime: 33.6% completion, ~$0.607 per callopenai realtime livekit: 29.4% completion, ~$0.077 per calllivekit gemini 3.1 live: 25.6% completion, ~$0.199 per callgemini 3.1 live mistral: 22.7% completion, ~$0.279 per callmistral gradium: 21.3% completion, ~$0.120 per callgradium vapi: 20.5% completion, ~$0.271 per callvapi deepgram: 19.8% completion, ~$0.193 per calldeepgram hugging face: 19.0% completion, ~$0.053 per callhugging face pipecat: 17.3% completion, ~$0.109 per callpipecat

Build this around your own systems

This board is one exam, on one utility’s policies. Veris builds the same thing around yours: your workflows, your systems, your rules, and the calls your agents actually get. Simulated callers, twinned services, and a verdict from the world the agent leaves behind rather than the transcript it produced.

Book a demo

Running a voice stack of your own? We’ll put it through the same 100 tasks and publish the result. Contact us at hello@veris.ai