Back to blogs
October 1, 2026

VAmoS Pro Bench: a harder exam for voice agents, and what fourteen stacks did with it

Veris AI

VAmoS Pro Bench is live. It is our second public voice-agent benchmark, and it is a much harder call than the first. See the full board, methodology and call recordings →

What a VAmoS Pro Bench call sounds like, and how it is graded.

What it asks of an agent

The agent is Rory, on the billing line of a regulated utility, where almost every call is one of two conversations: what is this bill and I cannot pay it. Each of the 100 tasks chains two to four requests into a single call. The caller is calm or angry. The line is clean, or there is street noise, café chatter, or a television playing in the room.

Behind the agent are sixteen tools, a Stripe billing twin and the real Apache Fineract loan engine, which enforce their own rules. Nothing is disclosed until the caller is verified. The account usage comes from NREL's public model of Pennsylvania homes, and the policy follows Pennsylvania's residential billing regulation.

The grade comes from what the agent left behind, not from what it said: the exact world effects, in order, with the same numbers, plus every figure it had to say out loud. Fourteen stacks, three runs each, 4,200 calls.

The results

Voice agentTask completionLatencyCost per call
grok voice44.7%2.64s~$0.210
gemini 3.8 live40.1%1.66s~$0.193
gpt-live 139.2%1.62s~$0.192
openai realtime 2.137.9%2.12s~$0.633
elevenlabs36.4%2.18s~$0.272
openai realtime33.6%2.32s~$0.607
livekit29.4%2.96s~$0.077
gemini 3.1 live25.6%2.66s~$0.199
mistral22.7%3.75s~$0.279
gradium21.3%2.42s~$0.120
vapi20.5%2.08s~$0.271
deepgram19.8%2.04s~$0.193
hugging face19.0%5.28s~$0.053
pipecat17.3%1.96s~$0.109

Completion over counted attempts, after removing calls the benchmark itself spoiled. Full columns, intervals and definitions are on the board.

10% 20% 30% 40% 50% $0.05 $0.10 $0.20 $0.50 Mean cost per call (USD, log scale) · Y: task completion (% of calls) Pareto frontier grok voice: 44.7% completion, ~$0.210 per call 3.8 Livegemini 3.8 live: 40.1% completion, ~$0.193 per call GPT-Livegpt-live 1: 39.2% completion, ~$0.192 per call Realtime 2.1openai realtime 2.1: 37.9% completion, ~$0.633 per call elevenlabs: 36.4% completion, ~$0.272 per call Realtime 2openai realtime: 33.6% completion, ~$0.607 per call livekit: 29.4% completion, ~$0.077 per call 3.1 Livegemini 3.1 live: 25.6% completion, ~$0.199 per call mistral: 22.7% completion, ~$0.279 per call gradium: 21.3% completion, ~$0.120 per call vapi: 20.5% completion, ~$0.271 per call deepgram: 19.8% completion, ~$0.193 per call hugging face: 19.0% completion, ~$0.053 per call pipecat: 17.3% completion, ~$0.109 per call grok voice: 44.7% completion, ~$0.210 per callgrok voice gemini 3.8 live: 40.1% completion, ~$0.193 per callgemini 3.8 live gpt-live 1: 39.2% completion, ~$0.192 per callgpt-live 1 openai realtime 2.1: 37.9% completion, ~$0.633 per callopenai realtime 2.1 elevenlabs: 36.4% completion, ~$0.272 per callelevenlabs openai realtime: 33.6% completion, ~$0.607 per callopenai realtime livekit: 29.4% completion, ~$0.077 per calllivekit gemini 3.1 live: 25.6% completion, ~$0.199 per callgemini 3.1 live mistral: 22.7% completion, ~$0.279 per callmistral gradium: 21.3% completion, ~$0.120 per callgradium vapi: 20.5% completion, ~$0.271 per callvapi deepgram: 19.8% completion, ~$0.193 per calldeepgram hugging face: 19.0% completion, ~$0.053 per callhugging face pipecat: 17.3% completion, ~$0.109 per callpipecat

Cost is logarithmic. The line runs through the stacks no other stack beats on both cost and completion.

Three things stand out.

Nobody is close to finished. The best stack completed 44.7% of calls; the field runs down to 17.3%. On VAmoS Bench, our card-support benchmark, the best completed 71%.

Verification is solved; acting is not. Agents pass the verification gate on 99.0% of graded calls, then produce the right world effects on 54.8%. On 44% of calls the agent verified the caller correctly and then wrote the wrong thing.

A television is the hardest condition by a wide margin. Pooled completion falls from 38.7% on a clean line to 8.6% with a TV in the room. Turn the noise up to the level of the caller's own voice and ten of the fourteen stacks differ significantly from their clean calls.

The agents are public, and so is what they got wrong

Every implementation is on GitHub; each one is only a transport around a shared prompt, the same sixteen tools and the same verification gate. The board carries a recording of one failed call per stack, with the line the agent said and what its trace shows it actually did.

Read the full benchmark →
Read the paper on arXiv →

Build this around your own systems

This board is one exam, on one utility's policies. Veris builds the same thing around yours: your workflows, your systems, your rules, and the calls your agents actually get. Simulated callers, twinned services, and a verdict taken from the world the agent leaves behind rather than the transcript it produced.

Build your own benchmark