VAmoS Pro Bench is live. It is our second public voice-agent benchmark, and it is a much harder call than the first. See the full board, methodology and call recordings →
The agent is Rory, on the billing line of a regulated utility, where almost every call is one of two conversations: what is this bill and I cannot pay it. Each of the 100 tasks chains two to four requests into a single call. The caller is calm or angry. The line is clean, or there is street noise, café chatter, or a television playing in the room.
Behind the agent are sixteen tools, a Stripe billing twin and the real Apache Fineract loan engine, which enforce their own rules. Nothing is disclosed until the caller is verified. The account usage comes from NREL's public model of Pennsylvania homes, and the policy follows Pennsylvania's residential billing regulation.
The grade comes from what the agent left behind, not from what it said: the exact world effects, in order, with the same numbers, plus every figure it had to say out loud. Fourteen stacks, three runs each, 4,200 calls.
| Voice agent | Task completion | Latency | Cost per call |
|---|---|---|---|
| grok voice | 44.7% | 2.64s | ~$0.210 |
| gemini 3.8 live | 40.1% | 1.66s | ~$0.193 |
| gpt-live 1 | 39.2% | 1.62s | ~$0.192 |
| openai realtime 2.1 | 37.9% | 2.12s | ~$0.633 |
| elevenlabs | 36.4% | 2.18s | ~$0.272 |
| openai realtime | 33.6% | 2.32s | ~$0.607 |
| livekit | 29.4% | 2.96s | ~$0.077 |
| gemini 3.1 live | 25.6% | 2.66s | ~$0.199 |
| mistral | 22.7% | 3.75s | ~$0.279 |
| gradium | 21.3% | 2.42s | ~$0.120 |
| vapi | 20.5% | 2.08s | ~$0.271 |
| deepgram | 19.8% | 2.04s | ~$0.193 |
| hugging face | 19.0% | 5.28s | ~$0.053 |
| pipecat | 17.3% | 1.96s | ~$0.109 |
Completion over counted attempts, after removing calls the benchmark itself spoiled. Full columns, intervals and definitions are on the board.
Cost is logarithmic. The line runs through the stacks no other stack beats on both cost and completion.
Three things stand out.
Nobody is close to finished. The best stack completed 44.7% of calls; the field runs down to 17.3%. On VAmoS Bench, our card-support benchmark, the best completed 71%.
Verification is solved; acting is not. Agents pass the verification gate on 99.0% of graded calls, then produce the right world effects on 54.8%. On 44% of calls the agent verified the caller correctly and then wrote the wrong thing.
A television is the hardest condition by a wide margin. Pooled completion falls from 38.7% on a clean line to 8.6% with a TV in the room. Turn the noise up to the level of the caller's own voice and ten of the fourteen stacks differ significantly from their clean calls.
Every implementation is on GitHub; each one is only a transport around a shared prompt, the same sixteen tools and the same verification gate. The board carries a recording of one failed call per stack, with the line the agent said and what its trace shows it actually did.
Read the full benchmark →
Read the paper on arXiv →
This board is one exam, on one utility's policies. Veris builds the same thing around yours: your workflows, your systems, your rules, and the calls your agents actually get. Simulated callers, twinned services, and a verdict taken from the world the agent leaves behind rather than the transcript it produced.
Build your own benchmark