last updated 2026-07-28

VAmoS Bench

Voice Agent Simulation Bench (VAmoS) runs a credit card support agent through simulated phone calls and judges them on end-to-end task completion.

Leaderboard

Voice agentTask completionConnectLatencyTurnsBarge-in/callCostSTTLLMTTS
pipecat71.0%100.0%1.95s90.17~$0.045Deepgram nova-3-generalgpt-4.1-miniElevenLabs
livekit70.3%100.0%2.37s70.07~$0.048Deepgram nova-3-generalgpt-4.1-miniElevenLabs
vapi69.3%97.7%2.25s130.54~$0.135Deepgram nova-3gpt-4.1-miniElevenLabs
elevenlabs67.3%99.7%1.19s70.27~$0.114Scribegpt-4.1-minieleven_flash_v2
cartesia67.0%100.0%2.23s70.02~$0.146Inkgpt-4.1-miniSonic
openai realtime67.0%100.0%1.53s70.15~$0.074gpt-realtime-2
gemini 2.5 native64.0%99.3%15.95s70.01~$0.023Gemini 2.5 Native Audio
gemini 3.1 live62.3%99.7%1.38s70.02~$0.016Gemini 3.1 Flash Live
retell61.3%100.0%5.76s91.26~$0.208Vendor ASRgpt-4.1-miniElevenLabs
openai realtime mini51.0%96.3%1.79s90.56~$0.032gpt-realtime-2.1-mini
nemotron*43.0%89.0%2.12s70.10n/aNemotronNemotron 3 NanoMagpie

Cost vs. completed tasks

45% 55% 65% 75% $0.00 $0.05 $0.10 $0.15 $0.20 Estimated cost per call (USD) · Y: task completion (% of calls) pipecat livekit vapi elevenlabs cartesia openai realtime gemini 2.5 native gemini 3.1 live retell openai realtime mini

Submit an implementation

This is a living benchmark. Every row is one implementation of Riley, the card-ops agent. Build your own on any framework, platform, or model, including ones already on the board. We’ll run it through the same 100 scenarios and publish the results. Feedback on the methodology is just as welcome.

Contact us