Back to VAmoS Pro Bench
October 8, 2026

Decision models for voice-agents: a benchmark

Joshua Meyer

We put decision models in charge of when a voice agent is allowed to talk. We ran 300 simulated phone calls with each of eight models on VAmoS Pro Bench, our voice-agent benchmark. We compared them with a baseline agent, which ends the caller's turn after 0.8 s of silence.

The baseline talked over half of all caller turns. With any of the eight models, that improved to between 5% and 21%. All but one model also made the agent measurably slower to reply. In general, the more often a model waited, the less the agent talked over callers.

0% 10% 20% 30% 40% 50% 60% 2.5 s 3 s 3.5 s 4 s 4.5 s 5 s Median reply time (caller finishes → agent speaks) Caller turns talked over ↙ better: fewer interruptions, faster replies Baseline (0.8 s silence): 50.5% talked over, 2.66 s Baseline Jev: 12.1% talked over, 3.28 s Jev Nimble: 4.6% talked over, 4.50 s Nimble Kev-4B: 16.9% talked over, 3.70 s Kev-4B Perplexity Decider v1: 20.9% talked over, 3.16 s Perplexity v1 Perplexity Decider v1.1: 17.9% talked over, 2.90 s Perplexity v1.1 Clef: 8.4% talked over, 3.21 s Clef Clef Flash: 4.6% talked over, 3.41 s Clef Flash GLiDE: 8.7% talked over, 3.40 s GLiDE

Each dot is one model, or the baseline, over about 300 calls. Lower means less talking over the caller; further left means faster replies.

The models

Jev started a family of what its makers call decision models. You send some text and a decision to make, with its options (yes or no, one of several, or a score), and you get back a probability for each option rather than generated text. We tested eight of them, each through a hosted API:

  • Jev (jev-1.13.0), from TypeSafe.
  • Nimble (nimble-latest), a 9B model from Bespoke Labs.
  • Perplexity Decider v1 and v1.1 (pplx-decider-v1-27b, pplx-decider-v1.1-27b).
  • Clef and Clef Flash (@cf/cloudflare/clef, @cf/cloudflare/clef-flash), on Cloudflare Workers AI.
  • GLiDE (fastino/GLiDE), from Fastino, which reasons further on decisions it is unsure about.
  • Kev-4B (jaredpalmer/kev-4b), an open-source model, through OpenRouter.

The setup

The agent is the Pipecat agent from VAmoS Pro Bench and the Jev post, with the same prompt, speech recognizer, language model, and voice. In the baseline, 0.8 s of silence ends the caller's turn. With a model, the model decides whether that silence ends the turn. If it decides the caller isn't finished, the agent waits. It ends the turn once the caller has been silent for about 5 s, or at the first pause after 10 s of waiting. If the model errors or hasn't decided within 1 s, the turn ends. If the caller, or a TV, starts talking before the model decides, the agent asks for a new decision at the next pause.

Every model had the same decision to make, with these instructions and options:

Instructions:
The user is talking to a voice assistant. Their words come from speech recognition, without punctuation, and may have been cut off. Decide whether the user's turn is complete. Complete means conversationally complete, not long: one word can be a complete answer, a question is complete, a correction is complete.

Options:
complete: the user has taken their turn and the assistant should answer
short: the user stopped mid-sentence and will continue in a few seconds: the last words leave a phrase open, such as ending on a conjunction, a preposition, an article, or an unfinished list or number
long: the user needs time to think or asked the assistant to wait, or has only acknowledged the question without answering it

Each model also had the same two simulated callers and the same mix of background noise: a quiet room, a street, a café, or a TV, a quarter of the calls each.

Waiting more, talking over less

Baseline Nimble Clef Flash Jev Clef GLiDE Kev-4B Perplexity v1.1 Perplexity v1 Waited – Nimble, waited: 73%73% Clef Flash, waited: 68%68% Jev, waited: 53%53% Clef, waited: 51%51% GLiDE, waited: 46%46% Kev-4B, waited: 39%39% Perplexity Decider v1.1, waited: 37%37% Perplexity Decider v1, waited: 37%37% Talked over Baseline (0.8 s silence), talked over: 50.5%50.5% Nimble, talked over: 4.6%4.6% Clef Flash, talked over: 4.6%4.6% Jev, talked over: 12.1%12.1% Clef, talked over: 8.4%8.4% GLiDE, talked over: 8.7%8.7% Kev-4B, talked over: 16.9%16.9% Perplexity Decider v1.1, talked over: 17.9%17.9% Perplexity Decider v1, talked over: 20.9%20.9% Reply time Baseline (0.8 s silence), reply time: 2.7 s2.7 s Nimble, reply time: 4.5 s4.5 s Clef Flash, reply time: 3.4 s3.4 s Jev, reply time: 3.3 s3.3 s Clef, reply time: 3.2 s3.2 s GLiDE, reply time: 3.4 s3.4 s Kev-4B, reply time: 3.7 s3.7 s Perplexity Decider v1.1, reply time: 2.9 s2.9 s Perplexity Decider v1, reply time: 3.2 s3.2 s Completed Baseline (0.8 s silence), completed: 17%17% Nimble, completed: 21%21% Clef Flash, completed: 22%22% Jev, completed: 23%23% Clef, completed: 24%24% GLiDE, completed: 22%22% Kev-4B, completed: 25%25% Perplexity Decider v1.1, completed: 23%23% Perplexity Decider v1, completed: 22%22%

Waited: of the times the agent asked the model for a decision, the share where the agent kept waiting, because the model decided the caller wasn't finished, or more speech came in before it decided. About two thirds of those requests came from the TV calls. Completed: calls that passed all three of the task's checks, as graded by an LLM verifier.

Nimble and Clef Flash waited most often and talked over callers least. The two Perplexity models and Kev-4B waited least often and talked over callers most. Jev is the one model out of order: it waited about as often as Clef but talked over callers more.

A model can make the agent wait in two ways: by deciding the caller isn't finished, or by deciding so slowly that the caller starts talking again first. Nimble and Clef Flash mostly decided to wait, most of all in the TV calls. GLiDE and Kev-4B decided to wait far less often, but they decided so slowly that callers often started talking again first.

No single model made a statistically clear difference to how many calls the agent completed.

The TV calls separate the models

0% 20% 40% 60% 80% Baseline (0.8 s silence), Quiet: 33% Perplexity Decider v1, Quiet: 8% Perplexity Decider v1.1, Quiet: 10% Kev-4B, Quiet: 5% Jev, Quiet: 9% GLiDE, Quiet: 5% Clef, Quiet: 7% Nimble, Quiet: 5% Clef Flash, Quiet: 4% Quiet Baseline (0.8 s silence), Street: 33% Perplexity Decider v1, Street: 12% Perplexity Decider v1.1, Street: 8% Kev-4B, Street: 7% Jev, Street: 8% GLiDE, Street: 4% Clef, Street: 6% Nimble, Street: 5% Clef Flash, Street: 5% Street Baseline (0.8 s silence), Café: 29% Perplexity Decider v1, Café: 10% Perplexity Decider v1.1, Café: 9% Kev-4B, Café: 6% Jev, Café: 9% GLiDE, Café: 3% Clef, Café: 5% Nimble, Café: 4% Clef Flash, Café: 7% Café Baseline (0.8 s silence), TV: 63% Perplexity Decider v1, TV: 31% Perplexity Decider v1.1, TV: 27% Kev-4B, TV: 27% Jev, TV: 16% GLiDE, TV: 14% Clef, TV: 11% Nimble, TV: 4% Clef Flash, TV: 3% TV Baseline (0.8 s silence), All calls: 50% Perplexity Decider v1, All calls: 21% Perplexity Decider v1.1, All calls: 18% Kev-4B, All calls: 17% Jev, All calls: 12% GLiDE, All calls: 9% Clef, All calls: 8% Nimble, All calls: 5% Clef Flash, All calls: 5% All calls Baseline (0.8 s silence) Perplexity Decider v1 Perplexity Decider v1.1 Kev-4B Jev GLiDE Clef Nimble Clef Flash

Share of caller turns the agent talked over, by background noise, about 75 calls per model in each condition. Each model's calls got their noise in a separate random draw. Bars keep the same order in every group; hover a bar for its value.

Without the TV, every model kept talk-over between 4% and 10% of caller turns. With the TV on, the models spread out from 3% to 31%. The speech recognizer wrote the TV into the caller's words, so in TV calls about half the words the model saw were never said by the caller. Nimble and Clef Flash waited through most of those turns, often until the 10 s limit. The agent rarely talked over callers, but took 6 to 7 s to reply. With the TV on, the agent completed almost no calls, with any model or the baseline.

What this does not show

This is one agent, one speech recognizer, and one flavor of background noise. A different prompt, or thresholds tuned on the models' probabilities, could change where each model lands, and the 10 s waiting limit shapes the TV results. Decision times include the network trip from our cluster to each provider. GLiDE's no-thinking version was not available over Fastino's API, so it is not in this test.

Run it on your own calls

We got these answers from 2,694 simulated calls. Veris builds the same kind of simulation around your callers, policies, and noise, so you can see what a new component does to your calls before your customers hear it.

Test it before your callers do

A new turn detector, model, or prompt can change how every call goes. Veris runs it against simulated callers and twins of your systems, and grades what the agent did.

Book a demo