We put decision models in charge of when a voice agent is allowed to talk. We ran 300 simulated phone calls with each of eight models on VAmoS Pro Bench, our voice-agent benchmark. We compared them with a baseline agent, which ends the caller's turn after 0.8 s of silence.
The baseline talked over half of all caller turns. With any of the eight models, that improved to between 5% and 21%. All but one model also made the agent measurably slower to reply. In general, the more often a model waited, the less the agent talked over callers.
Each dot is one model, or the baseline, over about 300 calls. Lower means less talking over the caller; further left means faster replies.
Jev started a family of what its makers call decision models. You send some text and a decision to make, with its options (yes or no, one of several, or a score), and you get back a probability for each option rather than generated text. We tested eight of them, each through a hosted API:
jev-1.13.0), from TypeSafe.nimble-latest), a 9B model from Bespoke Labs.pplx-decider-v1-27b, pplx-decider-v1.1-27b).@cf/cloudflare/clef, @cf/cloudflare/clef-flash), on Cloudflare Workers AI.fastino/GLiDE), from Fastino, which reasons further on decisions it is unsure about.jaredpalmer/kev-4b), an open-source model, through OpenRouter.The agent is the Pipecat agent from VAmoS Pro Bench and the Jev post, with the same prompt, speech recognizer, language model, and voice. In the baseline, 0.8 s of silence ends the caller's turn. With a model, the model decides whether that silence ends the turn. If it decides the caller isn't finished, the agent waits. It ends the turn once the caller has been silent for about 5 s, or at the first pause after 10 s of waiting. If the model errors or hasn't decided within 1 s, the turn ends. If the caller, or a TV, starts talking before the model decides, the agent asks for a new decision at the next pause.
Every model had the same decision to make, with these instructions and options:
Instructions:
The user is talking to a voice assistant. Their words come from speech recognition, without punctuation, and may have been cut off. Decide whether the user's turn is complete. Complete means conversationally complete, not long: one word can be a complete answer, a question is complete, a correction is complete.
Options:
complete: the user has taken their turn and the assistant should answer
short: the user stopped mid-sentence and will continue in a few seconds: the last words leave a phrase open, such as ending on a conjunction, a preposition, an article, or an unfinished list or number
long: the user needs time to think or asked the assistant to wait, or has only acknowledged the question without answering it
Each model also had the same two simulated callers and the same mix of background noise: a quiet room, a street, a café, or a TV, a quarter of the calls each.
Waited: of the times the agent asked the model for a decision, the share where the agent kept waiting, because the model decided the caller wasn't finished, or more speech came in before it decided. About two thirds of those requests came from the TV calls. Completed: calls that passed all three of the task's checks, as graded by an LLM verifier.
Nimble and Clef Flash waited most often and talked over callers least. The two Perplexity models and Kev-4B waited least often and talked over callers most. Jev is the one model out of order: it waited about as often as Clef but talked over callers more.
A model can make the agent wait in two ways: by deciding the caller isn't finished, or by deciding so slowly that the caller starts talking again first. Nimble and Clef Flash mostly decided to wait, most of all in the TV calls. GLiDE and Kev-4B decided to wait far less often, but they decided so slowly that callers often started talking again first.
No single model made a statistically clear difference to how many calls the agent completed.
Share of caller turns the agent talked over, by background noise, about 75 calls per model in each condition. Each model's calls got their noise in a separate random draw. Bars keep the same order in every group; hover a bar for its value.
Without the TV, every model kept talk-over between 4% and 10% of caller turns. With the TV on, the models spread out from 3% to 31%. The speech recognizer wrote the TV into the caller's words, so in TV calls about half the words the model saw were never said by the caller. Nimble and Clef Flash waited through most of those turns, often until the 10 s limit. The agent rarely talked over callers, but took 6 to 7 s to reply. With the TV on, the agent completed almost no calls, with any model or the baseline.
This is one agent, one speech recognizer, and one flavor of background noise. A different prompt, or thresholds tuned on the models' probabilities, could change where each model lands, and the 10 s waiting limit shapes the TV results. Decision times include the network trip from our cluster to each provider. GLiDE's no-thinking version was not available over Fastino's API, so it is not in this test.
We got these answers from 2,694 simulated calls. Veris builds the same kind of simulation around your callers, policies, and noise, so you can see what a new component does to your calls before your customers hear it.
A new turn detector, model, or prompt can change how every call goes. Veris runs it against simulated callers and twins of your systems, and grades what the agent did.
Book a demo