This is an extension of VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). We took one of its fourteen agents, the Pipecat one, let Jev decide when the caller had finished talking, and reran the same 300 calls. The agent cut its callers off on 52% of their turns before, and 11% after.
Share of the caller's turns that the agent started speaking over, on VAmoS Pro Bench: the same 100 tasks × 3 runs per agent, 75 calls per noise condition. The drop is significant in every condition (all p < 10−7).
If you build agents that work in text, you have never had to decide when the user has finished talking. They press Enter. A voice agent gets no Enter key. It hears a continuous stream of audio, and at any moment it has to decide whether it should talk. Decide too early and it talks over people. Decide too late and it leaves them in silence.
That decision is called turn detection, and most agent developers have never had to think about it. It is also a good job for Jev, the fast classifier model from TypeSafe that people have been talking about.
Most voice agents decide with a timer. A small acoustic model, called voice activity detection, marks whether anyone is speaking. When nobody has spoken for a set stretch, the turn is over and the agent answers.
A timer only hears silence. Callers pause in the middle of things: reading an account number in chunks ("four four seven… two one nine"), or saying "hang on, let me find the bill". The timer ends the turn, the agent answers half a sentence, the caller carries on, and now the two are talking over each other. On a phone call that snowballs into repeats, misheard numbers, and a caller who has to start again.
Background noise makes it worse. A television in the room is speech too, so the speech recognizer writes it down as if the caller said it.
Semantic turn detection looks at what the caller said as well as whether they are still making sound. "My account number is four four seven" is an unfinished thought; "that's everything, thanks" is a finished one. Several voice frameworks already ship a version of it.
Jev makes it a drop-in. It is a fast, general classifier: give it some text and a question, and in about 150 ms it returns a label. For turn detection, the text is the caller's words so far plus the agent's last line, and the labels are complete, short (cut off mid-phrase), and long (paused for something longer, like finding a bill).
Whether Jev helps depends on everything around it: the speech recognizer, the noise your callers have, how they pause, and how long the agent will hold a turn open. You find out the same way you choose a model or an architecture: with a benchmark.
We took the Pipecat agent from VAmoS Pro Bench and ran it twice through the same 100 tasks, three runs each, with the same simulated callers and the same noise: once with the plain timer, and once with Jev deciding. Nothing else changed. The benchmark grades each call on what the agent changed in the billing and loan systems behind it.
We tested one agent in one configuration, with hold times we chose. A fair benchmark of turn detectors would run all of them through the same calls, so these results say nothing about Jev in general. A different agent could land somewhere else.
We got these answers from 600 simulated calls. Veris builds the same kind of simulation around your callers, policies, and noise, so you can see what a new component does to your calls before your customers hear it.
A new turn detector, model, or prompt can change how every call goes. Veris runs it against simulated callers and twins of your systems, and grades what the agent did.
Book a demo