
Veris AI's simulated users score 84.8% on a five-check human-likeness rubric against τ²-bench-retail's 68.6%, a 16.3-point lead, ahead on all five checks and all three behavioural dimensions, judged by the same model on byte-identical prompts. The rubric, the method, the numbers, and where the rubric itself turned out to be wrong.

Six models, one production telecom assistant, three scenario sets. Three different winners, and the cheapest model in the pool won one of them outright.

Building an agentic self-improvement loop is easy: run → diagnose → fix → repeat. The scarce ingredient isn't the loop, it's a supply of real, safe, step-level failures to feed it. Here's one turn of a loop on a real agent: a 64% token-bill cut and two confidentiality leaks stopped before they left the building.

Veris comes up with realistic scenarios you wouldn't, runs them in parallel, and surfaces the failures you don't know to look for.

A practical guide to shipping enterprise agents that work: using AWS Bedrock AgentCore to build and deploy, and Veris AI to simulate, grade, and improve.

Our simulation sandbox that's been powering enterprise AI agents for a year is now fully self-serve.


We built SantaBench, a fun benchmark with a serious methodology. The task: play a cheeky Santa who researches users online and roasts them based on their social media presence. Adding a little sneer to your holiday cheer.
Turn production failures into improvements. See how Veris AI uses simulation, targeted evaluation, and automated prompt optimization to help agents self-improve without regressions.

See how reinforcement fine-tuning transforms small models into high-accuracy Sigma rule generators—delivering enterprise-grade detection at lower cost and latency.

Deep dive into how Veris AI uses reinforcement fine-tuning and simulation to train small models to generate accurate Sigma detection rules for enterprise security.


Agents live in messy systems with unpredictable users, interconnected tools, and cascading consequences. Unlike traditional software where inputs and outputs are fully controlled, an agent’s flexible I/O and its reasoning capability are both its superpower and its failure mode.



We're thrilled to share that Veris AI is launching with $8.5M in seed funding from an oversubscribed round led by Decibel Ventures and Acrew Capital.
