Automated testing for conversational AI
Simulate real customers and catch multi-turn failures, hallucinations, and bad tool calls before your customers do.
“Voxli becomes more valuable the longer you’ve used it. The tests you build become a custom dataset your competitors don’t have.”
Expertise.ai runs Voxli nightly to catch agent regressions before customers do. Built with our design partners.
Simulate multi-turn conversations at scale
Run hundreds of tests in parallel. Create impatient, confused, and combative personas to uncover failures a happy-path demo never reveals.
-
Catch hallucinations and bad tool calls
Every run flags hallucinated answers, broken tool calls, and wrong parameters, and ties the evidence to the exact turn.
-
One test, many personalities
Run the same scenario with different user personas to catch edge cases automatically.
Write tests in plain English
Describe what to test in plain language, or generate tests with Claude, ChatGPT, and more.
-
Replay production conversations
Turn a real conversation into a repeatable test and keep it as your regression suite.
-
Connect any agent
Chat, voice, or any framework. Connect in under an hour and run tests in CI or locally.
“The ability to replay real production conversations and turn them into repeatable tests solves a big and important problem for us.”
Agents break when something changes.
Start from the change you are about to make.
Swap models safely
Run the old model and the new one on the same suite and see exactly what changed.
See the use caseCatch hallucinations
Know when a change makes your agent invent facts, before a customer acts on one.
See the use caseProve launch readiness
Turn the use cases, brand rules, and safety limits you agreed on into a pre-launch suite.
See the use caseStop shipping regressions
Run the whole suite on every change, on a schedule or in CI, and get alerted when a flow breaks.
See the use caseTrust every tool call
Check that the agent calls the right tool with the right arguments and reports what it actually did.
See the use caseReproduce live issues
Import the conversation that went wrong, replay it until the failure shows itself, and keep it as a test.
See the use caseYour coding agent fixes what Voxli finds
Connect Claude Code, Cursor, or any MCP client to your Voxli workspace. Your assistant reads the failing run, applies a fix, and runs the tests again.
- Browse your scenarios and agents
- Trigger test runs from your editor
- Read detailed results, turn by turn
# Connect your workspace
$ claude mcp add voxli --transport http https://api.voxli.io/mcp
# Then ask your assistant
> Look up the latest run in Voxli and analyze why it failed.
“We had our development team spend three months building internal tools to handle what Voxli does out-of-the-box.”
Know what to fix before you deploy
Compare agent versions side by side after a model swap, prompt change, or tool update. Every test scores against your baseline.
Compare on the product page
Side-by-side comparison
Run two or more agents on the same suite and see exactly which tests moved.
Repeat runs, real signal
Run each test several times and read aggregated scores instead of a single lucky pass.
See Voxli in action
A short product walkthrough recorded by one of the founders.
When your agent speaks for a business, what it says matters.
We are early stage and founder led. Book a call and talk directly with the people building Voxli.