Voice-agent evaluation
Your agent changed. Which tests are now wrong?
VaaniEval connects production behavior to evaluation coverage, so your team can identify what is still representative, what may be stale, and what needs review.
A passing suite can still be the wrong suite.
A release can alter a prompt, model, tool, policy, or fallback path without turning an existing test red. Production calls can then diverge from the behavior your suite was written to protect.
Track what moved
Start with an agent release and the production behavior that deserves investigation.
Classify relevance
Review whether a test is still representative, possibly stale, missing, or redundant.
Decide with context
Use the transcript, available audio, trace, and outcome to decide what belongs in the suite.
VaaniEval is being built as a reviewable path from changed behavior to a coverage decision. The team responsible for the agent remains in control of the change.
Explore the evaluation scorecardBuild with VaaniEval
Make evaluation coverage part of how your agent ships.
We are working with production Voice AI teams that need a quality system to keep pace with releases and real call behavior.
