Metrics guide
Measure whether a voice-agent call actually worked.
Start with a small, reviewable scorecard. Aggregate metrics show where to look; the conversation, rationale, and evidence explain what to fix.
The current production scorecard
Task completion
Did the caller reach the intended customer and business outcome?
Intent understanding
Did the agent understand what the caller needed before responding or acting?
Required information
Did the workflow capture the information it needed to move forward?
Read scores with evidence
Scores are prioritization signals, not an automatic truth. Review the evaluator rationale and supporting conversation evidence, especially for uncertain or high-impact calls. Keep human review in the loop when the outcome matters.
Additional quality signals to define for your workflow
- Resolution quality: was the outcome correct, complete, and communicated clearly?
- Fallback behavior: did the agent recover safely when it could not proceed?
- Unsupported claims: did it make a statement not grounded in available policy, tools, or context?
- Operational context: latency, interruptions, silence, repeated turns, and premature termination where provider data makes them available.
VaaniEval currently starts with a focused evaluator scorecard. Teams helping shape the product can influence future custom-rubric and criteria-suggestion capabilities.
Help shape the roadmapBuild with VaaniEval
Make evaluation coverage part of how your agent ships.
We are working with production Voice AI teams that need a quality system to keep pace with releases and real call behavior.
