Evaluation & testing
Coval
Simulation and evaluation platform for voice and chat agents. It generates large volumes of simulated callers with accents, interruptions, noise and policy traps, scores production calls in real time on resolution and safety, detects regressions after prompt or model changes, and routes failures to human reviewers.
commercial · generally available · Research snapshot 2026-09-06
Visit the official product source ↗Where it fits
Evaluation & testing · Observability & traceability
Useful conversation with: Head of AI product, Contact centre operations lead, QA lead.
Ask for a demonstration
Demonstrate a pre-launch simulation of a thousand callers against my voice agent, then show the production scoring that flags a regression after a prompt change.
Capabilities and evidence
Support labels reflect the supplied research. Documentation and vendor claims are not independent product tests. “Not established” means the researcher did not find support; it does not prove a capability is absent.
Documented by provider
Coval documents simulating thousands of realistic conversations before launch against voice or chat agents, including inbound, outbound and voice-to-voice connections, with edge cases, background noise, accents and IVRs.
Limit: Docs do not describe capture of internal agent spans or tool-call traces beyond conversation-level testing.
Source s1
Documented by provider
The platform scores every production call in real time, standardises metrics, surfaces regressions and can route failures to human reviewers whose feedback returns to the evaluation loop.
Limit: No documented retention period, immutable record or access-control model for stored call evaluations.
Source s1
Vendor claim
The product page states voice agent evaluation considers timing, turn-taking, interruptions, audio issues, tool calls, caller emotion and the final transcript, and that the same scenarios can be run across competing voice AI vendors.
Limit: Marketing page; per-dimension scoring methodology is not documented.
Source s2
Limitations to discuss
- No documented audit trail, RBAC, retention policy or PII redaction for call transcripts, which matters for regulated contact-centre use.
- Evidence is limited to voice/chat conversational agents; no coverage of back-office or coding agents.
Sources
- Welcome to Coval - Coval Documentation · Coval · official docs
Access date reported by researcher: 2026-09-06 - Coval: Voice AI Testing & Evaluation Platform · Coval · official product
Access date reported by researcher: 2026-09-06
Listing does not imply partnership, supplier status, a working DutyGraph integration, or a compliance certification.
Suggest a correction