DOCS · VOICE · EVALUATION REFERENCE
Voice Agent Production Evaluation Standard
A technical evaluation reference for deciding whether a voice workflow has enough retained evidence to be treated as production-ready for its stated environment.
Evaluation reference · draft 1.0 · 2026-09-19
Evaluation contractLatencyTurn takingInterruption & cancellationTools, handoff & fallbackConcurrency & language1. Evaluation contract
Every reported result should identify the metric definition, test method, pass/fail rule, environment and retained evidence. Repository targets and laboratory fixtures remain targets or test evidence unless a current production measurement set is explicitly attached.
- VERIFIED. Retained current evidence supports the stated test in the stated environment.
- TARGET. A desired threshold or engineering objective, not a production-result claim.
- UNKNOWN / NOT CERTIFIED. This reference is not asserting current retained evidence for that dimension.
2. TTFA, STT, TTS and end-to-end latency
TTFA · assistant-audio onset
Metric definition. Declare the exact timing boundary used for the test and the assistant-audio onset event counted by the harness.
Test method. Capture cold and warm paths separately, with provider, transport state, sample size and percentile.
Pass/fail contract. Set the threshold before the run; do not compare two TTFA numbers when their start or end boundaries differ.
Current verified result. The repository contains TTFA instrumentation, engineering targets and voice test infrastructure. This draft does not promote those target values into a current production benchmark.
STT behavior
Measure endpointing, finalization, transcript stability, partial-result behavior and failure handling under declared audio conditions and provider configuration.
TTS behavior
Measure audio-start timing, continuity, cancellation behavior, warm/cold transport state and failure recovery. Report these dimensions separately so one latency number does not hide an interruption or continuity failure.
End-to-end latency
Measure the complete user-turn to assistant-audio path and retain stage timing for VAD, STT, model or tool execution and TTS where the configured path exposes those events.
3. Turn taking
Test normal completion, short backchannels, premature endpointing, overlapping speech and ambiguous silence. Each fixture should declare the expected listening or speaking state and the failure that would make the turn unstable.
4. Interruption, barge-in and cancellation
Interruption / barge-in
Start user speech while assistant audio is active, then measure detection, cancellation propagation and the time until assistant audio becomes silent.
Cancellation
The current native voice_core includes a cancellation token with deterministic lifecycle tests. A production evaluation still needs to verify that every downstream generation and audio path used by the configured deployment respects cancellation.
Stale-response prevention
After a turn is invalidated, delayed audio or tool output from that turn must not reappear. The production fixture should deliberately create delayed work after cancellation and fail if stale output becomes visible or audible.
5. Tool execution, handoff and fallback
Tool execution
Measure tool-call correctness, policy behavior, added latency, error handling and whether the spoken response reflects the verified tool result rather than a proposed action.
Handoff
Test the transfer trigger, context package, destination behavior and caller-facing confirmation. A spoken transfer message is not itself proof that the receiving path accepted the handoff.
Fallback
Declare expected behavior when STT, model, TTS, telephony or a required tool fails. The result should distinguish a recovered conversation from an unresolved termination.
6. Concurrency, language and code-switch behavior
Concurrency
Test the configured deployment under increasing simultaneous sessions and report saturation behavior separately from single-session latency. Concurrency targets in planning documents are not treated here as measured capacity.
Language / code-switch behavior
UNKNOWN / NOT CERTIFIED in this reference. Model capability alone is not enough. Certification requires a retained language mix, audio fixture set, scoring method and pass criteria.
7. What this standard does not claim
- It does not turn engineering targets into customer-facing guarantees.
- It does not certify every telephony provider, language, deployment or concurrency level.
- It does not replace field monitoring after deployment.
8. Evaluate a real voice workflow
Use the commercial Voice page to define the workflow and authority boundary, then attach a test contract for the configured call path.
Review Voice Agents →Action-level governanceTools & MCPEvaluate a voice workflow