DOCS · VOICE · EVALUATION REFERENCE

Voice Agent Production Evaluation Standard

A technical evaluation reference for deciding whether a voice workflow has enough retained evidence to be treated as production-ready for its stated environment.

Evaluation reference · draft 1.0 · 2026-09-19

Evaluation contractLatencyTurn takingInterruption & cancellationTools, handoff & fallbackConcurrency & language

1. Evaluation contract

Every reported result should identify the metric definition, test method, pass/fail rule, environment and retained evidence. Repository targets and laboratory fixtures remain targets or test evidence unless a current production measurement set is explicitly attached.

2. TTFA, STT, TTS and end-to-end latency

TTFA · assistant-audio onset

Metric definition. Declare the exact timing boundary used for the test and the assistant-audio onset event counted by the harness.

Test method. Capture cold and warm paths separately, with provider, transport state, sample size and percentile.

Pass/fail contract. Set the threshold before the run; do not compare two TTFA numbers when their start or end boundaries differ.

Current verified result. The repository contains TTFA instrumentation, engineering targets and voice test infrastructure. This draft does not promote those target values into a current production benchmark.

STT behavior

Measure endpointing, finalization, transcript stability, partial-result behavior and failure handling under declared audio conditions and provider configuration.

TTS behavior

Measure audio-start timing, continuity, cancellation behavior, warm/cold transport state and failure recovery. Report these dimensions separately so one latency number does not hide an interruption or continuity failure.

End-to-end latency

Measure the complete user-turn to assistant-audio path and retain stage timing for VAD, STT, model or tool execution and TTS where the configured path exposes those events.

3. Turn taking

Test normal completion, short backchannels, premature endpointing, overlapping speech and ambiguous silence. Each fixture should declare the expected listening or speaking state and the failure that would make the turn unstable.

4. Interruption, barge-in and cancellation

Interruption / barge-in

Start user speech while assistant audio is active, then measure detection, cancellation propagation and the time until assistant audio becomes silent.

Cancellation

The current native voice_core includes a cancellation token with deterministic lifecycle tests. A production evaluation still needs to verify that every downstream generation and audio path used by the configured deployment respects cancellation.

Stale-response prevention

After a turn is invalidated, delayed audio or tool output from that turn must not reappear. The production fixture should deliberately create delayed work after cancellation and fail if stale output becomes visible or audible.

5. Tool execution, handoff and fallback

Tool execution

Measure tool-call correctness, policy behavior, added latency, error handling and whether the spoken response reflects the verified tool result rather than a proposed action.

Handoff

Test the transfer trigger, context package, destination behavior and caller-facing confirmation. A spoken transfer message is not itself proof that the receiving path accepted the handoff.

Fallback

Declare expected behavior when STT, model, TTS, telephony or a required tool fails. The result should distinguish a recovered conversation from an unresolved termination.

6. Concurrency, language and code-switch behavior

Concurrency

Test the configured deployment under increasing simultaneous sessions and report saturation behavior separately from single-session latency. Concurrency targets in planning documents are not treated here as measured capacity.

Language / code-switch behavior

UNKNOWN / NOT CERTIFIED in this reference. Model capability alone is not enough. Certification requires a retained language mix, audio fixture set, scoring method and pass criteria.

7. What this standard does not claim

8. Evaluate a real voice workflow

Use the commercial Voice page to define the workflow and authority boundary, then attach a test contract for the configured call path.

Review Voice Agents →Action-level governanceTools & MCPEvaluate a voice workflow