The GenAI Field Guide

Voice and conversational agents

Build real-time conversation that can listen, respond, and act safely.

A voice agent is a conversational AI system that people talk to out loud, on a phone line, in an app or through a device, and that can answer questions and take actions while the conversation is happening. It uses the same language models, tools and guardrails as a text agent, but the medium changes almost every engineering decision. Speech arrives as a continuous audio stream rather than a finished message, nobody presses Enter, people talk over the agent, and a pause of two seconds feels broken. Anything the agent says is heard once and cannot be scrolled back, so a wrong date or a long list costs far more than it does on a screen.

The useful mental model is a real-time loop with a latency budget. Audio comes in, a detector decides whether someone is speaking and whether they have finished, speech becomes meaning (through a transcription step or directly inside an audio model), the model decides what to say or which tool to call, and synthesized speech goes back out while the system keeps listening for interruptions. Each stage spends part of a budget of roughly one second between the caller stopping and the agent starting to speak. Most quality problems in voice are timing problems: cutting people off, waiting too long, or continuing to talk after being interrupted.

The Basic questions walk through that loop: the overall architecture, cascaded pipelines versus speech-to-speech models, voice activity detection, turn detection, interruption handling and tool use mid-call. The Advanced questions cover what separates a demo from a production line: keeping latency low, confirming actions, testing with realistic audio, sounding natural, preventing wrong actions and handing over to a human. Three added questions cover the phone network, capturing names and numbers accurately, and the metrics that tell you whether a deployed agent is working.

How does a voice agent work?

A voice agent runs a continuous loop: capture audio, detect speech and the end of the user's turn, turn speech into meaning, let a model decide what to say or which tool to call, then stream synthesized speech back while still listening.

Most production voice agents are built as a cascaded pipeline of streaming components. Audio arrives in small frames (typically 10 to 30 milliseconds each) from a phone line, browser or device. A voice activity detector (VAD) marks which frames contain speech. A speech-to-text (STT, also called automatic speech recognition or ASR) engine turns that speech into text, emitting partial transcripts as the person talks and a final transcript when they stop. A turn detector decides when the user has finished, and only then does the system commit to a reply.

The transcript goes to a language model along with the system prompt, the conversation so far and the tool definitions. The model either streams back text or requests a tool call, such as checking appointment slots. Its text output is fed, a sentence or phrase at a time, into a text-to-speech (TTS) engine, which streams audio back to the caller. Streaming at every stage is what makes the loop feel conversational: the agent starts speaking its first sentence while the model is still writing the second.

An orchestrator ties these pieces together. It holds conversation state, routes audio between components, cancels the model and the speaker when the user interrupts, runs tools, and enforces timeouts. This is where most of the real engineering lives, because each component is fairly good on its own and the failures happen at the joins: a turn detector that fires too early, a tool call that takes four seconds of dead air, or a history that records words the caller never heard.

The alternative is a speech-to-speech model that takes audio in and produces audio out directly, which collapses the middle of the pipeline. Either way, the surrounding concerns are the same: turn-taking, interruptions, tools, confirmation of actions and handoff to a human. The rest of this chapter takes those one at a time.

What is speech-to-speech AI?

Speech-to-speech AI takes spoken audio in and produces spoken audio out. It can be a cascade of speech-to-text, a language model and text-to-speech, or a single native audio model that reasons over audio directly without an intermediate transcript.

The term covers two quite different architectures. A cascaded system chains three models: STT converts audio to text, a text language model produces a reply, and TTS converts that reply back to audio. A native speech-to-speech model (sometimes called an end-to-end or realtime audio model) represents audio as tokens inside the model itself, so it hears the input and generates output audio in one pass. Several providers now offer native audio models over streaming APIs, and some research systems go further with full-duplex designs that listen and speak at the same time.

Native models have real advantages. They avoid the hops between three services, which can cut latency. They also hear things a transcript throws away: tone, hesitation, emphasis, laughter and the difference between a sarcastic "great" and a sincere one. Their output can carry matching prosody, so they can sound warmer and react more naturally.

Cascades keep advantages that matter in production. Every step leaves a text trail you can log, search, redact and evaluate. You can swap the STT engine for one tuned to your domain vocabulary, pick any text model with the strongest tool calling, and choose a TTS voice independently. Guardrails and policy checks written for text agents work unchanged. Debugging is easier because you can see whether a failure came from mishearing, reasoning or speaking. Native models often expose a transcript too, but it is typically produced alongside the audio rather than being the thing the model reasoned over, so it can differ from what was actually said.

Neither wins everywhere. As of late 2026, many teams run cascades for transactional phone work where accuracy, auditability and tool reliability dominate, and native models for companion, coaching or language-practice products where expressiveness matters most. Some hybrid designs use a native model for conversation and route high-stakes actions through text-based tool calls with strict validation.

What is voice activity detection?

Voice activity detection (VAD) classifies each short frame of audio as speech or non-speech, so the system knows when someone starts and stops talking and can ignore silence, line hiss and most background noise.

Audio reaches the agent as a continuous stream, much of it silence or noise. A VAD looks at each frame (commonly 10 to 30 milliseconds) and outputs a speech probability. Older detectors used signal energy and simple statistical models; most current systems use a small neural network trained to separate human speech from noise, music and silence. These models are cheap enough to run on every frame, on a CPU, in real time.

A raw per-frame probability is too jumpy to use directly, so VAD output is smoothed with a few parameters. A threshold decides what counts as speech. A minimum speech duration (for example 100 to 250 ms) ignores clicks and coughs. A minimum silence duration decides how much quiet ends a speech segment. Padding or pre-roll keeps a little audio before the detected start so the first syllable is not clipped. Using different thresholds for starting and stopping (hysteresis) stops the state flickering on and off around the boundary.

VAD matters for three reasons. It gates what goes to STT, which saves cost and reduces hallucinated transcripts from noise. It provides the earliest signal that the user has started talking, which drives interruption handling. And its end-of-speech signal is the raw input to turn detection. It is not the same as turn detection: VAD says the person stopped making sound, while turn detection decides whether they have finished their thought.

Tuning is a trade-off. A sensitive VAD catches soft speakers but fires on a television in the background or the agent's own voice leaking back through a speakerphone. A strict one misses quiet callers. Echo cancellation upstream of the VAD, and tuning on audio from your actual channel, matter more than the choice of model.

What is turn detection?

Turn detection, also called endpointing, decides when the user has finished speaking and the agent should respond. Good systems combine silence duration with signals about whether the words so far form a complete thought.

Human conversation moves fast. Studies of turn-taking across many languages find the typical gap between one speaker stopping and the next starting is around 200 milliseconds, which means people predict the end of a turn rather than waiting for silence. A voice agent has to make the same call, and it faces a direct trade-off. Respond too early and it talks over someone who was only pausing ("I'd like to book for... next Tuesday"). Respond too late and the caller sits in silence, often repeating themselves, which then collides with the agent's late reply.

The simplest approach is a silence timeout: once the VAD reports, say, 700 ms without speech, the turn is over. It is predictable but crude. Short timeouts cut off slow speakers, people reading out numbers and anyone thinking aloud; long timeouts make every reply feel sluggish.

Better systems add a semantic signal. A small classifier or language model looks at the partial transcript and estimates whether the utterance is complete. "My account number is 4 4 7" is clearly unfinished; "Yes, that's right" is clearly done. The system then uses an adaptive timeout: short (200 to 400 ms) when the words look complete, long (1 second or more) when they trail off mid-phrase or mid-number. Some models also use audio cues such as falling pitch at the end of a statement. Native audio models often do turn detection internally, with settings for how eager they are.

Context changes the right answer. After the agent asks for a 16-digit card number, a longer wait is right. After a yes-or-no question, a fast response is. Good orchestrators let each conversational state set its own turn-taking policy, and treat a false start as recoverable: if the user resumes speaking just after the agent begins, interruption handling cancels the reply and the turn continues.

What is interruption handling?

Interruption handling, often called barge-in, is how the agent reacts when the user starts speaking while it is talking: stop audio quickly, cancel pending generation, record what was actually heard, and respond to the new input.

People interrupt constantly, usually for good reasons: to correct a detail ("no, Wednesday"), to skip a long explanation, or because they already know the answer. An agent that keeps talking feels like a recorded menu. Supporting barge-in means the orchestrator listens during its own playback and, when the VAD detects real user speech, stops the TTS audio, cancels any in-flight model and TTS requests, and flushes buffered audio that has not yet played.

The step most teams miss is truncating the history. The model generated a full reply, but the caller may have heard only the first half. If the full text goes into the conversation history, the model will believe it already told the caller the cancellation policy, and later turns will be wrong. The orchestrator needs to track playback position and store only the words actually played, often marked as interrupted, so the model knows where it was cut off.

Not every sound is an interruption. Callers say "mm-hmm" and "right" while listening (called backchannels), cough, or talk to someone else in the room. Stopping for all of these makes the agent halting. Common defences are a minimum speech duration before cutting off, requiring a recognised word or two, and ignoring a short list of backchannel words. Echo is the other hazard: on a speakerphone the agent's own voice comes back through the microphone, so acoustic echo cancellation has to remove it before the VAD sees it.

Some moments should not be interruptible, such as a legally required disclosure or the final read-back before a payment. In those states the agent can keep speaking, or pause and resume the required text afterwards, and it should log that it did.

How can voice agents use tools?

Voice agents use tool calling like text agents: the model requests an action such as checking slots or looking up an order, the application runs it, and the agent speaks the verified result. Voice adds latency masking, spoken summaries of results and stricter confirmation.

Tool calling works the same way it does in text. The model is given tool definitions with JSON schemas, emits a structured call such as find_slots(date="2026-10-13"), and the application executes it and returns the result to the model. The model never touches the booking system directly; your code does, with the caller's identity and permissions. Everything in the tools chapter about schemas, validation, idempotency and authorization applies unchanged.

What changes is time. In a chat interface a three-second lookup is a spinner; on a call it is dead air, and after about two seconds callers start asking "hello?". Common patterns: have the agent say a short filler before slow tools ("Let me check that for you"), start safe read-only lookups speculatively as soon as intent is clear, and run independent lookups in parallel. For genuinely slow work, such as a back-office check that takes a minute, run it asynchronously, keep talking or offer a callback, and let the result arrive as an event.

Results must be turned into speech, not read out. A tool might return 14 available slots as JSON; the agent should offer two or three and ask a narrowing question. Identifiers should be normalised for speech ("order ending 4 7 2") and raw IDs, URLs and tables kept off the audio channel. Instruct the model explicitly that its output will be spoken.

Write tools need extra care because speech is ambiguous and the caller cannot see what is about to happen. A common split is: read tools run freely, write tools require a spoken confirmation of the exact parameters, and the server re-validates everything before committing. Tool results that come from external content are untrusted input, the same as in any agent.

How do you keep latency low?

Stream every stage, overlap work instead of running it in sequence, keep the first spoken sentence short, colocate services, keep prompts small and cached, and measure voice-to-voice latency per stage so you know where the time goes.

The number that matters is voice-to-voice latency: the time from when the caller stops speaking to when they hear the first sound of the reply. Under about 800 ms feels responsive; past about 1.5 seconds callers start talking over the agent or assume it has failed. That figure is the sum of several stages, and the main technique is to stop them adding up by running them as overlapping streams.

Streaming STT produces partial transcripts while the user talks, so the final transcript is ready almost as soon as they stop. Turn detection is often the largest single contributor, because any fixed silence wait is pure delay; semantic endpointing that ends confident turns after 200 to 300 ms saves more than most model changes. The model's time to first token depends on prompt size, model size and provider load, so keep the system prompt and history compact, use prompt caching for the stable prefix where the provider supports it, and summarise long histories. Streaming TTS should start on the first phrase, not the full reply, and the prompt should encourage a short opening clause ("Sure, Tuesday works.") so there is something to say quickly.

Overlap goes further. Some systems start the model speculatively on a confident partial transcript and discard the result if the user keeps talking. Read-only tool lookups can start as soon as the intent is clear. Network hops count too: placing STT, model and TTS in the same region as the telephony media server can save 100 ms or more per round trip, and keeping connections warm avoids repeated TLS handshakes. Phone networks themselves add delay you cannot remove, which is one more reason to win back time elsewhere.

Measure per stage and per percentile. Averages hide the long tail, and callers remember the four-second pause, not the median. Instrument each turn with timestamps for end of speech, final transcript, first token, first audio byte and playback start, then track p50 and p95 for each. Smaller or faster models are a valid trade if evals show task success holds; a faster model that needs one more clarifying turn is slower overall.

How should a voice agent confirm actions?

Before committing an action, the agent reads back the exact details in plain spoken form, asks for an explicit yes, and only then calls the write tool with those same confirmed values. Confirmation strength should scale with the cost of being wrong.

Speech recognition is fallible, and callers misspeak. "Fifteen" and "fifty" sound alike, "next Friday" is ambiguous, and names are frequently misheard. On a screen the user would see the mistake; on a call they cannot. Confirmation is the step where the agent turns its interpretation into words the caller can check: "Moving your appointment to Wednesday the 14th of October at 3 PM with Dr Rao. Shall I go ahead?"

Read-backs should use resolved, absolute values, not the caller's words. Repeating "next Tuesday" confirms nothing; saying "Tuesday the 13th" exposes a misunderstanding. Numbers should be grouped for listening ("ending 4 7 2 9"), money stated with currency, and anything irreversible flagged ("This cannot be undone"). Keep it short: one action, the key fields, one question.

Dialogue design distinguishes explicit confirmation (a direct yes or no question before acting) from implicit confirmation (weaving the value into the next question: "Tuesday the 13th, and what time suits you?"). Implicit confirmation is fine for low-stakes slots the caller can correct as the conversation continues. Explicit confirmation is required for anything that moves money, cancels, deletes, shares data or is hard to undo. The answer itself must be interpreted conservatively: "yes" confirms; "um, I think so, but..." or silence does not.

Confirmation must be bound to the action. The safest pattern has the model propose an action, the orchestrator store the exact parameters as a pending action, the agent read those stored values back, and a confirmed yes trigger execution of that stored action rather than a fresh model-generated call. That prevents the model from subtly changing a value between read-back and execution. Server-side validation still runs afterwards; confirmation guards against misunderstanding, not against unauthorised requests.

How do you test voice agents?

Test at three layers: components on recorded audio, whole conversations with simulated callers speaking through realistic audio conditions, and production calls through sampling and review. Score task success, actions taken, turn-taking and latency, not just transcripts.

A voice agent can fail in places a text agent cannot: mishearing, cutting people off, talking over them, freezing during a slow tool, or saying something that reads fine but sounds wrong. A text eval of the model alone catches none of these. Testing needs to cover the audio path end to end, under conditions that resemble real calls.

Component tests use recorded audio with known labels. Measure STT word error rate (WER, the share of words substituted, deleted or inserted) on your own call audio, and more importantly the accuracy on the entities that matter: dates, amounts, names, account numbers. Test the VAD and turn detector on labelled recordings that include long pauses, digit strings and background speech. These tests are cheap and catch regressions when you change a vendor or a threshold.

Conversation tests use a simulated caller: a language model given a persona and a goal ("you want to move your Tuesday appointment to Wednesday, you are in a car, you correct the date once") whose text is synthesised to speech, passed through degradations, and played to the agent in real time. Degradations should include the 8 kHz phone codec, background noise at several levels, varied accents and speaking rates, mid-sentence pauses and deliberate interruptions. Tools are mocked, including slow responses, errors and empty results. Each run is scored with deterministic checks on the trajectory (right tool, right arguments, confirmation before writes, no forbidden actions) and on timing (voice-to-voice latency, premature endpoints, stop latency on barge-in), with a calibrated judge for softer qualities such as clarity.

Production review closes the loop. Sample calls, especially transfers, hang-ups and long calls, have people listen to some of them, and turn every real failure into a new scripted scenario. Because each simulated run is stochastic, run each scenario several times and track pass rates rather than single outcomes. Regression-test before every prompt, model, voice or threshold change.

What makes a voice agent sound natural?

Natural voice agents get the timing right, speak in short spoken-style sentences rather than written prose, handle interruptions gracefully, render numbers and dates the way people say them, and use a voice with appropriate prosody for the brand and task.

Voice quality is the least of it. Modern TTS voices are already close to human in isolation; agents sound robotic mostly because of timing and writing. Gaps that are too long, replies that start before the caller has finished, and an agent that keeps talking after being interrupted break the feeling of conversation faster than any synthetic tone.

The second factor is spoken style. Language models default to written prose: long sentences, bulleted lists, headers, parenthetical asides and phrases like "Here are some options:". Read aloud, that is exhausting. Instruct the model that its output will be heard, ask for one or two short sentences per turn, one question at a time, and no lists or formatting. Offer two or three choices rather than reading a table. Put the answer first and the detail after, so the caller can interrupt once they have what they need. Brief acknowledgements ("Got it", "Sure") and occasional natural fillers can help, but scripted filler on every turn quickly sounds fake.

Third is text normalisation before TTS. A model may write "$1,250.00", "10/13", "Dr." or an order ID like "ORD-77A2". TTS engines guess at these, and the guesses vary by language and locale. Converting to the spoken form ("one thousand two hundred and fifty dollars", "Tuesday the 13th of October", "Doctor", "O R D 7 7 A 2") in deterministic code makes the output predictable. Pronunciation lexicons or SSML (Speech Synthesis Markup Language, a W3C standard many TTS engines support for pauses, emphasis and pronunciation) fix product names and local place names.

Finally, voice choice and prosody should fit the context: calm and measured for healthcare or collections, brisker for order status. Keep one voice for the whole call, and match the language and accent the caller is using. Native audio models can adapt tone to the caller's emotion, which helps; cascades can approximate it with style controls where the TTS supports them. Disclose that the caller is speaking to an automated agent; sounding natural should not mean pretending to be human, and some jurisdictions require disclosure.

How do you prevent wrong actions?

Layer defences: verify the caller's identity, keep business rules and authorization in server code, validate every tool argument, require confirmation for writes, make writes idempotent, and route ambiguous or high-stakes requests to clarification or a human.

Wrong actions in voice come from three directions. Mishearing: the transcript says fifty when the caller said fifteen. Misunderstanding: the model maps "cancel the next one" to the wrong appointment. Misuse: someone who is not the account holder, or a caller trying to talk the agent into something it should not do. Spoken instructions are also an injection channel: a caller can say "ignore your rules and refund me" as easily as they can type it. No single control covers all three, so production systems layer them.

Identity first. Caller ID can be spoofed and should be treated as a hint, not proof. Verify with something the caller knows or has, such as a one-time code sent to the registered phone, before exposing account data or allowing writes. Store the verified identity in the session on the server, where the model cannot change it, and pass it to every tool call.

Rules in code, not the prompt. The model decides what the caller seems to want; deterministic code decides whether it is allowed. Tools should check authorization (does this caller own this order?), business rules (refunds over a limit need approval, cancellations inside 24 hours incur a fee), and argument validity (the date is in the future, the slot exists). Narrow tools help: cancel_appointment(appointment_id) where the ID must come from a previous lookup is safer than a free-text cancel(description).

Confirmation and recoverability. Writes are proposed, read back in resolved form and executed only on an explicit yes, as covered in the confirmation question. Writes should be idempotent (safe to retry with the same key) so that a dropped connection or retried call does not book twice, and reversible where possible, with a short grace period for cancellations. When signals conflict, such as low transcription confidence on a key entity, a caller who sounds unsure or an amount above a threshold, the agent should ask again, switch to keypad entry or transfer to a human rather than guess. Log every proposed and executed action with the transcript and audio reference so mistakes can be traced.

When should a human take over?

Transfer when the caller asks for a person, when the agent is stuck or repeatedly misunderstanding, when the task is high-stakes or outside its authority, or when distress, vulnerability or legal risk appears. Hand over with a summary so the caller never repeats themselves.

A voice agent is valuable for the calls it can complete, and harmful on the calls it should not attempt. Handoff (also called escalation or transfer) should be designed as a normal outcome with clear triggers rather than an apology at the end of a failed call.

Triggers fall into four groups. Explicit request: the caller asks for a person. Honour it promptly; making people repeat "agent" five times is one of the most disliked patterns in phone automation. One attempt to resolve quickly is reasonable if the caller agrees, but do not trap them. Agent struggle: repeated failure to understand the same slot, the caller rephrasing several times, low transcription confidence on key entities, loops in the dialogue, or rising frustration. Policy and stakes: disputed payments, legal threats, complaints, anything outside the tools' authority, or amounts above an automation limit. Safety and vulnerability: signs of distress, medical emergency, threats of self-harm, or a caller who appears confused or vulnerable. Some of these, especially safety, should have fixed scripted responses and immediate routing rather than model judgement.

The quality of a transfer is as important as the decision. A warm transfer passes the human a short structured summary: verified identity status, what the caller wants, what the agent already did, and why it is transferring. The caller should hear what is happening ("I'm connecting you to a colleague who can sort out the disputed charge; I've passed on the details") and roughly how long it will take. If no one is available, offer a callback with a captured number and time rather than an endless queue.

Track transfer rate and reasons together. A low transfer rate is not automatically good; it can mean callers are hanging up instead. Review transferred calls regularly, since they show where to add tools or improve dialogue, and review a sample of non-transferred calls to find the ones that should have been.

How does a voice agent connect to the phone network?

Phone calls reach a voice agent through a telephony provider or SIP trunk that bridges the public phone network to a media server, which streams the call audio, usually 8 kHz narrowband, to your orchestrator and plays synthesized audio back.

The public switched telephone network (PSTN) does not speak WebSockets. To answer phone calls, a voice agent needs a bridge: a telephony provider or a SIP trunk (SIP, the Session Initiation Protocol, is the signalling standard for internet telephony) that owns phone numbers, receives calls and hands them to software. From there, either the provider streams raw call audio to your server over a WebSocket, or your own media server terminates the SIP call and handles the audio directly. Browser and app clients usually connect over WebRTC, which handles jitter, packet loss and echo cancellation for real-time media.

Phone audio is very different from microphone audio. Classic telephony uses the G.711 codec at 8 kHz sampling (narrowband), which cuts off frequencies above roughly 3.4 kHz. Sibilants like "s" and "f" lose detail, which is why "fifteen" and "fifty" or "S" and "F" are hard to tell apart. Mobile and internet calls may arrive as wideband audio but are often transcoded down along the way. Your STT model should be one that performs well on 8 kHz audio, and your test audio must go through the same codec. TTS output must be resampled and encoded to match the line.

Telephony brings its own features. DTMF (the keypad tones) is a reliable channel for digits like PINs and account numbers, and can be offered when speech recognition struggles. Answering machine detection matters for outbound calls so the agent does not deliver a conversation to voicemail. Call transfer, either via a SIP REFER message or the provider's API, is how a human handoff actually happens. Recording, consent announcements and retention rules are set at this layer too, and vary by country and industry.

The network itself adds latency (often 100 ms or more each way on mobile calls) and occasional packet loss, so the budget for your own pipeline is tighter on phone than on web. Place the media server and the AI services close together, and treat jitter, dropped audio and one-way audio as failure modes to monitor.

How do you capture names, numbers and addresses accurately by voice?

Treat spoken entities as the hardest part of the call: bias recognition toward expected values, collect them in a dedicated state with longer turn timeouts, validate with checksums and database matches, read them back in chunks, and fall back to spelling or keypad entry.

General transcription is good enough for intent, but structured values are where voice agents fail most. Names have unusual spellings, email addresses mix letters and symbols, account and policy numbers are long, and narrowband audio blurs similar sounds (B, D, P, T, V; M and N; fifteen and fifty). A transcript with a low overall word error rate can still get one digit wrong in a 12-digit account number, and that one digit makes the whole interaction fail.

Bias the recogniser. Many STT services support custom vocabulary, keyword boosting or phrase hints. When the agent is collecting a known type, pass expected values: the caller's own name from the account, the street names in their postcode, the product catalogue. Some services also accept a hint that the next answer is digits, a date or a spelled sequence. Give the caller time: switch the turn detector to a longer timeout while collecting digits or spelling, since people pause between groups.

Validate in code. Card numbers have a Luhn checksum, many account and policy numbers have their own check digits or formats, postcodes and phone numbers have known patterns, and dates must be plausible. Match against what you already know: if the account has three orders, fuzzy-match the spoken number to those three rather than searching everything. A failed check is a cue to ask again, not to guess.

Read back and offer alternatives. Read long numbers in groups ("4 4 7 2, 9 1 0 3") and names letter by letter when it matters, optionally with a phonetic alphabet ("S for Sierra"). Offer the keypad for digits on phone calls, or send a link or text message for values like email addresses that are painful to say aloud. Never read back full sensitive values such as card numbers; confirm the last four digits instead.

Which metrics show whether a voice agent is working in production?

Track outcomes first (task completion, containment with resolution, transfers, repeat calls), then conversation quality (latency percentiles, interruptions, re-prompts, entity accuracy), and connect each call's metrics to its trace so a bad number leads straight to the calls behind it.

Voice agents are often judged by containment rate, the share of calls handled without a human. On its own that number misleads: a caller who hangs up in frustration is contained, and so is one who gives up and calls back tomorrow. Pair containment with task completion (did the booking, payment or answer actually happen, verified from system records), repeat contact within a few days on the same issue, and abandonment (hang-ups before resolution). Together they distinguish resolved calls from calls that merely ended.

Conversation quality metrics explain why outcomes move. Voice-to-voice latency at p50 and p95. Premature endpoint rate: how often the caller resumes speaking within a second of the agent starting. Barge-in rate and stop latency. Re-prompt rate: how often the agent says it did not catch something. Entity accuracy on sampled calls for names, dates and numbers. Silence events: gaps over two or three seconds, usually from slow tools. Each of these maps to a specific component, which makes them actionable.

Safety and cost metrics complete the picture: actions executed without confirmation (should be zero), authorization denials, handoff reasons, cost per completed task (STT minutes, model tokens, TTS characters and telephony minutes add up differently from text), and average call length. Post-call surveys add a human view, but response rates are low and skewed, so treat them as a signal rather than ground truth.

All of this depends on per-call traces: one record per call with every turn's timestamps, transcript, model input and output, tool calls, confirmation events and a reference to the audio. Dashboards then let you click from a bad p95 day to the specific calls that caused it. Sampled human listening remains essential, since many voice failures, such as an odd tone, a mispronounced name or an awkward pause, are obvious to a listener and invisible in the numbers.