Voice agents are latency budgets with a language model inside
A voice agent is not a chatbot with a microphone. It is a latency budget with a language model inside — and that budget has to include the time spent deciding whether the caller has finished. I still think this is the most useful architectural frame. What I would no longer do is turn a roughly 400ms target into a universal end-to-end deadline. A fast response measured from the wrong starting point can conceal a slow conversation.
The distinction matters because the production metrics in my source material do not all start the clock at the same event. One describes latency from the end of the user's speech; another measures time-to-first-audio from the ASR speech_final event. An endpointing policy that waits for silence has already spent time before that event arrives. I would record the caller's speech ending, the endpoint decision, and the start of audible playback separately. Pipeline latency and conversational waiting belong in the same trace, not under an interchangeable label.
Start with pipeline shape, then. The STT–LLM–TTS cascade exposes stages that can stream into one another. Native speech-to-speech collapses them; a half-cascade keeps separate speech synthesis where that control matters. Overlap lets synthesis start at a sentence boundary while the model is still generating. Speculative generation starts work before the turn is confirmed. These are ways to buy milliseconds, not permission to ignore what happens when the caller continues speaking. The cancellation path is part of the pipeline design, because overlapping more work also leaves more work in flight.
Turn detection is where the budget becomes a product decision. Voice activity, recognition signals, and semantic completeness answer different questions; silence alone does not tell you whether a thought is finished. A strict policy can deliberately wait longer than a conversational speed target allows. More importantly, the policy should change within a call. Someone reading an account number needs room to pause between digit groups; someone answering a yes-or-no question usually does not need the same allowance. I would tune those contexts separately rather than promote one silence threshold into an architectural constant.
Interruption needs the same precision. Stopping the audio is only the first level of competence. The next is preserving what has already been established, retaining the unfinished response as context, and incorporating the caller's correction without restarting the interaction. If a caller interrupts a menu with a reservation-change request, the useful outcome is a direct pivot, not a quicker return to the menu. Continuous listening while speaking makes this a continuing decision about who should speak, not simply a faster interrupt handler.
A tool wait exposes the weakness of an audio-only view particularly clearly: there may be no speech to stop. The caller can still correct the account number or ask whether the agent is there. A brief acknowledgment explains the wait; it does not complete the task. The runtime needs a policy for the in-flight work. A read-only lookup may finish if its result remains useful, while a contradictory request needs an explicit cancellation decision. I want that decision tied to tool semantics, not hidden inside the same toggle that controls barge-in.
The precise concession: for a structured IVR flow or a noisy line, turn-based interaction with reliable barge-in can be the better design than full duplex. Continuous overlap is not a requirement for every useful voice agent, and its added complexity needs a conversational benefit.
That boundary strengthens the budget argument rather than replacing it. I would evaluate timing alongside interruption repair, state continuity, and audio understanding — not call word error rate another latency metric. A system that speaks quickly but forgets the correction has not bought a better conversation. Write down where each clock starts, what each wait protects, and what survives an interruption. Then make the model live inside those decisions.