Glossary
Voice latency
Voice latency is the delay between a caller finishing speaking and hearing the reply; for AI voice agents it decides whether a call feels like a conversation.
Also called response latency, voice AI latency, turn latency, end-to-end latency, call delay.
In telephony the term also means the time sound takes to travel from one end of a call to the other. When latency is short, the call feels like a conversation. When it is long, callers repeat themselves, talk over the agent or hang up.
How it works
On an ordinary call, delay comes from the network: encoding the audio, sending it across operators and decoding it at the other end. The ITU’s Recommendation G.114 (2003) advises that one-way delay should not exceed 400 ms for general network planning, and notes that highly interactive tasks, including many voice calls, can be affected by much lower delays.
An AI voice agent adds its own steps on top of the network:
- End-of-turn detection. The agent has to decide the caller has finished, not just paused. Waiting longer avoids interruptions; waiting less feels quicker. This trade-off shapes how the whole call feels.
- Understanding. The agent works out what the caller said.
- Deciding. The agent chooses the reply, sometimes after looking something up in a brochure or the CRM.
- Speaking. The agent’s voice starts. In well-built systems the first words play before the whole sentence is ready.
- The trip back. The audio returns through the SIP trunk and the carrier network to the caller’s phone.
End-to-end latency is the sum. Agents that listen and speak in a single step remove some of the hand-offs between steps.
Why it matters
Conversation has rhythm. A pause that would be normal in a text chat feels like a dropped call on the phone, and a caller who hears silence tends to say “hello?” into it. A slow agent also invites overlap: the caller starts again just as the agent starts speaking, and both stop.
Latency is not only speed. An agent that answers instantly but cuts the caller off mid-sentence is worse than one that waits a moment longer. The goal is replies that arrive when a person would reply.
Example
A buyer says, “Main Saturday ko aa sakta hoon, but morning mein nahi.” (I can come on Saturday, but not in the morning.) A well-tuned agent waits through the short pause after “aa sakta hoon” (I can come), hears the rest, and replies with an afternoon slot. A badly tuned one jumps in at that pause with “Great, Saturday 10 baje?” (Great, Saturday at 10?) and has to be corrected.
Common confusions
- Latency vs speed of speech. A fast-talking voice with a long wait before it starts still feels slow. What callers notice is the gap.
- Average vs worst case. An agent that is quick most of the time but occasionally stalls feels unreliable. Look at the slowest replies, not just the typical one.
- Latency vs IVR menus. An IVR plays a message and waits for a keypress, so timing is set by the menu. An agent has to get the timing right on every turn.
- Network vs processing. Moving the agent closer to the caller cuts network delay, but it does nothing for the time the agent spends detecting the turn, thinking and starting to speak.