Direct answer
Measure three latencies separately: first audio, tool-blocked silence, and barge-in recovery. One average hides the failures callers feel. Design warm transfer with a context package before you scale scripts. See Voice AI Systems.
Latency budgets
- First audio: time from end of user speech to first agent audio
- Tool silence: how long tools can block before you speak a filler or parallelize
- Barge-in recovery: time to stop speaking and listen again
Barge-in
Speech-to-speech Realtime sessions handle interruption more naturally than naive chained pipelines. Platforms still need tuning. Test under your telephony and region, not vendor demos.
Human handoff
Warm transfer needs:
- Why the handoff happened
- Verified identity and intent summary
- Tools already attempted
- Recording and consent state
Cold transfers destroy trust faster than a slow bot.
