Direct answer
Inbound and outbound voice AI share models, but they do not share operating constraints. Inbound optimizes for queue wait, barge-in, and warm transfer. Outbound optimizes for pacing, consent, and CRM outcomes. Design the call path and risk map before picking Realtime or a voice platform. See Voice AI Systems.
Inbound architecture
Inbound agents sit on a live queue. Callers interrupt. Context from IVR and CRM must arrive before first reply.
- Number routing and hours of operation
- Identity lookup without leaking PII into prompts
- Tool latency budgets so silence does not feel broken
- Warm transfer packages for humans when confidence or policy requires it
Outbound architecture
Outbound is paced work. Compliance and list quality dominate.
- Consent and quiet hours
- Retry and voicemail policies
- CRM disposition write-back that sales trusts
- Campaign-level cost and connect-rate monitoring
Shared production requirements
Both paths need transcripts, retention policy, eval suites, and incident runbooks. Speech-to-speech Realtime fits conversational quality. Chained STT/LLM/TTS fits heavier text control. Many products mix both.
Proof: Apptension ships production AI systems handling 1.5M+ calls/month with senior engineers only.
Decision checklist
- Is the primary job inbound containment, outbound conversion, or both?
- What must hand off to a human with full context?
- Which tools block the turn, and what is the silence budget?
- What recording, retention, and PII rules apply?
