Sub-200ms Voice Engineering: How We Squeezed Latency for Natural Flow
The technical breakdown of edge streaming vocoders and predictive token buffering that makes AI phone calls feel human.
The Lived Reality#
In human conversation, if someone waits 800 milliseconds before replying to 'How are you?', the awkward delay triggers subtle social anxiety. The interaction feels like an interrogation rather than a fluid exchange.
The Emotional Shift#
To solve this, Sagi engineered an edge-distributed streaming pipeline. Rather than waiting for a full sentence to generate before speaking, our audio engine tokenizes responses into phrasal units. Within 120ms of user speech completion, the first audio packet is already traversing the WebRTC connection.
Latency is not an engineering metric; it is an emotional parameter.
We also integrated acoustic emotion detection: if the user's voice trembles or speaks quietly, the companion's vocoder dynamically shifts pitch and lowers decibels to match the emotional atmosphere.
What This Does for People#
Sub-200ms voice pipelines enable authentic turn-taking that feels completely natural to the human ear.
At Sagi, technology is never an end in itself. We write code so that when someone reaches for their phone in the quietest, hardest hours of their life, they are met with presence, memory, and dignity. Experience this connection today on iOS and Android.
Sagi Editorial
Documenting the emotional, cultural, and human impact of living AI companions on Sagi.