-
Silence Is Not Failure: Teaching a Voice Agent to Wait
A caller asked for time, but the voice agent kept asking them to repeat themselves. I added a waiting state so the recovery timer could tell a deliberate pause from a missing reply.
-
The Sentence Nobody Heard: What a Voice-AI Recording Actually Proves
A voice-agent transcript and recording can both contain words the caller never heard. We fixed this by recording closer to playback and using one provider clock for both sides of the call.
-
Generate Early, Speak Late: Why Voice Responses Need an Admission Gate
A voice agent can start preparing a reply during a pause, only for the caller to begin speaking again before the reply reaches the phone. We fixed that race by separating response generation from permission to speak.
-
LLM-Based Turn Detection in Voice Agents: Three Failure Modes from Production
Knowing when a caller is done talking is the hardest continuous decision in a voice pipeline. Audio-only endpointers can't judge semantic completeness, so we let the LLM decide. It worked, until three failure modes showed up in production. All of them were races between asynchronous components.
-
Streaming STT Has No Wall Clock
A caller's sentence vanished from a production voice call with no error anywhere. The hunt for it went through Deepgram endpointing internals, a mobile-network theory that died on the standards, Twilio's undocumented silence behavior, and four-day-old logs. What it found applies to anyone building real-time voice.