Voice
How AI Voice Agents Work
A voice agent is a text system with two conversions bolted to its ends and a stopwatch running. Understanding why latency, rather than accuracy, is the binding constraint explains almost every design decision in one.
Drafted with AI assistance and edited by the Bralak engineering team. No client examples, performance figures or benchmark comparisons appear in these pieces — every technical claim is one a reader can check independently.
Strip the marketing away and a voice agent is a text agent with speech recognition in front of it and speech synthesis behind it. Everything difficult about building one comes from what those two conversions cost in time, and from the fact that a phone call gives you nowhere to hide the cost.
In chat, a system that takes three seconds to reply looks like it is thinking. On a call, three seconds of silence is a dropped line, and the person says hello? and starts again.
The pipeline, from ring to reply
- InputIncoming call
- SystemSpeech recognition
- DecisionIntent
- ReasoningAgent reasoning
- RetrievalKnowledge retrieval
- ActionBusiness action
- ActionSpoken response
- Human escalationWarm transferAvailable at any step
The call arrives over the phone network or a SIP trunk and audio starts streaming. This layer also owns the things that turn out to matter operationally — recording consent, call transfer, DTMF tones, and what happens when the network degrades.
Audio becomes text, streaming rather than after the fact. It works on partial audio and revises as more arrives, which is why a good transcript appears to correct itself mid-sentence. Waiting for a complete, final transcript before doing anything else is the single most common cause of a sluggish agent.
Deciding the caller has finished speaking. This is its own hard problem and it is where a lot of perceived rudeness originates: cut too early and the agent talks over someone drawing breath mid-sentence, wait too long and every exchange gains a beat of dead air. Someone reciting a long account number is the case that breaks naive settings.
The text agent, exactly as it would be in chat — context, tools, retrieval, a decision about what to do and say. The only difference is that its time budget is a fraction of what a chat interface would tolerate.
Text becomes audio, streamed out as it is produced rather than generated whole. Streaming here is what lets the first syllable start while the rest of the sentence is still being written.
Note that every stage is streaming and overlapping. A pipeline built as five sequential steps, each waiting for the last to finish, produces a system that is correct and unusable.
Why latency is the binding constraint
Ordinary human conversation runs on turn gaps of roughly a couple of hundred milliseconds — a well-documented finding in conversation analysis, and consistent across languages. That number is not a target anyone will hit with this architecture, but it is the yardstick your caller is unconsciously applying.
The consequence is a fixed budget shared across the whole pipeline, and every component spends from the same pot. Recognition needs to commit, the model needs to think, synthesis needs to start, and the network needs its round trips. Add a retrieval call and a database lookup and the budget is gone before the reasoning starts.
In text, quality is the constraint and latency is a nicety. In voice they trade directly against each other, and the trade is made on every single turn.
This is why voice agents are architected differently from their text equivalents even when they do the same job:
- Smaller, faster models on the conversational path, with the larger model reserved for the steps that genuinely need it.
- Speculative work — starting retrieval on a partial transcript, because a wasted lookup is cheaper than a pause.
- Filler that is honest. Let me pull that up buys real time and is true. Silence and stalling that is not true are the two ways to spend it badly.
- Aggressive scoping. Every extra tool call is time the caller spends listening to nothing.
Interruption is a feature, not an edge case
People interrupt. They interrupt to correct a wrong assumption, to answer before the question finishes, to say no, the other order. A system that cannot be interrupted forces the caller to sit through a sentence they already know is wrong, and it is the fastest way to make an agent feel like an obstacle rather than a service.
Handling it — barge-in — means listening while speaking, detecting that the caller has started, stopping synthesis mid-word, and then dealing with the awkward part: the agent’s own state. It said half a sentence. The caller heard half a sentence. What the model believes it said and what the caller actually received have diverged, and the context has to be corrected to what was truly heard, or the rest of the call is built on a false transcript.
There is also a discrimination problem underneath it. A cough, a colleague talking nearby, a door — none of these are the caller taking a turn. Cutting off mid-sentence for background noise is its own failure, and tuning this is genuinely per-deployment work: a call centre and a car have different acoustics and want different thresholds.
Transfer design decides whether people trust it
Every voice deployment needs an answer to what happens when this does not work, and the quality of that answer shapes the caller’s opinion of the whole system more than any successful call does.
Explicit request, obviously, but also repeated failure to progress, detected distress, and any topic the deployment has decided is out of scope. A loop that tries the same clarifying question a third time should be a transfer, not a fourth attempt.
The human should receive who the caller is, what they wanted, what has already been tried and what was already said. Making somebody repeat everything they just explained is the moment the caller concludes the agent wasted their time — and they are right.
When the pipeline itself breaks — recognition unavailable, model timing out — the fallback is a queue, not an apology loop. This path is easy to leave untested because it only runs on a bad day, which is precisely when it is being judged.
What voice is genuinely good for
Voice is the right channel when the caller’s hands or eyes are busy, when they do not have an app and will not install one, when the interaction is short and transactional, or when the phone is simply the channel your customers already use. Booking, rescheduling, order status, qualification, routing, outbound reminders and confirmations all sit comfortably here.
It is the wrong channel when the information is dense or needs to be re-read, when the caller must compare options side by side, when what is needed is a document, or when the matter is emotionally difficult enough that a person is the point rather than the fallback. A long list of choices read aloud is worse than a screen showing them, and no amount of engineering changes that.
The most reliable voice deployments are narrow on purpose. One or two intents, done well, with a fast route to a human for everything else, beats an agent that attempts the full contact-centre surface and is mediocre across all of it. Narrow also means the latency budget is spent on one path you can actually tune.
This piece supports our AI Voice Agents page, which covers how we design, build and run these systems in production.
Your Next Intelligent System Starts Here.
Tell us what you’re trying to improve, automate or build. We’ll help you identify the right AI strategy and engineering path.