A voice agent is judged on its worst turn, not its average one.
Latency is the whole experience. A second of silence reads as thinking, two reads as broken, and a reply that lands while the caller is still speaking reads as rude. Under that sit five hard problems — recognition on a bad line, barge-in, endpointing, an answer that is grounded rather than invented, and a handoff that arrives with context. We build the turn. The voice itself is a vendor choice and nearly solved.
FACT 01A voice agent we built runs SPIN-method sales calls for a B2B team — qualifying, proposing slots and booking meetings in natural conversation, with voicemail, rescheduling and human handoff handled.
FACT 02A digital employee that speaks and writes, with an optional speaking face — and that acts across systems at the end of the conversation instead of merely concluding it.
FACT 03There are no call metrics on this page — no connection, booking or deflection rates, no containment, no word error rate. Latency appears only as a design threshold you can drag yourself, below.
30 minutes, working session. Play us thirty seconds of a call that went badly and you leave with a one-page read either way. Nobody follows up more than once.
02 / CONDITIONSAS HEARD
Why the last voice bot you used was unbearable
Almost none of it is the voice. Six failures, and every one of them is engineering.
LATENCY
The pause that kills it
A second and a half of silence after the caller stops. Short by any technical standard, unbearable in a conversation. People start talking into it, which breaks the turn, which makes it worse.
BARGE-IN
It cannot be interrupted
The agent is halfway through a sentence the caller has already answered, and there is no way to stop it. Everyone has been in this call and everyone remembers it.
VOCABULARY
It mangles the important word
Names, order numbers, addresses, drug names — the exact tokens the call exists to capture. Generic recognition is fine at conversation and poor at your vocabulary, which is the part that matters.
HANDOFF
The transfer arrives cold
The agent gives up and puts a human on the line with none of the context, so the caller repeats everything. Deflection improves; the experience it was measuring gets worse.
GROUNDING
It invented something
Confidently, in a voice, to a customer, with no transcript anyone reviews. Voice hides this failure better than text does, which is exactly why it is more dangerous here.
COST
Cost per minute was never modelled
Recognition, model, synthesis and telephony, each billed separately and discovered at volume rather than designed for.
03 / INSTRUMENTINTERACTIVE
Drag the silence until the conversation breaks
Two tracks: the caller above, the agent below. The slider is the whole round trip — recognition, model, synthesis and network. Nothing here makes a sound; the argument is entirely in the shape.
ILLUSTRATIVE — A WRITTEN EXCHANGE, NOT A CLIENT RECORDINGTURN: NATURAL
CallerAgentBoth speaking at once
End-to-end turn320 ms
Caller talks over the agentno
What the human calls itnatural
The exchange interleaves the way people actually speak. Nobody notices the machine, which is the only review a voice agent ever wants.
Our own agent, on a real call — sound off until you press play
A recording from the outbound agent below, already published on its case page. It is here so the schematic has something to be measured against. Voice agent case →
The agent, running
A call that ends in a calendar entry
Qualification, objection handling and a booked slot, logged where the sales team already works — built on n8n Cloud with Aircall and VAPI, writing outcomes to Pipedrive and Google Calendar in real time. Voice agent case →
SPIN-method outbound calls for a B2B sales team: qualifying, proposing slots, booking into Google Calendar and logging outcomes to Pipedrive as the call happens, with voicemail, rescheduling and human handoff for the edges. A sales call punishes every weakness in a turn — objections, interruptions, hesitation, and a person who hangs up the moment it feels mechanical.
NATURAL CONVERSATION — NOT A PHONE TREE WITH BETTER SPEECH
Natural spoken and written conversation with website visitors, with an optional 2-D or 3-D speaking face, answering questions and guiding people through products. The hard part is that it has to do something at the end — which turns a chat problem into a permissions-and-actions problem, with confirmation on anything consequential.
AI VIDEO AGENT SPEECH IN, SPEECH OUT LIVE CHARACTER
Recognition and synthesis around a language model, driving an animated 3-D character in a browser for children learning French. The speech loop shares its budget with animation and head tracking, and children are the least patient users available — a stall does not read as latency to them, it reads as the teacher freezing.
Tell us the call you want handled, or play us thirty seconds of one that went badly.
Speech in, speech out — AI video agentRecognition, a language model, synthesis and a 3-D character sharing one frame budget, live in a browser. AI video agent case →
05 / SCOPETYPICAL RANGES
The work, as we actually sell it
Timelines are indicative ranges and depend on scope. We name our tools plainly — ElevenLabs and Wav2Lip on the synthesis and lip-sync side, Vapi and Aircall on the call side, n8n where orchestration belongs outside the app — and we hold no partnership or reseller status with any of them.
A turn-latency read
1–2 WEEKS
Where your current experience spends its time — recognition, model, synthesis, network — and what it would take to get under the threshold where a conversation stops feeling mechanical. You keep the breakdown either way.
A voice agent for one defined conversation
8–16 WEEKS
One call type, end to end: telephony, recognition, grounding, synthesis, barge-in, handoff and logging. Deliberately one conversation — an agent that does everything badly is worse than a human queue.
Domain vocabulary work
SCOPE-DEPENDENT
Recognition tuned for the proper nouns your calls actually turn on — order numbers, product names, addresses. Usually the largest single quality gain available, and almost always the cheapest.
Grounding and refusal
SCOPE-DEPENDENT
Answers retrieved from your own content with a citation trail, and an agent that says it does not know instead of improvising. On voice this is not a nicety.
Handoff that carries context
2–5 WEEKS
A transfer where the human receives the transcript, the intent and what has already been attempted — so deflection stops degrading the experience it is measuring.
Cost and concurrency modelling
1–2 WEEKS
Recognition, model, synthesis and telephony modelled per minute at your real volume and your real concurrency, before you commit to a vendor mix.
A speaking face, where it earns its place
SCOPE-DEPENDENT
Audio-driven lip sync with the offset calibrated — the misalignment a person feels as wrong long before they can name it. And the honest advice when the interaction is better without a face.
A voice of your own
WITH CONSENT ONLY
We have built a replica of a person’s voice so their own digital employee could speak in it — with that person’s explicit consent, for their own product. That is the only circumstance in which we will do it.
The speaking face — video
When a face helps, and when it does not
A digital employee holding a spoken conversation with an optional speaking avatar. A face raises the bar rather than lowering it: lip-sync offset that is a few frames wrong reads as uncanny long before anyone can say why. Digital employee case →
Where a voice is chosen — digital employeeA language, a stored voice, a default or an adapted one, the line it will say and the speed it says it at. Everything on this screen is a setting somebody has to be accountable for, which is why the next section exists. Digital employee case →
06 / DISCLOSUREWHAT WE WILL NOT BUILD
Disclosure, consent and the work we turn down
This is engineering practice, not legal advice — requirements vary by jurisdiction and we are not your lawyers. It is also a position, and stating it publicly is worth more to us than any capability claim on this page.
We build in
Disclosure that the caller is speaking with an automated system, said at the start rather than admitted when challenged.
We build in
Consent handling for recording, an auditable record of what was said and what was done, and confirmation before any consequential action.
We build in
A route to a human that carries the transcript and the intent, available before the caller has to demand it.
With consent
A replica of someone’s voice, for their own product, with their explicit agreement — we have done exactly this, so that a founder’s digital employee could speak in their voice.
We will not
Replicate the voice of a person who has not agreed to it, whoever is asking and whatever the intended use.
We will not
Build an agent whose purpose is to be mistaken for a human being, or a robocall operation. This is not a pricing conversation.
07 / OBJECTIONSANSWERED PLAINLY
What you are probably thinking
“Our customers will hate talking to a robot.”
Some will, and the design decides how many. Almost all of the hatred comes from three things: latency, not being able to interrupt, and a handoff that loses context. Those are engineering problems, and they are the ones we build first.
“What if it says something wrong?”
Then it is a defect with a transcript attached. Answers are grounded in retrieved content rather than improvised, the agent is built to say it does not know, consequential actions require confirmation, and every call leaves an auditable record. Voice hides this failure better than text — which is why it needs more discipline, not less.
“Will it sound like a person?”
Close — and that is not the question we would ask. Voice quality is a vendor choice and nearly solved. What decides whether people tolerate the call is turn latency and interruption handling, the parts nobody demos.
“Is this even legal?”
Requirements vary by jurisdiction and we are not your lawyers. What we build in as engineering practice is disclosure that the caller is speaking with an automated system, consent handling for recording, and an audit trail. We will not build an agent designed to be mistaken for a person.
“We tried a voice bot and it was terrible.”
Almost certainly a phone tree with better speech attached: no barge-in, no grounding, and a handoff that dropped context. The read we do first measures those three specifically, so you find out whether the problem was the idea or the build.
“We tried an outside dev shop and it went badly.”
Usually the same three causes: no senior person accountable, a demo-grade codebase, and a handover that never happened. Here you get direct access to the engineers doing the work, no account-manager layer, a senior architect signing off every project, and tests and documentation as part of the deliverable.
How the read starts
Thirty seconds of a bad call tells us most of it
A working session in Los Angeles. The fastest way into this problem is to listen to one call that went wrong together and name where the turn failed — before anybody writes a requirement document about it. More about how we work →
Many of our clients have been with us for 7+ years straight. The Upwork and Clutch records are independently verifiable.
09 / ARTEFACTKEEP EITHER WAY
Not ready to talk? Take the checklist.
One page, 13 questions about your voice agent you should be able to answer before it calls a customer — end-to-end turn latency, whether a caller can interrupt, how proper nouns are handled, what the agent does when it does not know, what the human receives on handoff, whether disclosure happens, and cost per minute at volume.
10 / NEXTONE STEP
Play us thirty seconds of a call that went badly.
A free 30-minute working session, not a sales call. We tell you where the turn is failing and what it would take to fix it. You keep a one-page read either way: the latency budget broken into stages, the two vocabulary items recognition will get wrong, and the handoff rule we would write before anything else.