Insights // Operations2026-08-2912 min read

AI Voice Agents Answer the Phone Now. What That Actually Does to Your Support Org.

Enterprise voice agents crossed from novelty to normal this year. The technology question is largely settled. The interesting questions are what happens to your escalation path, your metrics, and the people left holding the hard calls.

Varun Raj Manoharan
Varun Raj ManoharanFounder & Principal Engineer
AI Voice AgentsCustomer SupportContact Center AIAgentic AIEnterprise AI

Key takeaways

  • Voice agents work well enough now that deployment is an operations decision, not a technology bet. The failure modes have moved to handoff, metrics, and staffing.
  • The calls left for humans are harder, longer, and more emotional, which means the remaining team needs to be more senior, not less, and average handle time will get worse by design.
  • Handoff quality decides customer sentiment more than voice quality does. Making someone repeat themselves after four minutes with a bot is the thing they will remember.
  • Containment rate is the wrong headline metric. It rewards a bot that refuses to transfer, which is the single most damaging behaviour a voice agent can have.

Call a mid-sized company today and there is a decent chance an AI answers, holds a natural conversation, looks something up, and resolves your problem without a human ever joining. Two years ago that was a demo. This year it is unremarkable.

Which means the interesting questions have moved. Nobody sensible is still asking whether the voice quality is good enough or whether latency will ruin the interaction. Those got solved. What did not get solved is everything downstream of the call: who handles what the bot cannot, what your metrics mean now, and what happens to a support team whose easy work has been removed.

The metric that quietly ruins deployments

Containment rate, the share of calls resolved without transferring to a human, is the number most voice agent vendors lead with and the number I would most like to see removed from dashboards.

The problem is that it can be optimised in two ways. You can make the agent better at resolving things, which is what everyone intends. Or the agent can become reluctant to transfer, which produces the same number and is a much easier local optimum to fall into.

A customer trapped in a four-minute loop with an agent that keeps rephrasing rather than escalating is a contained call. It is also a customer who will tell people about it. I have listened to recordings of these and they are genuinely painful, because the agent is being polite and helpful and completely failing to do the one thing the person needs, which is to be handed to someone with authority.

What I would measure instead: resolution rate, meaning the caller did not call back about the same issue within a week. Transfer-with-context rate, meaning when it did escalate, the human received the full picture. And time to human for the calls that needed one, which should be short and which containment optimisation actively makes worse.

Containment is worth tracking. It is not worth targeting.

Handoff is where the customer experience is decided

If I could get one thing right in a voice agent deployment it would be this, ahead of everything about the agent itself.

When the agent transfers, the human should receive the full context: the transcript, what was already tried, the customer's account state, and what the agent believes the problem is. The human should open with something that demonstrates they have it, not with "how can I help you today?"

Being asked to repeat everything after four minutes of talking to a machine is the specific experience that turns a neutral interaction into a complaint. It is also entirely a systems integration problem, not an AI problem, which is why it gets under-resourced. The AI part is exciting and the CRM screen-pop part is not.

Two practical points. The transfer should carry a structured summary rather than a raw transcript, because a human under time pressure will not read four hundred words. And the agent should say what it is doing, by name where possible: "I am transferring you to our billing team and passing on everything we have discussed" sets an expectation that the next person can meet.

What happens to the people who are left

This is the part that is consistently under-planned and it has real consequences for the business case.

When a voice agent takes the routine calls, what remains for humans is the hard residue. Complex account situations. Angry customers who have already tried the bot. Cases with no clean answer. Anything requiring judgment, an exception, or an apology that means something.

Three consequences follow.

Average handle time goes up, and it should. The easy three-minute calls have gone. What is left takes twelve minutes. If your operations reporting treats rising AHT as a problem, you will spend six months investigating a metric that is behaving exactly as intended. Change the target before you deploy, not after the first monthly review.

The team needs to be more senior. The work is harder and the people doing it need more authority to resolve things, because a customer who has escalated past an AI has little patience for a human who also has to check with someone. This partially offsets the headcount saving, and pretending otherwise makes your business case wrong.

And the training pipeline breaks. Junior support staff historically learned the product and the customer base by handling hundreds of easy calls. That ladder has been removed. Nobody has a great answer for this yet. The best I have seen is deliberate rotation, where newer staff spend time reviewing agent transcripts and handling a curated mix rather than only the hard queue, which is slower and more expensive than the old way and better than the alternative of having no path to senior.

The realistic economics

Reports of seat reductions in the region of 10% across large enterprise support estates are consistent with what I see, and I think the honest framing is that the saving is real, meaningful, and considerably smaller than the vendor arithmetic suggests.

Where it goes: some of the saved cost reappears as more senior staff on the remaining queue. Some reappears as the ongoing engineering to maintain integrations, update the agent when products change, and monitor quality. Some reappears as the review function, because somebody has to listen to calls and check the agent is still behaving.

Where the value shows up that the naive model misses: after-hours coverage that used to be unavailable or outsourced. Absorbing volume spikes without a hiring cycle. Consistency, since the agent does not have a bad Tuesday. And speed to answer, which affects satisfaction more than most support leaders' models account for.

My rough guidance for a business case: assume 50 to 65% of call volume can be handled without a human within a year, assume you keep 40 to 50% of the headcount you would have modelled removing, and count the coverage and speed benefits separately rather than folding them into the labour saving. That produces a case that survives the first annual review, which is more than most of them do.

The failure modes worth designing against

A few specific things that go wrong, from deployments I have watched.

The agent confidently gives wrong policy information. This is the most damaging single failure because it creates a commitment the company then has to either honour or retract. The mitigation is to make policy answers retrieval-backed against a maintained source rather than something the model produces from training, and to have the agent decline rather than improvise when the retrieval comes back empty.

The agent cannot recognise distress. A caller who is upset, or whose situation is serious, needs a human quickly regardless of whether their query is technically resolvable. Sentiment-triggered escalation is a small feature with a large effect on the incidents that end up in front of an executive.

Accents and speech patterns that degrade recognition. This is better than it was and it is not solved, and the failures are not evenly distributed across your customer base, which makes it a fairness issue rather than just a quality one. Test with real recordings from your actual customers, across the range of how they actually speak, not with clean studio audio.

Product changes that the agent does not know about. A pricing change ships, marketing updates the website, and the voice agent tells customers the old price for three weeks. Whatever your product change process is, the agent's knowledge source needs to be in it.

Where I would start

Not with the whole queue. Pick one call type: order status, appointment scheduling, balance enquiries, password resets. Something high-volume, low-emotion, and with a clean definition of resolved.

Run it in parallel first, with the agent handling calls that opt in or that arrive outside staffed hours. That gives you real recordings to evaluate without risking your main volume.

Fix the handoff before you widen scope. If the transfer experience is not good, widening scope just increases the number of people having a bad one.

Then expand by call type, one at a time, with the emotional and financially consequential categories last or never. Some call types should stay with humans permanently, and deciding that deliberately is a sign of a well-run deployment rather than a failure of ambition.

The thing I keep coming back to

Voice agents are now good enough that the technology is not the interesting part of the project. What determines whether a deployment is a success is the operational design around it: what gets escalated and how fast, what the human receives when it does, what you measure, and whether you were honest about what happens to the team.

That is not a satisfying conclusion for anyone hoping the hard part was the AI. It is consistently what I see, and it is also good news, because operational design is a problem your organisation already knows how to solve.

We build voice agents and, more often, we get called in when one is live and the operations around it are not working. If either is where you are, let's talk it through.

Available for new projects

Let's build something great.

Have a project in mind? We are an elite software and AI development studio ready to bring your ideas to production. Let's talk about your roadmap.

See our work