I spent forty minutes last month arguing with a chatbot about a refund that should have taken four minutes with an actual human, and somewhere around minute thirty I realized I was typing in short, angry sentences like the bot would somehow understand my frustration better if I used fewer words. It didn’t. It just kept offering me a discount code instead of the refund I asked for three separate times.
That experience is the one most people bring up when this topic comes up, and fair enough, it’s the one that sticks. But it’s not really the whole picture anymore, and a friend of mine who runs support for a small e-commerce brand set me straight on that pretty quickly when I complained to her about it.
Her team switched to an AI-handled first line of support about a year ago, mostly out of necessity, not excitement. They were drowning in repetitive tickets – order status, return policy questions, the same five things over and over, and hiring more people wasn’t really in the budget. The bot picked up that slack almost immediately. Order status checks that used to take a human two minutes to look up now resolve instantly, no waiting, no queue. Nobody misses answering “where’s my package” for the two-hundredth time that week.
Where it falls apart is exactly where mine did. Anything that requires judgment, an exception to policy, a genuinely upset customer who needs to feel heard before they’ll accept any resolution, the bot handles badly or not at all. My friend’s team learned this the hard way early on, when a handful of angry customers escalated to social media because the bot kept looping them through the same three scripted responses without ever actually solving anything. That’s expensive in a way that’s hard to put a number on. An annoyed customer complaining publicly does more damage than the cost of just paying a human to handle the call properly in the first place.
So they built in an escalation trigger, admittedly after the fact rather than from the start, where certain keywords or a repeated unresolved question automatically bumps someone to a real person. Sounds obvious in hindsight. It wasn’t obvious when they were in the middle of rolling the thing out and mostly just excited about the ticket volume numbers going down.
The tone problem is the part nobody really prepares for either. A bot can be polite without being warm, and customers notice the difference even if they can’t quite articulate it. My friend described getting feedback that the chatbot felt “fine but cold,” which is a strange thing to try to fix with code. They spent weeks tweaking response phrasing, adding small acknowledgments, little things like confirming frustration before jumping straight to a solution. It helped some. It didn’t fully close the gap. There’s apparently something about knowing you’re talking to an actual person that changes how forgiving customers are willing to be, even when the actual information given is identical.
Staffing changed in a way that surprised her too. She didn’t end up with a smaller team, which is what everyone assumed going in. She ended up with a differently skilled one. Fewer people were doing repetitive lookups, and more people were trained specifically to handle escalations well, since those are now the only calls that actually reach a human, and they tend to be the hardest ones by default. Training for that is a different skill than training someone to read a script quickly. It took longer to hire for than she expected.
Cost savings were real, just smaller than the sales pitch implied going in. The bot handles volume cheaply, no argument there. But the harder tickets now take more time per case than they used to, because the person handling them is dealing exclusively with the complicated, emotionally charged stuff that used to get diluted across a mixed queue of easy and hard calls. Net savings exist. They’re just not the dramatic number that gets quoted in a vendor’s demo deck.
I still think about that forty-minute refund argument sometimes, mostly because it’s a decent reminder that the technology works great right up until it doesn’t, and the gap between those two states isn’t always obvious from the outside. My friend’s honest take, after a year of living with this, is that the bot earned its place handling the boring 80% of tickets. The other 20% is still very much a people problem, and probably always will be.
Funny enough, she told me the bot actually got better after they stopped trying to make it handle everything. Scaling back what it was allowed to touch, and being upfront in the chat window that a human is one message away if things get complicated, did more for customer satisfaction than any script rewrite ever did. Nobody minds talking to a bot first, it turns out, as long as it’s honest about its own limits instead of pretending to be something it’s not.

