When a Chinese telecom company serving more than ten million users deployed a conversational agent in late 2018, the researchers tracking the rollout recorded a result that looks at first like a contradiction. Daily call volume to human agents didn't change significantly. But the average call ran 7.23 seconds longer, and during peak hours as much as 26 seconds longer. The automated system was resolving cases. The human operation was getting slower.
The explanation is almost mechanical. Once you take password resets, balance checks, and tracking inquiries out of a queue, what remains is a concentrated residue of ambiguity, policy judgment, and situations that match no template. Average difficulty rises not because new problems appeared but because the easy cases that used to dilute the queue are gone. A team that once alternated between routine lookups and complex disputes now handles disputes all day. The total may hold roughly steady while the mix underneath it changes completely — and the mix is the part nobody is counting.
That change has consequences for the people still working the queue, and they are not subtle to anyone standing on the floor. A field study of 531 agents at an Italian telecom compared two groups: one handling technical problems and complaints, the other providing directory information. The technical group reported substantially higher customer aggression and more emotional dissonance — the sustained work of managing your own reactions while absorbing someone else's frustration. Automating the easy work collapses those two groups into one. Whatever variety and cognitive pacing the old mix provided is replaced by a steady diet of the hardest cases the organization encounters. Whether that produces measurable "decision fatigue" is unsettled; a 2025 systematic review found only 45% of 130 quantitative tests supported the hypothesis. The specific framework isn't necessary. The well-documented association between sustained difficult service work and exhaustion, error, and absence is enough.
The metric organizations report is structured to miss all of it. Klarna's 2025 annual report says its assistant handled 80% of chats — 31 million conversations — with no drop in consumer satisfaction. The report does not publish the difficulty distribution, handle time, error rate, or turnover for the remaining 20%. Commonwealth Bank announced 45 redundancies after deploying an AI voice system, then reversed them within a month as call volumes rose and team leaders went back on the phones. In both cases the composition of the remaining human work was the variable nobody had tracked. It can be tracked: an Amazon research team has demonstrated that contact complexity can be scored independently of volume or handle time, validated against senior-agent judgments. The measurement exists; the incentive to use it mostly doesn't, because automation rate justifies the spending and difficulty composition complicates the story.
Automation doesn't remove a proportional slice of the work. It removes the predictable middle and leaves the rest — and the rest is not a scaled-down version of the whole.
Support queues make this visible earlier than most functions do, because the residue is a person on the phone rather than a slower quarterly number. A team reporting 80% automation may be running a human operation that is functionally harder than the one it replaced, not because the system failed but because it succeeded at precisely the work that made the queue survivable. Automation rate is the easiest number to produce and among the least informative. What happened to the people holding everything the system couldn't close is harder to measure, and it is the number that would tell you whether the deployment worked.
-
Escalation type changes outcomes: A 2026 randomized field experiment on a large e-commerce platform found that emotional escalations from AI agents took 40.8% longer and produced worse ratings than matched human-handled chats, while technical escalations preserved service quality but still ran 19.1% slower.
-
Scoring complexity independently: Amazon researchers built a contact-complexity measure from 450,000 transcripts and 152 issue codes, combining transcript length, model uncertainty, and classification difficulty into a score that aligned with senior agents' own complexity judgments once it passed a 0.6 threshold.
-
Production agents are shorter-loop than expected: A peer-reviewed study of 86 deployed agent systems found that 68% executed ten steps or fewer before human intervention and 80% used structured workflows rather than open-ended planning, suggesting that deliberate constraint is the norm in production, not the exception.
-
Staffing models miss the residual: Commonwealth Bank reversed 45 AI-related job cuts within one month as call volumes rose and team leaders returned to phones, after the bank acknowledged it had not adequately considered all relevant business factors in its workforce reduction.

