The Census Bureau has been asking firms what they actually do with generative AI, and the answers are narrower than the conversation around them. Among firms whose workers use it, 85% report use for writing or editing documents. About half use it for search or for interpreting documents. And 57% of adopting firms have it running in three or fewer business functions. Text production, search, document interpretation — the places where language models have an obvious foothold.
Automation doesn't shrink an organization's work evenly. It takes the parts that match a template and leaves the rest.
The rest is not a scaled-down version of the whole. Every organization's work runs from the standard and repeatable to the cases that force someone to stop and reason: the invoice that doesn't reconcile, the customer situation nobody wrote a policy for, the edge case that has to be argued about before it can be resolved. Routine and exception have always sat side by side, in the same budgets, the same roles, the same afternoons.
Pull out the routine end and what remains becomes denser. Each exception now costs more relative to the automated baseline around it. A peer-reviewed invoice-processing study shows the dynamic in miniature: automate extraction and most of the effort disappears, but the human work left over is disproportionately the hard cases, corrections running several times longer than the automated norm. Nobody has a reliable cross-industry figure for what exception handling costs as a share of operations, and that absence tells you something. The two kinds of work have historically shared a budget line. The concentration only becomes visible once automation pulls them apart.
That's the cost side. A second concentration is happening at the same time and is harder to see. Routine work was also a curriculum. Processing standard invoices taught you what a non-standard one looked like; handling ordinary requests built the pattern recognition you drew on when something extraordinary arrived. Learning used to be distributed across the work, an unbilled input organizations consumed without ever tracking it. As the routine becomes infrastructure, that curriculum thins out. The exceptions stay. The preparation for them erodes.
So exceptions end up holding both: the most expensive manual work left, and the last place where anyone routinely meets a situation the existing rules don't cover.
Two findings sharpen this. A meta-analysis of 106 experiments on human-AI collaboration found that combined teams tend to underperform whichever performer, human or machine, would have done better working alone. Coordination overhead, miscalibrated trust, unclear boundaries between who is responsible for what — these cost something. Exception work is where you'd most want a human and a system working together, and where the division of responsibility is least clear. The collaboration penalty falls hardest on the work that most needs collaboration. A separate preprint found heavy AI users increasing their individual productivity actions by 21% while their communication actions rose only 7%. Capacity to produce is outrunning capacity to coordinate. And exceptions are, close to by definition, what happens when one person's or one system's scope runs out and someone else has to be brought in.
Organizations need exception work to be efficient, because that's where the remaining costs collect. They also need it to be generative, because that's where the remaining learning happens. These are different demands on the same work, and they pull against each other.
Efficiency wants exceptions closed quickly and handled the same way twice. Learning needs someone to sit in the ambiguity long enough to notice what's actually new about it.
Which leaves a question I can't answer. When the routine becomes infrastructure — reliable, invisible, no longer a place where people practice anything — where does an organization's capacity to learn come from? I don't think anyone knows. The question is only now becoming askable.
-
Collaboration design is underexplored: The meta-analysis finding that human-AI teams underperform the better solo performer drew on experiments where more than 95% left the final decision with a human, while only three tested a predetermined division of subtasks — suggesting the field has barely begun designing collaboration around comparative strengths.
-
Learning gains depend on use: A randomized experiment found that students who used generative AI for explanations showed persistent knowledge gains a week later, while those who used it to generate text saw quality gains disappear when the tool was removed — a distinction that matters for how organizations structure exception work as training.
-
AI adoption remains shallow: The Census working paper found that among firms using AI, nearly 65% limited generative AI to three or fewer worker-task categories, and comprehensive adopters represented just 4% of functional AI users — meaning the routine-exception concentration described here is still in early stages.
-
Production outpacing coordination: The workplace preprint measuring Microsoft 365 activity associated heavy AI use with a 21% rise in document-creation actions but only a 7% rise in communication actions, a gap the authors warn could weaken the diffusion of diverse information across organizations.

