The week after Labor Day, half the organizations I talk to are still catching up on what their agents did while everyone was at the lake. Which tells you most of what you need to know about where enterprise agent deployment has landed: agents do things, humans check the things. The checking is the relationship.
That's roughly right as description, and I've said so myself. In "The Work That Isn't Strategy" earlier this year I argued that capacity freed up by agents doesn't drift toward creative or strategic work on its own. It drifts toward writing specs, fixing output, and deciding whether to accept what came back. Sevda Polat named the broader version "review-shaped work". Nobody made this up.
What bothers me is that it has hardened into the only frame available. Oversight shapes the tooling, the budget line, the conversation. That's a ceiling.
Two recent studies are worth sitting with, both designed as research rather than as vendor case studies.
A preregistered field experiment at Procter & Gamble put nearly 800 professionals through controlled conditions and found that individuals working with generative AI reasoned across functional boundaries that normally required pairing a commercial professional with an R&D specialist. One person, with AI assistance, produced ideas combining technical and market thinking at a level that had previously taken a cross-functional team. The AI didn't make them faster at the job they already had. It widened what a single person could hold at once.
A two-year qualitative study of perfumers at a fragrance company found something comparable in a domain that could hardly be further away. A perfumer takes a vague brief, has some sense in the nose of what it should become, and feeds that into a generative system. What comes back is a set of alternatives, and the alternatives change what the perfumer thought they were after. Then they translate, recombine, redirect. The human contribution wasn't approval. It was an evolving sense of what "right" meant, revised in contact with possibilities nobody had considered at the start. The search space got wider and the judgment steering through it stayed human.
I don't want to oversell this. A randomized experiment across 66 firms found that AI tools gave workers back about two hours a week on email with no measurable change in how they spent their days. Anyone who has ever automated a painful runbook knows this feeling. The time comes back and gets eaten by whatever was already in the queue. Free hours are not different work. Cappelli, Tambe, and Jiang put it precisely in their Annual Review article this August:
"The destination of freed capacity is a design choice — it depends on which parts of jobs get transferred, how remaining work is structured, what people are expected and enabled to do with what's left."
The P&G result wasn't luck. Someone built the conditions for it: AI assistance, problems structured to reward range, explicit permission to reason outside your function. The perfumers got there through two years of deliberate integration into how the creative work already ran. In both cases somebody had to design for a kind of work that didn't exist before the system showed up.
Most organizations are not doing that. They're buying review infrastructure, because that's the frame they have, and because a review queue is something you can put on a budget line while a boundary-spanning capability is not. Meanwhile the work these studies found goes unnamed inside the org, which means it also goes unstaffed.
Your agents need oversight. Fine. The question I'd be asking this week is whether anyone in the building is designing for what agents make newly possible, and whether they have any language for it yet.
- Where saved time goes: A Danish study linking surveys and administrative records for roughly 25,000 workers found that new AI-related tasks split across content generation, output review, and workflow integration — all three growing together rather than oversight giving way to creative work.
- Production agents stay short: A peer-reviewed ICML study of 20 deployed agent systems found that 68% executed no more than ten steps before human intervention, with most relying on simple architectures and human evaluation rather than extended autonomy.
- Human-AI teams underperform expectations: A meta-analysis of 106 experiments found that human-AI combinations beat humans alone on average but performed worse than whichever party — human or AI — was already better at the task, suggesting that adding a checkpoint is not automatically an improvement.
- Workers want more agency than experts prescribe: A Stanford audit across 1,500 workers and 104 occupations found that "equal partnership" was the most desired collaboration level in 47 occupations, with workers preferring more human involvement than technical experts considered necessary on nearly half of tasks evaluated.

