Vision

Vision

The Work That Oversight Can't See

Most organizations have built their entire agent strategy around one relationship: agents do things, humans check the things. That's necessary, and it describes what most deployments actually look like. But two recent field studies — one at Procter & Gamble, one following perfumers over two years — turned up something else. People doing work that wasn't available to them before: cross-functional reasoning a single role couldn't hold, exploratory search through possibilities nobody had considered. Almost nobody is designing for it.

The Work That Oversight Can't See
Most organizations have built their entire agent strategy around one relationship: agents do things, humans check the things. That's necessary, and it describes what most deployments actually look like. But two recent field studies — one at Procter & Gamble, one following perfumers over two years — turned up something else. People doing work that wasn't available to them before: cross-functional reasoning a single role couldn't hold, exploratory search through possibilities nobody had considered. Almost nobody is designing for it.
Two Modes Mapped

What the Documents Should Become
Agents can produce the competitive scan, the feasibility memo, the positioning draft, each adequate on its own terms. Together they don't yet say anything. The emerging skill is deciding what these outputs should become — for whom, and in what order. Research across hundreds of thousands of enterprise AI conversations suggests this assembly work already takes up more collaboration time than generation does.

Affordable Curiosity
When running an analysis cost a person two days, you ran the one you were fairly sure you needed. When it costs twelve minutes, you can afford the one you're merely curious about. Cheap agent execution is lowering the threshold for justifiable curiosity across enterprise teams. But wider exploration doesn't automatically improve decisions, and the constraint that replaces affordability turns out to be navigation.

A Documentary Editor on Why the Best Moment in 500 Hours of Footage Never Announces Itself
CONTINUE READINGEvidence Note

The week after Labor Day is a fair moment to reconsider what "productive" means. A Stanford study of nearly 250,000 AI conversations found that high-stakes tasks averaged 12.4 turns — almost double the 6.5 for routine work. As stakes rose, users modified AI output more, accepted it unedited less, and challenged the model's reasoning successfully over 80% of the time. People used the exchange to develop their own thinking.
Enterprise AI is still evaluated mostly by time saved. For routine tasks, that captures real value. But the most consequential collaboration — longer sessions, more friction, more human intervention — registers, by that metric, as inefficiency. What organizations choose to measure will determine what kind of human-AI work survives.
Further Reading








