In July 1998, a researcher at State Farm sent the National Highway Traffic Safety Administration a pattern he had found in insurance claim files: twenty-one reports involving Firestone ATX tires, fourteen of them on older Ford Explorers. Against the number of those tires on American roads, twenty-one is nothing. NHTSA did not open an investigation.
Warranty claims on one tire line had already risen sixfold between 1995 and 1997. But the industry's standard performance measure, historical adjustment data, kept returning reassuring answers. The monitoring system was working. It was answering the question it had been built to answer. The claims were answering a different question, one nobody had formally asked.
NHTSA opened its investigation in May 2000, twenty-one months after State Farm's notice. Firestone recalled 14.4 million tires that August.
The case predates AI by decades, which is part of why it is useful. It isolates a dynamic that gets sharper as automation expands. When a system handles the predictable middle of a distribution efficiently, the cases that don't fit become the main channel through which an organization finds out that its assumptions have drifted from reality. Dashboards measure performance against categories somebody already chose. Exception volume, resolution time, cost per case: these tell you how well you are managing the unexpected inside the frame you already have. The exceptions worth attention are the ones indicating the frame has stopped describing the world.
A fraud pattern the model was never trained on arrives this way. So does a customer population whose needs have moved somewhere the segmentation can't reach, or a regulatory change that turns a routine approval into a consequential one. None of it registers as performance degradation. It arrives as cases that don't fit, worked by people who may or may not recognize what they are looking at.
Separating signal from noise in exception data is genuinely hard, but the Firestone pattern had features that, in hindsight, distinguished it from random variation. The claims clustered by product configuration, specific tire lines on specific vehicles. They grew over time instead of appearing as isolated spikes. And they diverged from the metric the industry was actually using to judge safety. Any one of those, on its own, might mean nothing. Together they described a shift the monitoring categories had no way to express.
Noticing that required someone treating claims as evidence about the world rather than as cases to close. Every formal method for detecting anomalies rests on prior choices about what to measure and where the boundary sits between expected variation and something worth investigating. The adjustment data functioned exactly as designed. It had been designed around a question that did not capture the failure mode actually occurring.
Current banking model-risk guidance takes some of this seriously, calling for investigation of persistent deviations and for monitoring changes in clients, exposures, and market conditions. That is a real step. It applies to models, though, not to the broader organizational question of whether exception data reaches anyone with the authority and the context to act on what it shows.
As automation absorbs more routine decisions, the volume of decisions rises while the number of people who see any individual case falls. I wrote in an earlier piece about what happens when automated production outruns an organization's capacity to absorb corrections. The risk here is quieter. It isn't that corrections pile up. It's that the cases carrying real information about a changed environment get processed and cleared without anyone registering what was in them. An organization can run its exception queue efficiently for years and, in the process, close off the one route by which the news would have arrived.

