Vision

Vision

The Credence Problem

When you send an agent to do something you cannot do yourself, you also cannot check whether it did it correctly. That is not a defect in the design. It is the same condition that made the agent worth using. Economists identified this structure decades ago in auto repair and medical care: goods whose quality the buyer cannot assess even after consumption. A growing share of agent output belongs in that category, and the research on what happens when people try to compensate through transparency suggests the problem does not yield to the fixes we reach for first.
The Credence Problem
When you send an agent to do something you cannot do yourself, you also cannot check whether it did it correctly. That is not a defect in the design. It is the same condition that made the agent worth using. Economists identified this structure decades ago in auto repair and medical care: goods whose quality the buyer cannot assess even after consumption. A growing share of agent output belongs in that category, and the research on what happens when people try to compensate through transparency suggests the problem does not yield to the fixes we reach for first.

Two Cases

The Users Who Got There First
Blind and low-vision users have spent years depending on AI descriptions they cannot independently verify. They've developed sophisticated checking strategies — and new research shows those strategies still miss confident wrong answers at alarming rates. What this community has learned about trusting an AI intermediary, and where that trust still breaks down, matters for every organization now deploying AI as its default first reader.

When Nobody Reads the Original
When an AI summarizes a contract, a security incident, or a recorded engineering conversation, the summary becomes the event for everyone downstream. Enterprise organizations are acquiring the same dependency on unverifiable AI accounts that the accessibility community has navigated for years — but without the verification instincts. And the expertise needed to catch errors is quietly thinning as the AI displaces the practice that built it.

Standards Sidebar

The W3C's accessibility guidelines encode a principle that agent builders are arriving at from the other direction: the adequacy of a description depends on what someone needs it for. A photo of a bird on a birding site needs species and plumage. The same photo on a parks page needs "bird at the lake." A printer icon used as a button should say "print," not "image of a printer."
The same logic applies to any agent-generated output. A contract summary can be thorough, well-sourced, and coherent while missing what the reader actually needed. Were they assessing risk? Checking dates? Looking for a reason to walk away? Thoroughness is not the same as adequacy, and the source material alone cannot tell you which details matter.
The W3C addressed this by requiring a human author who understands the context to decide what counts. For agents, the interesting question is what happens when that contextual judgment is precisely the thing being automated. Most agent systems today optimize for plausible completeness because completeness is measurable and purpose-fitness is not. The incentive structure rewards looking thorough over being useful.
Further Reading








