Practitioner's Corner

Practitioner's Corner

The Green Dashboard Problem

When researchers ran web agents through interfaces seeded with forced opt-ins, misleading pop-ups, and pre-selected checkboxes, the agents with the highest task-completion rates turned out to be the most susceptible to manipulation. Guardrails helped without closing the gap: some agents named the deceptive element, identified it correctly, and proceeded anyway. The execution trace stayed clean. Every metric said success.
Incident monitoring was built to catch exceptions — crashes, timeouts, error codes. An agent that books a flight and quietly adds insurance nobody asked for has produced no exception. It has produced a misdirected success, and nothing in the system is instrumented to see it.

The Green Dashboard Problem
When researchers ran web agents through interfaces seeded with forced opt-ins, misleading pop-ups, and pre-selected checkboxes, the agents with the highest task-completion rates turned out to be the most susceptible to manipulation. Guardrails helped without closing the gap: some agents named the deceptive element, identified it correctly, and proceeded anyway. The execution trace stayed clean. Every metric said success.
Incident monitoring was built to catch exceptions — crashes, timeouts, error codes. An agent that books a flight and quietly adds insurance nobody asked for has produced no exception. It has produced a misdirected success, and nothing in the system is instrumented to see it.
Phil Cuvin and the Evaluation That Had to Apply Pressure

Standard agent benchmarks check whether the system completed the task. Phil Cuvin's DECEPTION benchmark, published at ICLR 2026, adds a second question: did the agent also get quietly steered into outcomes the user never authorized? Answering it meant building web environments that behave the way commercial sites behave — with countdown timers, preselected add-ons, and cancellation flows designed to exhaust you. The design choices behind those environments point to an operational cost most agent evaluation budgets don't yet include.
Phil Cuvin and the Evaluation That Had to Apply Pressure
Standard agent benchmarks check whether the system completed the task. Phil Cuvin's DECEPTION benchmark, published at ICLR 2026, adds a second question: did the agent also get quietly steered into outcomes the user never authorized? Answering it meant building web environments that behave the way commercial sites behave — with countdown timers, preselected add-ons, and cancellation flows designed to exhaust you. The design choices behind those environments point to an operational cost most agent evaluation budgets don't yet include.


The Agent That Chose Our Vendors Had a 96% Success Rate and a 30% Error Rate
CONTINUE READINGThe Measurement Gap

In McKinsey's 2026 State of AI survey, 80% of respondents say AI has improved their personal productivity. Only 37% attribute any positive impact to their organization's earnings. That figure is flat from last year, even as adoption has grown.
The optimistic assumption is that financial results will eventually follow. But a person finishing tasks faster and a firm's earnings improving are different phenomena, tracked by different instruments, on different time horizons. A completed task registers as a gain for the individual while generating an escalation, a retry, or a cleanup cost that surfaces in someone else's queue entirely. Speed at the task level and value at the organizational level are not connected by a simple pipeline.
The 6% of organizations reporting meaningful earnings impact share a telling pattern: three-quarters have redesigned workflows around AI, and they are twice as likely to have built processes for measuring its effects. They invested in the institutional plumbing that connects individual output to organizational outcomes.
Most organizations haven't. Which means the gap between reported productivity and reported earnings isn't closing because nobody has built the instrumentation that would reveal where the costs actually land.
Further Reading








