Foundations

Foundations

What the Guardian Sees

The OWASP Agent Control Standard ships with two demo scenarios: an agent tries to delete a filesystem and gets blocked, and an agent asks for an entire database table and gets a narrower query back. Both look like the control system saying no, with the second saying it more politely. But refusing an action and rewriting one draw on different things. Refusal needs recognition. Rewriting needs a correct replacement, which means the system has to know something about what the agent should have asked for — and where that knowledge comes from is the whole architecture.

What the Guardian Sees
The OWASP Agent Control Standard ships with two demo scenarios: an agent tries to delete a filesystem and gets blocked, and an agent asks for an entire database table and gets a narrower query back. Both look like the control system saying no, with the second saying it more politely. But refusing an action and rewriting one draw on different things. Refusal needs recognition. Rewriting needs a correct replacement, which means the system has to know something about what the agent should have asked for — and where that knowledge comes from is the whole architecture.
When Human Oversight Works Against You

Adding a human checkpoint to an agent workflow feels like adding safety. But a meta-analysis of 106 experiments found that human-AI teams performed worse than the better solo performer on decision tasks, and in more than 95% of those setups the human held the final call. Research on clinical alerts and security warnings deepens the problem: review accuracy falls sharply under repetition, and production agent queues manufacture repetition by design. Most teams have never tested whether their approval gates are still functioning as judgment.
When Human Oversight Works Against You
Adding a human checkpoint to an agent workflow feels like adding safety. But a meta-analysis of 106 experiments found that human-AI teams performed worse than the better solo performer on decision tasks, and in more than 95% of those setups the human held the final call. Research on clinical alerts and security warnings deepens the problem: review accuracy falls sharply under repetition, and production agent queues manufacture repetition by design. Most teams have never tested whether their approval gates are still functioning as judgment.

Decision Architecture Reading








