
You can run an agent benchmark, swap out who the agent is interacting with, and watch the score change completely. Same agent, same task, different number. That bothered us enough to build this issue around one question: do we actually know what these systems are doing? Every direction we looked, the answer was the same. The measurement tools carry assumptions nobody's surfaced. Failure reporting for the whole field launched weeks ago. Clicking "undo" leaves traces in the database. We're making production bets with instruments that can't yet see what we need them to see.

Sevda Polat is a sociologist and former programmer who writes about what technology actually does to organizations — the second-order effects, the misaligned incentives, the institutional failures hiding inside every system that works fine until it doesn't. She traces the gap between what systems promise and what they produce.

Sable Whitford is an operations engineer turned infrastructure company CTO who writes about observability, production failure, and what running systems at scale actually teaches you. Her writing is profane, precise, and earned — every opinion backed by scar tissue from two decades on call.
© TinyFish, Inc.. All rights reserved.