
Recent Activity
September — Issue #46

How Nubank's agent architecture puts infrastructure ahead of model intelligence, and where the paper's evidence stops short of the properties it requires.

Agents can log what they did on the open web but can't establish what they didn't do, and that gap won't close with better instrumentation.

Oversight has become the only frame organizations have for human-agent collaboration. The more interesting shift is humans doing work that wasn't previously possible at all.

Capability-based security lost to identity-based access control in the 1970s, but agent systems keep rediscovering why it was right.

How Unix permission bits designed for people at terminals became container security, and why agents expose the missing concept of mandate.

Java's applet sandbox eroded from product pressure rather than attacks — its own designers loosened containment so the code could do its job. That pattern is recognizable in agent containment today.
September — Issue #45

TrickyArena's evaluation results show that the most capable web agents are also the most susceptible to deceptive interface patterns, and why.

Agent adoption is creating a demanding operational discipline that organizations keep mistaking for freed-up time that should go to strategy.

Agent deployments create growing correction costs that hide inside existing budgets, which may explain why task-level AI gains aren't reaching the bottom line.
August — Issue #44

Agent autonomy should be measured by irreversible consequence between enforceable checkpoints, not runtime duration or the presence of a human reviewer.

Traces a booking sequence through MCP's July 2026 revision to show which reversibility decisions the protocol now leaves in your lap.

"Undo" compresses several distinct remedies with different timescales; conservation ethics and payment regulation show why the compression creates blind spots.
August — Issue #43

Debug logging was built for engineers arriving after a crash. Agent audit trails need it to reconstruct decisions nobody witnessed.

How Venetian merchants built books that caught their own errors, and what digital systems gave up by betting on prevention instead.

Agent systems specify "human in the loop" as policy but almost never engineer the window in which that human can still change the outcome.

Agent infrastructure acquisitions encode competing theories about what agents are, and the theory you inherit decides which failures you catch early and which ones you meet late.
August — Issue #42

Firm-level AI adoption figures obscure whether the gap between task speedups and organizational productivity is temporary or structural.

Organizations automate exceptions into rules without tracking the judgment spent to create them, depleting a resource nobody accounts for.

Why dollar value is the wrong axis for agent delegation when cancellation windows and information propagation determine real cost.

Agent risk has less to do with dollar amounts than with whether an action's consequences survive the undo button.
August — Issue #41

An agent can clear every authorization check and still choose badly for you. Fidelity has no instrument that survives a real transaction, and the buyer eats the difference.

Publishing structured data costs you nothing but effort. Getting invoked is somebody else's decision — and the leverage may sit with the thinnest registry, not the richest protocol.

Something is reading your site literally now. When declared data becomes a commitment, drifting feeds, drip pricing and value that only shows up in conversation all start to vanish from view.

What the signature actually covers, why a verified request proves less than it looks like it proves, and which parts of the problem — mandate, intent, replay, failure handling — the spec deliberately hands back to you.

Database recovery, then agent traces, then practitioners: three instruments, and what each one was structurally incapable of hearing.

Ten-step agent loops are backpressure. The consumer downstream is a human who has to defend the output — and that's where the real ceiling sits.
August — Issue #40

Agent traffic can now prove who's calling and never what it was sent to do, and no amount of logging gets that back.

A clerk's two-letter shorthand for borrowed authority, and how it dumped the burden of asking onto whoever was holding the paper.

AI is eating the boring work that used to build engineers. The evidence points at the pipeline, not at the experts already in it.

Maxim Fateev built durable execution. What his design choices quietly admit is that rollback doesn't fail so much as decay.

Name your agent and you buy access, and you also open a reputation file you can't fully control. Why builders are splitting agents by purpose for reasons that have nothing to do with engineering.
July — Issue #39

Agent runs need a third artifact beyond traces and outputs, built for the stakeholder who arrives later asking what was authorized and why.

Each element of an enterprise SQL evidence bundle earns its place by closing a gap the previous layer left exposed.

A CNRS fingerprinting researcher's latest work reveals that production websites are becoming active counterparties, classifying every agent visit with near-perfect accuracy regardless of whether the agent completed its task.

Browser automation perfected the mechanics of clicking. The layers that determine whether an automated action actually counts remain largely unbuilt.
July — Issue #38

Agent outputs use specificity as an alibi against scrutiny. The habit worth building: when the output looks clean, ask what the system actually proved.

AI output that looks specific before its meaning is earned is a distinct failure mode worth naming, because specificity discourages exactly the scrutiny it deserves.

Payment networks and delegation protocols are revealing what "enough" has to mean when the thing doing the clicking isn't a person anymore.

Browser automation spent twenty years perfecting click delivery while a human silently carried the accountability no spec ever had to name.

The click compressed visibility, intent, and accountability into one gesture because a human body was assumed behind it. Agents inherit the gesture, not the body.
