For a long time, the web ran on an implicit deal. A click came from an authenticated session, and that was enough. Enough to place an order, authorize a payment, trigger whatever the button said it would trigger. The session proved the human, the human supplied the intent, and the intent justified the action. Nobody had to define "enough" more precisely than that, because nobody needed to.
The terms of that deal are being renegotiated in public right now, and the early drafts are revealing.
In June 2026, AP reported that Visa embedded its payment network inside ChatGPT, enabling AI agents to recommend and complete purchases on behalf of users at any merchant that accepts Visa. Spending limits. Required approval steps. Approved merchants. Most early transactions would keep humans in the loop through notifications. But Visa's Jack Forestell also acknowledged the harder case: consumer intent and merchant processing are both correct, but "something happened in the middle."
That middle is where a click stops being enough. And that middle is where I've spent my whole career, watching things go wrong between "we sent it" and "it did what we wanted."
Visa's public rules from April 2026 go further. They define an "Agentic Transaction" as e-commerce undertaken by an Agentic Payment Provider on behalf of a cardholder, using cardholder-defined payment instructions, completed without direct cardholder-merchant interaction. Before acting, the provider must obtain consent, state the expiration date of the payment instruction, verify the cardholder's identity, and provision a token through Visa. After acting, it must keep order confirmations available for at least 120 days.
Read that list again. Expiration dates on payment instructions. Identity verification before token provisioning. 120-day records. Those aren't features of a click. They're features of a mandate.
Mastercard's Agent Pay, announced in April 2025, pulls a similar thread from a different angle. Agents must be registered and verified before making payments. Transactions must be recognizable as agent-facilitated to consumers, issuers, and merchants. The system needs to see and name the actor type. Not just "someone clicked." WHO clicked, and WHAT are they.
The infrastructure layer is still catching up. MCP's tool-call schema sends a tool name and arguments. No field for mandate, approval chain, actor identity, or dispute linkage. The spec recommends human-in-the-loop controls, but that's implementation guidance, not protocol structure. (I have watched this movie before. "We recommend you do the right thing" is not the same as making the wrong thing impossible.)
OAuth 2.0's actor claim gets closer. It gives you delegation semantics: this token was issued to this party acting on behalf of that party. Useful. But delegation and mandate are not the same thing. Knowing who asked is different from knowing what they were allowed to ask for, under what constraints, and what happens when it goes wrong. Delegation says "on behalf of." Mandate says "within these bounds, until this date, subject to this review."
What's emerging, unevenly and incompletely, is the outline of what "enough" likely requires when the actor isn't a person:
- Actor type. Human or agent, and if agent, registered by whom?
- Durable mandate. What was the agent authorized to do, when does that authorization expire, who granted it?
- Contestable records. Can the human reconstruct what happened, challenge it, get a remedy?
None of these systems solve the problem yet. Visa's rules are early. Mastercard's technical specs aren't fully public. MCP carries no institutional semantics at all. But their design choices, taken together, trace the shape of what a click never needed to carry.
And now does.

