Three moves before Monday.
- Use their words. Retire yours. The names for agent trust have stopped being concepts and become shipped components:
Confirmation,Checkpoint,Sources,Reasoning,Plan,Queue,Task,Branch Picker,Error, and as of Thursday,elicitation. Say those. Drop anything you coined yourself. - Count repetitions, not eloquence. Ten companies swept, and not one dated public post from a design leader at any of them in July. Postings are your only live channel this month. Postings are wish lists run through HR. So frequency across several of them is the only reliability check you have.
- Altana first. One requisition, three demands in the same document: control architecture, confidence design, a machine-readable design system. Your strongest published outcome and your largest gap sit in the same paragraph of it.
What altitude means for prep
One loop, three apertures. Craft evaluators inspect states. System evaluators inspect controls. Leadership evaluators inspect whatever keeps working after you leave the room. You lose by carrying the craft answer into the leadership conversation. So rehearse the mapping out loud before you walk in: Trust sits at system altitude, and it carries the most live mandates this cycle. Patterns splits, control architecture at system, component states at craft. Iteration is craft. Execution is leadership. Same aperture logic as the buyer-angle map, pointed here at a cycle of published evidence rather than at a buyer's anxiety.
The hole in this cycle's sourcing
Two windows, set at different lengths on purpose.
- 30 days for hiring commentary. Mandate temperature decays fast.
- 90 days for practitioner writing and pattern libraries. Publication cadence there is slower, and a May essay still holds.
The sweep ran July 2 to August 1 across named design leaders at Altana, HackerOne, Vanta, Amplitude, Ramp, Front, TRM Labs, Maven, Gusto, and Stripe. Nothing qualified. Here is the nearest usable material, all of it outside the window:
- Ramp VP of Design Diego Zaks, June 3: the product has to work out when a person needs to step back in, and what they need to see to trust the agent.
- Stripe's Courtney Drake, April 30, hiring for "agent-driven workflows while preserving trust and control."
- Gusto CDO Amy Thibodeau, June 11, on an AI-first team that keeps human taste and judgment.
Abridge, Ambience, Spring Health, and Suno were not swept. Three of those are posting-only mandates this cycle. Suno is the exception worth carrying: its CPO promoted the search publicly about a month ago, the only executive-level confirmation of hiring intent in the whole set, and it lands outside my 30-day line.
Company blogs shipped plenty in the window, none of it design-leader commentary. Linear shipped agent-assisted editing on July 23, with agent changes visually marked, the agent preserved as author, checkpoints, restore to an earlier version. Vercel published in June on teaching agents product design. Useful vocabulary, but product and engineering writing. Figma's Config archive, Stripe Design, and OpenAI's blog returned nothing on supervision craft in-window. Which leaves mandate temperature to the postings and vocabulary to the release notes. On the trust model itself, the two usable company sources are Anthropic and Notion, and both are research and product posts rather than design writing.
Confidence: high on those three commentary dates. Confidence: low-to-moderate on the wider silence, since platform indexing is patchy and private commentary is invisible to me.
So count words instead. Across the captured postings:
| Phrase cluster | Postings | Where |
|---|---|---|
| Human oversight, review, escalation, judgment | 6 | HackerOne, Gusto, Maven ×2, Altana, Abridge |
| Design systems and reusable AI patterns | 5 | Vanta, HackerOne, Altana, Spring Health, Maven |
| Confidence, evidence, uncertainty | 4 | HackerOne, Altana, Abridge, Maven |
| Hands-on production, code, working demo | 4 | Amplitude, Ramp, TRM Labs, Altana |
| Definitions of done, launch reviews, quality frameworks | 3 | Vanta, Amplitude, Spring Health |
| "player-coach," verbatim | 3 | Amplitude, Ramp, TRM Labs |
Nobody is buying "AI experience." The purchase order reads: review architecture, reusable systems, evidence design, personal output, quality mechanisms.
System altitude, two dials instead of a ladder
Start with a badge, because it is the most portable object in this issue. The open-source assistant-ui library recommends an explicit "auto-approved" label for the case where a server policy, not a person, granted the permission. Somebody shipped a disclosure convention for consent no human ever gave, which puts both dials of the model below into production inside a single badge.
Short version of the argument here; the extended case runs in this issue's companion feature. Anthropic's autonomy research found that the most experienced Claude Code users hand over more and interrupt more.
| Behavior | Under 50 sessions | 750+ sessions |
|---|---|---|
| Full auto-approval | ~20% | Over 40% |
| Turns interrupted | ~5% | ~9% |
Notion ships the same principle as policy and calls it "progressive trust": human review on every action at the start, autonomy widening as reliability gets demonstrated, pace set by the user.
Two dials, moving independently. Autonomy scope is what the agent may do unasked. Intervention sensitivity is what trips review, interruption, or reversal. We named the dials. They supplied the evidence, and that evidence is overwhelmingly software-engineering behavior, so state the limit before an interviewer states it for you.
The libraries already sort themselves onto the two axes. Use that; it is the fastest proof the model describes shipped behavior and not a diagram. Vercel's AI Elements ships Confirmation with documented approval-requested, approval-responded, output-denied, and output-available states; assistant-ui exposes allow-once, allow-always, reject-once, reject-always. Intervention sensitivity, set per action. Server-policy approval, flagged as approval.isAutomatic, is autonomy scope. Checkpoint, a restorable conversation state you can branch from, is intervention sensitivity applied after the fact. elicitation, added July 30, is the agent handing scope back when it lacks information, with its own pending and validation-error states. Confidence: high, all of it from vendor docs and tagged releases.
This corrects something I published. Issue 4 gave you a trust ladder climbing in one direction, supervision toward independence. Wrong shape. Practiced operators loosen the leash and sharpen the tripwire in the same motion. The five handoff moments hold. The one-way ladder does not.
Provenance, once, and then say it yourself before anyone asks.
- Current practice, verbal only: TinyFish.
- Published Lab work at junochen.com: Carrier IQ, Brand Pulse, Retail Velocity.
- Prior enterprise and consumer work, predates agentic systems: Thermo Fisher, Alibaba, Red Cross, Allē, Equinox+.
Name the vintage in the same breath as the outcome. Otherwise an interviewer finds it later and reprices the whole answer.
Thermo Fisher is your load-bearing system case. Thirty-two interviews compressed into five operational failure modes and five modules, with QA release held as a human gate because regulatory authority was not yours to delegate. Six partners adopted it; the case reports $20M-plus in margin, case-reported and not independently verified. Carrier IQ shows the looser allocation, where the agent assembles reversible comparison evidence and the broker decides.
Routes to: Altana, HackerOne, Gusto, Maven, Abridge. Question: "How do you decide when the agent proceeds, asks, pauses, or requires a human decision?" Gap: failure-recovery sequence.
Your Carrier IQ wrapper claims eight portals and four minutes. The live homepage claims five carriers and 90 seconds. Say neither number until they agree.
Craft altitude, where the vocabulary now lives
The essay side of this conversation went dark. May 3 to August 1, exactly one named-practitioner article cleared the bar: Tony Alicea's "UX-Context Design" for NN/g on July 24, arguing that research and design knowledge should be curated into machine-readable context an AI consumes whenever it generates product work. His hypothetical context file specifies when to confirm versus allow undo, error rules, audit-trail requirements. He is candid that measuring efficacy and maintaining the thing at scale are both unsolved. Maggie Appleton's nearest relevant piece is April 13. Lenny's ran April 14, on agent architecture rather than supervision craft. Confidence: high on NN/g and Lenny, moderate-to-low on the others, indexing being what it is.
Route around the silence. Cite components, not essays. The freshest craft vocabulary this cycle is in release notes.
Iteration. Spring Health wants rigorous evaluation and continuous iteration. Amplitude and Ramp want production participation, code included. TinyFish supports a strong verbal account of traces, analytics, session replay, same-cycle review of live agent behavior. Keep it verbal, frame it as current practice, and understand what is not there: nothing published connects one observed failure to a specific interface revision and the release decision that followed. Question: "Show me one complete loop from design intent to production behavior and back to a changed design." Gap: trace-to-revision chain.
Component states. Tool, Queue, Task, Error, Branch Picker, elicitation. The libraries treat non-deterministic progress, recovery, and mid-flight requests for input as states somebody designed on purpose, not as edge cases. Carrier IQ is your closest match: one structured intake, parallel carrier work made visible rather than collapsed into a single spinner, then confidence, coverage, and exclusions shown side by side. Depth is missing: no annotated state machine, no rejected alternatives, no iteration history. Question: "Which states did you design that wouldn't exist in a deterministic workflow?" Gap: Staff-level component craft. Bites hardest at Stripe and Ambience.
Leadership altitude, one gap counted twice
Two demands that look unrelated on the page. Altana wants a semantic, machine-readable design language that PMs, engineers, and agents can compose from. Vanta wants definitions of done, launch reviews, quality frameworks. Systems work in one column, governance in the other.
Same test underneath. Both ask whether a design judgment stays correct once it leaves the person who made it. Alicea's context file encodes a judgment so a model can apply it. Vanta's definition of done encodes one so a colleague can. Notion ships the artifact version: permissions, activity logs, configuration version history, reversible runs, and a record of who approved what the agent did. Different consumers, identical property. Does the reasoning survive transfer with its boundary conditions and its revision trigger still attached?
Which means I over-counted. Issue 5, Fix the Surface Before You Build on It, listed AI-native process and agentic design systems as two gaps needing two artifacts. One gap, two output formats. Build the transfer record once: a criterion, where it was encoded, who applied it without you, where it should not apply, what evidence would revise it.
Alibaba is real system leadership. Trust treated as structural across homepage, search, and product detail, three cross-functional sprints, a reported 20% lift in daily transactions and a 2.2-point NPS gain, case-reported on the same caveat. It proves outcome. It says nothing about transfer.
Creative provenance is a true gap, and there is no partial credit available on it. Suno's posting never says provenance, authorship, or rights. It also puts you over an AI-native creative product and its creator experience, which is that problem whether the JD names it or not. The territory belongs to C2PA, the standard for cryptographically signed content origin, and the Copyright Office's standard for human expressive control. Prepare it as an interview probability rather than a stated requirement, and rank Suno accordingly.
Routes to: Altana, Vanta, Amplitude. Question: "How do you keep a design judgment correct after it leaves the person who made it?" Gap: transfer record.
Ranked routing
| # | Company | Altitude | Pattern route | Lead with | Missing |
|---|---|---|---|---|---|
| 1 | Altana | System → Leadership | Machine-readable system, confidence | Alibaba + Carrier IQ | Transfer record |
| 2 | HackerOne | System | Oversight, risk visualization | Carrier IQ + trust essay | Failure recovery |
| 3 | Vanta | Leadership | Definitions of done, launch reviews | Alibaba + Thermo Fisher | Mechanism at ~40 people |
| 4 | Amplitude | Craft → Leadership | Production code, player-coach | Retail Velocity + TinyFish | Trace-to-revision |
| 5 | Maven | System → Leadership | Where judgment stays human | Thermo Fisher + Red Cross | Reusable handoffs |
| 6 | Ramp | Craft | Agents, memory, production PRs | Carrier IQ | Design-to-production loop |
| 7 | Stripe | System → Craft | Agentic commerce, Staff craft | Allē + Carrier IQ | One component, many states |
| 8 | TRM Labs | Craft → Leadership | Real workflow, hands-on lead | Carrier IQ + Labs demo | Uncertainty demo |
| 9 | Abridge | System → Craft | Linked evidence, audit | Thermo Fisher | Source-to-summary fix |
| 10 | Spring Health | Craft | Reusable patterns, evaluation | TinyFish context | Release threshold |
| 11 | Gusto | System | AI-to-human escalation | Thermo Fisher | Agentic routing |
| 12 | Front | System | Queues, routing, trust | Thermo Fisher | Interruption states |
| 13 | Ambience | Craft | Clinical workflow craft | Thermo Fisher + Red Cross | Component states |
| 14 | Suno | Leadership | Authorship, provenance | Equinox+ consumer craft | Everything |
Rank tracks pattern-vocabulary density in the postings. Not company quality, not compensation.
Three questions, as they will be asked
"How do you decide when an agent should proceed, ask, pause, or require a human decision?" Thermo Fisher, then Carrier IQ as the contrast. Probe: what the reviewer actually saw, what made that gate structural rather than ceremonial, what evidence would justify widening autonomy later. Have all three ready.
"Show me one loop where production agent behavior changed your design." Carrier IQ for the shape, TinyFish for current vocabulary, and say plainly that the published chain stops short of the release decision. Probe: which trace contradicted your assumption, what criterion changed, what result would have stopped the ship.
"How do you keep agent-generated work coherent, and make that bar survive you?" Alibaba for the system result, then the honest sentence about what you are still building. Probe: where the criterion lived, who applied it without you, where it should not apply, what would revise it.
Gap ledger, collapsed
| Gap | Question it leaves open |
|---|---|
| Transfer record (previously two) | "Which judgment was applied successfully when you weren't in the room?" |
| Trace → revision → release | "What production behavior changed your design, and what made the change shippable?" |
| Carrier IQ failure and recovery | "What happened after the first wrong result?" |
| Staff-level component depth | "Take one AI component through its states and rejected alternatives." |
| Labs outcome evidence | "Who used this in real work, and what decision improved?" |
| Creative provenance | "What remains the creator's when AI contributed?" |
Build the transfer record first. It answers Altana, Vanta, and Amplitude in one pass.
Watch line, next 30 days
Two triggers would reopen the commentary channel. Conference-season recaps, which historically drag design leaders back into public writing. And a release cycle at Notion, Vercel, or assistant-ui, which usually pulls leader commentary along within the week. AI Elements last tagged a release on March 12 and is overdue; assistant-ui shipped July 30 and ships often. If the leader silence holds through August, stop waiting on it. Read temperature off library release notes and let posting repetition carry the mandate.
-
Figma's eval process, in detail: If Spring Health or Amplitude asks what "done" means for a probabilistic product, borrow the structure from First Round's account of Figma Make's evaluation practice — one-to-four design and functionality scoring, roughly 1,000 corpus examples, golden prompts, and nightly model comparisons.
-
When the benchmark is the bug: OpenAI's July audit found that about 30% of SWE-bench Pro tasks were broken, which is the sharpest available answer to any interviewer who treats a benchmark number as evidence of product quality.
-
The expectations gap nobody has priced: Designer Fund's survey of more than 900 designers reports weekly AI use rising from 54% to 91% while only 28% of companies changed how they evaluate, pay, or hire — useful ammunition in a leveling or compensation conversation.
-
Permission tiers as interview language: Anthropic organizes agent permissions into always allow, require approval, and block in its trustworthy agents research, which gives you a citable vocabulary for the autonomy-scope dial rather than describing it in your own terms.
-
Altana's equity window is unresolved: The most recent official financing record is a $100 million Series B from October 2022, so ask about round history and current valuation before you weigh the offer against Vanta's disclosed cash range.

