Three moves, in this order. First, strip the decorative access gate off your case pages today, because it costs an afternoon and it currently contradicts the artifact you're about to publish. Second, build one artifact instead of five. Third, publish the three finished essays on the schedule I set at the bottom.
The finding is small and a little awkward. The product you'd write your best artifact from has an approve button and no reject button. You can't write that artifact from memory. The event hasn't happened yet.
The gate, and a retraction
As of this morning: all five full case pages ship the complete case HTML to the browser before any access check runs. The private essay does the same. The server hands over the whole document, TinyFish pricing material included, and then an overlay asks for a password. Curtain, not door. I made this argument in Trust Calibration Starts at Home. What's changed since is my read on whether it survives contact with a hiring panel. In the Vanta positioning sheet I told you the tiered gating read as a deliberate trust ramp and could be presented that way. Retract it. You are about to publish work whose entire claim is that you design real authority boundaries, and anyone who runs view-source finds a decorative one in ten seconds.
So: delete the overlay from the five case pages, and take the private essay off the public web altogether. Not a server-side gate. Both are defensible in principle, but a real gate is engineering work on a calendar you don't control, and it wants the same weekend as the correction build, which outranks it. Deleting a fake boundary is an afternoon. The essay is the one page where "public, no gate" is the wrong answer, because the TinyFish pricing material inside it isn't yours to publish. That one comes down rather than opens up.
What a control record is, and why one beats a checklist
A control record follows a single causal chain across time. One wrong machine output. The surface that made it visible. The correction a human made. What the system did differently afterward. The verification step that showed up in the next run. The release criterion that changed because of it. Six links, one event, more than one run.
Why one of those outranks a portfolio of separate artifacts is a mechanical question, not a taste question. Happy-path case studies can be staged. You pick the screens, you pick the flow, and a missing control stays invisible because nothing in the story ever reaches for it. A contested run can't be staged. Put a wrong output on screen and every absent control becomes conspicuous in the same frame. There is no version of this artifact that survives a hollow product underneath it, which is exactly why it reads as proof instead of assertion.
There's a market reason too, and it's the one I'd lead with. Line up the Act-tier postings on your board and they're all describing one surface in four vocabularies: the moment a person decides whether to trust what the machine produced, and what they can do if they don't. That moment is the sorting variable in senior AI design hiring right now. Visual craft, system rigor, breadth of surface: table stakes at your level, because every candidate in the slate has them. Owning the trust-and-override moment isn't, and almost nobody in the slate can prove they've designed one.
Which brings us to Carrier IQ. There's more inspectable machinery in it than there was in July: canonical intake, parallel carrier runs staged as Session, Navigate, Fill, Extract, Verify, four post-extraction review states, coverage deltas, per-carrier evidence sessions that hold a session ID, a live URL, a proof attachment, a case note, and an Approve Bind decision that writes down provider, timestamp, operator note, and the evidence session it was approved against.
There is no control for "this number is wrong."
So artifact one is a build, then a run, then a write-up, in that sequence. Add the correction event. Run the contested case. Record the pair. Publish. Everything below ranks on distance from that spine. Does it visualize a moment inside the record, document the process that produced it, name the authority state it passes through, define the shipping criteria it changed, or defend it under pressure in a live room? Extensions beat standalones, in every domain. Same aperture discipline as The Same True Story, Told Four Ways: one event, several framings, one underlying truth.
Calibration before the queue
Quotes are from postings only. Anything in quotation marks below is language lifted from a primary employer record. Where I've written an interview question, that's my synthesis from the other four source types: company design blogs, practitioner publications, hiring-manager commentary, and the emerging agentic pattern libraries. I haven't dressed synthesis up as a quote. For which companies have actually shipped intent architecture, provenance, human gates, and correction versus which only publish about them, go to Said vs. Shipped. That piece is your vocabulary source. This one is your build order.
Tier labels are a week stale, and I'm not guessing at the gaps. Act means outreach this week. Watch means monitor without spending. Last complete table is the Design Context File, July 24, amended the 27th and 29th.
| Company | July label | Where it stands |
|---|---|---|
| Vanta, Amplitude, Ramp, Maven VP | Act | Outreach this week |
| HackerOne | Verify-to-Act | Live direct posting has since surfaced; treat as Act on verification |
| Gusto | Probe | No final label |
| Ambience, Rubrik | Watch | Monitor without spending |
| Abridge, Stripe | Language-harvest | Not promoted |
| Altana | None | Posting postdates the table outright |
| Front, TRM Labs, Spring Health, Maven Senior Staff | None | Never got an assignment in the July record |
Unassigned is not unimportant. Build for them; verify timing before you reach out.
Every mapping below names the rubric dimension the artifact actually moves. Three are in play.
- AI centrality — whether AI is the product or a feature bolted onto one.
- Design influence ceiling — how far a design mandate travels before it hits a wall.
- AI exposure quality — how close your own hands have been to shipping agent behavior in production, as against reading about it.
Your board is strong on the first, thin on the second, and the control record is the only artifact that decisively settles the third. Tags are shorthand: [centrality], [ceiling], [exposure].
Tier 1 — this week
1. Carrier IQ longitudinal control record (Trust Design; the spine itself) [exposure]
- Asked as: "Tell me about a time a system you designed produced a wrong output. What happened next?"
- Unlocks: Amplitude (Act), which asks the leader to "help define how agents and humans work together in the same workflows" and will press hard on which AI-native surface you personally built. HackerOne (Act on verification), which reserves "trust, oversight, and strategic decision-making" for humans while agents run operations. Gusto (probe), which frames the leader's job as deciding "when a human is needed, and when outputs should not be sent out at all." Altana (unassigned), which lists correction, escalation, interruption, and graceful failure as evaluated surfaces.
- Build: new build first, then writing. The correction control doesn't exist yet. Nothing else in this queue is blocked by it. Everything else is weaker without it.
- Status: pending. App is live. Record isn't.
2. Two-control autonomy and intervention map (Trust Design; names the authority state the record occupies) [centrality]
- Asked as: "How do you decide how much autonomy to give an agent?"
- Unlocks: HackerOne on inspectability plus human intervention. Gusto on send, suppress, escalate. Maven VP (Act), which wants clarity on "what AI enables and where human judgment matters most." Maven Senior Staff (unassigned), Abridge, Ambience, and Spring Health each name intervention triggers in their own vocabulary, which tells you they've stopped asking where the handoff sits and started asking how it's tuned.
- Build: writing plus one diagram. No product work.
- Status: pending. It ranks second because it corrects a ladder you've already published, and a public revision is stronger evidence of judgment than any first draft. Draft below.
3. A judgment you can hand over (Presentation and Credibility; the criterion the record changed, generalized) [ceiling]
- Asked as: "How do you set a quality standard across a team without becoming the bottleneck?"
- Unlocks: Vanta (Act), whose success criteria name "design system, definitions of done, and launch review frameworks." Amplitude, on holding a bar without turning into an approval gate. Maven VP, on standards that make good work repeatable across product, brand, and research. Altana, which asks outright for a "structured, semantic, machine-readable language" of patterns usable by non-designers and by tooling. Gusto, where an 80-plus person design org means personal taste doesn't scale by definition.
- Build: writing only. Cheapest item in the tier, biggest gap closed.
- Status: pending. Draft below.
Blunt version. Your craft altitude is not in question anywhere on this board. Your leadership altitude is in question at both Act companies with large design orgs. Item 3 costs you one evening and it's the item most likely to change a hiring conversation.
If the correction control isn't running against a real contested case by August 15, stop waiting on it. Publish artifacts 3 and 2 standalone, ship the record as a follow-up when it lands, and lead outreach with the revision instead of the proof.
A window that closes while you build is worse than a queue published out of order.
Tier 2 — everything here extends the record
4. Failure-and-recovery sequence (Trust Design; visualizes a moment inside the record) [exposure]. "Walk me through your recovery design." Same event, tighter aperture: error state, correction affordance, recovery path, consequence. Gusto and Altana both name graceful failure as an evaluated surface, and Gusto specifically wants to know whether the operator gets a usable path out or a dead end. Needs screenshots. Pending.
5. Eval-forward release criteria (Design Patterns; the sixth link, standalone) [ceiling]. "What would have to be true for you to ship this?" Six columns: failure mode, trace evidence (the log of what the agent actually did), design decision, eval criterion (the automated test that gates a release), release gate, production monitor. Spring Health (unassigned) pairs conversation design with "rigorous evaluation, and continuous iteration" and names versioning, staged rollouts, regression testing; Vanta wants launch review frameworks. Writing plus screenshots. Pending, and already sitting in your own gap register as the missing bridge between design intent and shipping discipline.
6. Seven-state agent state machine (Design Patterns; names every authority state the record passes through) [centrality]. "What states does your agent have, and who can move it between them?" HackerOne on human-in-the-loop oversight and agent-to-agent interaction; Gusto on the send, suppress, escalate, hand-to-expert decision. One page answers both, and no other artifact does. Needs new build (diagram). Pending.
7. Carrier IQ live-demo script (Presentation and Credibility; defends the record under pressure) [exposure]. Nobody asks this question. It gets asked at you, live, at the moment the run breaks on the call. All four Act companies will put you in front of a screen, and the criterion under evaluation is your composure when the demo dies, not the demo. One page: the three questions you want them to ask, and the two you can't answer yet. Writing only. Pending.
8. Annotated trace-to-design record (Design Iterations; documents the evidence layer beneath the record) [exposure]. "How do you use production data in your design process?" Amplitude will ask what you built and what you learned, and expects the learning to trace back to observed behavior. One real trace, annotated at the points where it changed a decision. Needs screenshots plus writing. Pending.
9. Visual Trust Pattern Library (Trust Design; generalizes the record's components) [ceiling]. "Show me your patterns, not your screens." Ambience (Watch) on scalable systems that don't cost you quality; Abridge on grounding output so a provider can verify it. Prose for five agentic component patterns already exists in the last queue. Nothing visual is published. Needs new build. Partially drafted.
10. Machine-readable design-system proof (AI-Native Execution; encodes the criterion) [ceiling]. "Can something that isn't you use your system correctly?" Four panels: unconstrained result, encoded constraint, constrained result, prevented drift class. This is Altana's literal portfolio requirement, a system actively used by non-designers and ideally by AI tooling to ship production experiences. Needs new build. Pending.
11. Carrier IQ evidence-run demonstration (AI-Native Execution; prerequisite, not sibling) [exposure]. Evidence sessions, live URLs, proof attachments, case notes, approval notes: roughly implemented already. It serves Ramp and Amplitude on the hands-on production expectation, which is precisely why it shouldn't stand as its own page. Fold it into artifact one rather than build it twice. Partially implemented.
Tier 3 — hold, but know what each one buys
Two of these are finished. Ship them on the record's schedule anyway. A process essay published before the record is an assertion; published in the same week, it's an index to evidence.
| Artifact | Domain | Dim. | Asked as | Unlocks / criterion | Build | Status |
|---|---|---|---|---|---|---|
| "How I Design AI Products" essay | Design Iterations | exposure | "How do you work?" | Amplitude, hands-on production participation | Writing only | Complete draft published as draft material in the July 24 queue; not on junochen.com. Ship with #1 |
| Agentic Labs decision-moment annotation | Design Patterns | exposure | "Which decisions were yours?" | Ramp (Act), hands-on production expectation | Writing only | Same source and status as above. Ship with #1 |
| Intent → code → run → trace → revision loop | Design Iterations | exposure | "Show me one full cycle." | Amplitude, systems-level thinking and pixel-level craft in a single artifact | Needs screenshots | Pending |
| Constraint-first design-system before/after | Design Iterations | ceiling | "How do constraints show up in your system?" | Altana, patterns and interaction grammars | Needs new build | Pending |
| Figma-to-production-code craft record | AI-Native Execution | exposure | "Do you ship code?" | Maven Senior Staff (unassigned), production code fluency | Needs new build | Pending |
| Agentic component deep dive | Design Patterns | centrality | "Show me the multi-state component." | Ambience (Watch), "from whiteboard to launch" ownership | Needs new build | Pending |
| Staff-level component craft deep dive | Presentation and Credibility | centrality | "Take me to the pixel level." | Stripe (harvest), Abridge (harvest), Ambience on high-stakes craft | Needs new build; same session as above | Pending |
| Multi-sided AI service-design artifact | Design Patterns | ceiling | "Who else is in the loop?" | Gusto, two-sided operational handoffs | Needs new build | Pending. Allē and Red Cross carry the non-AI version |
| High-stakes regulated-AI trust artifact | Trust Design | centrality | "What changes when it's regulated?" | Abridge, output that "maps AI-generated summaries to ground truth" | Needs new build | Pending. Thermo Fisher carries the non-AI version |
| Career-narrative visual | Presentation and Credibility | ceiling | "Walk me through your path." | All Act companies, first five minutes of the review | Needs new build | Pending |
| First-person failure narrative | Presentation and Credibility | ceiling | "What's a decision you got wrong?" | Ambience, critique and postmortem culture | Writing only | Pending. Distinct from #4: that one is the product's failure, this one is yours |
Two standing rules. TinyFish is currency, not collateral. It buys you credibility as someone who reads agent traces today, and it never carries a design claim, a metric, or a case study. The brackets in Draft 1 are not prose gaps to smooth over. Each one is a fact only the contested run can produce. Don't publish with a bracket unfilled, and don't publish with a bracket filled by an estimate.
Publish order for the three drafts. Draft 3 goes up this week, independent of whatever state Carrier IQ is in, because it's a criterion with exclusions and a kill condition rather than a process story, and it closes the leadership-altitude gap that's currently throttling your read at Vanta and Gusto. Draft 2 follows inside the same week, as a public revision of a claim you've already published, which stands on its own for the same reason. Draft 1 publishes the day after the contested run produces the before-and-after pair, not one day sooner. That's the bracket rule stated positively: a first draft with an open bracket is finished writing waiting on evidence, not a publishable page.
Three drafts follow, in your voice.
Draft 1 — The Reject Button I Didn't Build
I built a quoting surface for insurance brokers that drives carrier portals in parallel, five by default and eight available. The broker enters the client once: identity, vehicle, prior carrier, continuous coverage history. Then the agents fan out. Every carrier run moves through the same stages, and I put those stage names in the interface on purpose: Session, Navigate, Fill, Extract, Verify. What comes back lands in a comparison view carrying coverage deltas, so the broker sees where a carrier diverges from the baseline instead of reading eight PDFs against each other.
Then the surface routes every result into one of four states. Bindable. Normalize. Referral. Call review.
That routing is the design. It's the sentence the product says to the broker: here's how much of this you can trust, and here's what you're allowed to do about it.
Open any carrier record and you can launch an evidence session against that carrier's live quote flow. The session holds its ID, its status, its execution mode, the URL it worked against, a summary of the last step, the parsed outcome, a proof attachment, a case note, and a follow-up prompt if the broker wants to send the agent back for more. A bindable quote can then take an Approve Bind decision, which writes down the provider, the timestamp, the broker's note, and the evidence session it was approved against.
I was proud of that record. It took me too long to notice what wasn't in it.
There is no control that says this number is wrong.
The broker can edit the intake and rerun. Leave a case note. Send the agent back to look again. Approve. What the broker cannot do is take a single extracted value, mark it incorrect, and have that assertion mean anything at all to the system. I built the approval and skipped the refusal.
That's a claim about who is in charge, and I made the wrong one. An approval-only interface tells the operator their judgment is a rubber stamp. It collects consent and quietly removes the thing that makes consent real, which is the ability to withhold it at some cost to the system. Say a deductible comes back at [BRACKET: wrong value as extracted] while the carrier's own page says [BRACKET: correct value, with source location]. Under the design I shipped, the broker can work around it or sign it. That's the menu.
Here's what I'm building, and the record I'm keeping while I build it.
A correction is an event, not an edit. When the broker marks a value wrong, the original stays. The corrected value gets appended with the broker's identity, the timestamp, and the evidence session it was corrected against. Overwriting destroys the only thing that makes the correction useful later, which is the pair.
A correction changes the route. A record with an operator-corrected field cannot be Bindable on that run. It drops to Normalize at minimum. The whole point of the routing vocabulary is to say how much verification a result has survived, and a corrected field has survived less than an uncontested one.
A correction changes the next run. Verify already exists as a stage, and today it runs generically. After a correction on a specific field for a specific carrier, Verify picks up a targeted check on that field, and the check appears in the run log, because that's how the broker learns the system heard them. [BRACKET: run N and run N+1 side by side, targeted check visible in N+1.]
A correction changes what shipping means. Ready used to mean the four states resolved without an error. Now a build can't ship if a field corrected in a prior run comes back uncorrected and unflagged in the next one. Three weeks ago that criterion did not exist in my head. [BRACKET: criterion as written into the release checklist, dated.]
Two limits I won't paper over. Confidence in this product is derived from intake completeness, not from prior-run agreement, and I haven't calibrated it against anything. And the broker's authority here is a product authority, not a legal one. Approve Bind is a designed decision moment. I'm not claiming a regulatory workflow I haven't built.
I'm publishing the correction event rather than the comparison view because the comparison view is easy to make look good. Any of us can stage a happy path. Nobody can stage a run that disagreed with itself, or the small, boring, expensive machinery that has to exist before the disagreement can go anywhere.
Draft 2 — Two Dials, Not Three Rungs
I've been drawing the same ladder for two years. Watch, then verify, then delegate. The agent earns autonomy, oversight recedes, everybody's happier. Three rungs, one direction, fits on a slide.
It's wrong. I found out by building on it.
I built the ladder into the carrier quoting surface, where four review states carry different amounts of trust. Then into Brand Pulse. Then into Retail Velocity. The autonomy setting travelled between all three without much friction. The intervention design travelled not at all. Every time, I rebuilt it from scratch, differently, worse. That's the signature of a model with a missing variable: the thing you never named is the thing you can't reuse. I see the same shape every day in production traces I can't show you, which makes it corroboration and not evidence, so treat it that way.
Here's the variable I was missing. There are two controls, and they move independently.
Authority scope is what the agent may do without asking. Everybody designs this one. It's the approve step, the auto-run toggle, the don't-ask-me-again checkbox.
Intervention trigger is what evidence or consequence pulls a human back in, and what that human can do once they're there. This dial has three settings that get collapsed into the single word "oversight," and they shouldn't be. Review is looking at something after it happened. Interrupt is stopping something mid-flight. Reverse is undoing something that already landed. A system can offer any combination. Most offer review and stop.
The failure mode is moving both dials the same direction. Autonomy up, intervention down, which is what the ladder instructs. The result is a recognizable species of product: cautious and irritating while nobody trusts it, then quietly dangerous the moment somebody does. Interruption affordances get designed for the novice who doesn't yet know when to interrupt, and they atrophy exactly when the expert starts needing them.
Treat the dials as separate and a few things fall out.
Scope autonomy by reversibility, not accuracy. A quote extraction can run unsupervised, because a wrong number gets caught downstream. A bind cannot, because there is no downstream left. I don't need to know how often the agent is right to make that call. I need to know what happens when it isn't.
Tier intervention triggers by consequence, not confidence. Confidence is a property of the model. Consequence is a property of the world. When I routed carrier results into four named states, what set the routing wasn't how sure the system was. It was what a wrong answer in that state would cost, and who would pay. Bindable means an error is expensive and unrecoverable, so a human decision is required. Call review means the evidence is thin in a way that no additional automated run will fix. Both are consequence statements wearing confidence clothing.
Make interrupting socially free. If stopping a run means the operator loses their place, re-enters their inputs, or looks to a colleague like someone who doesn't trust the tool, they won't do it, and the oversight is theater. The cost of interrupting is a design variable. Almost nobody treats it as one.
Give reversal a window, and show it. Not an "are you sure" dialog. A stated period during which the operator can pull the action back, displayed before they commit, with enough state preserved that pulling it back is cheap.
I wrote in Trust Is the New Interface that trust is what we're actually designing now. I stand by the claim and I'd revise the mechanism. Trust doesn't slide along one axis from suspicion to surrender. It's two settings moving independently, and the systems I've watched work best hold both high at once: the agent does a great deal, and the human can stop it fast, on a thin signal, at no cost.
The ladder let me skip the second dial because the third rung implied it disappears. It doesn't disappear. It gets sharper.
Draft 3 — A Judgment You Can Hand Over
Most of my best design decisions died with the project. They weren't wrong. They lived in my head as taste, and taste doesn't transfer. Somebody asked me why a state was called Normalize instead of Review. I gave a good answer in the room. Six weeks later a different surface shipped with a different word carrying a different meaning, and nothing in the system objected.
A design judgment is only real if it survives the person who made it. Here's one of mine, written so it doesn't need me.
The rule. Any machine-produced result a human will act on gets routed into exactly one named state before it reaches them. Each state has an entry condition, one permitted action, and one escalation path. Three states minimum, five maximum. In the carrier quoting surface I built, they're Bindable, Normalize, Referral, and Call review.
The entry conditions. A result is Bindable when every field the decision requires was extracted and none was corrected on this run. Normalize when a field diverges from the baseline in a way a defined transform can reconcile; deductible drift is the ordinary case. Referral when resolving the divergence needs a party outside the current session. Call review when the evidence available to software is insufficient in kind rather than in quantity, which is the precise condition under which no further automated run will help.
The permitted action. One advancing action per state, and it's the only one the state promotes. Bindable can be approved. Normalize can be transformed and re-checked. Referral can be routed. Call review can be assigned to a person. Inspection stays available everywhere: pull the evidence, read the deltas, leave a note, because looking is never the action that needs constraining. I know this matters because I built the permissive version first, all four advancing actions live on every record, and then I watched myself sitting at my own screen with a drifted deductible, reaching for Approve because it was the leftmost button and the number looked close enough. If the man who wrote the states can pick the wrong action inside his own interface, four actions per state isn't flexibility; it's a design that can't explain itself.
The escalation path. Every state escalates in exactly one direction, toward more human involvement, never less. Nothing auto-promotes into Bindable. Promotion takes a human action with a recorded identity.
That's the criterion. A product manager can apply it. An engineer can implement it without asking me what I meant. A model can be handed it as a constraint and produce a routing I'd recognize, which is a test I've started running and a low bar I've watched systems fail.
The two parts people skip are the parts that make this a criterion instead of a preference.
Where it must not apply. Any workflow where the wrong route is unrecoverable inside the same session. If routing to Normalize instead of Call review means a patient receives a summary nobody verified, then the rule is too coarse and the state count is too low. It also shouldn't apply where the operator has no authority to act on the escalation, because then the states are labels and you've built a taxonomy while telling yourself you built a control. I learned that boundary expensively on a regulated supply-chain program, where a status nobody downstream was empowered to act on turned out to be worse than no status at all. That work is in the mySupply case.
What would invalidate it. If operators systematically override one route, the route is wrong, not the operators. A standard has to name the evidence that would kill it or it isn't falsifiable, and an unfalsifiable design standard is the loudest person's preference with a document wrapped around it. Concretely: if more than a small minority of Bindable results get manually downgraded before approval, the entry condition is too permissive, and I change the condition rather than the training.
The value here isn't the four states. Somebody else's product needs different ones. The value is that the judgment left my head in a form with entry conditions, exclusions, and a kill condition, which means a team can inherit it, argue with it, and prove it wrong without me in the room.
As far as I can tell that's the whole job at this level. Not making the good decision. Making it portable.
-
Carrier count contradiction on live pages: The homepage says five carriers in 90 seconds while the Carrier IQ wrapper says eight portals in four minutes, and the live app resolves the count but not the timing — fix one before Draft 1 ships, because a panel that reads both pages will read the gap as carelessness.
-
The confidence basis is unresolved: The wrapper derives confidence from data completeness while the standing portfolio record described prior-run consistency, so decide which is true before any interviewer asks how the score is computed and whether it has ever been calibrated against the Anthropic autonomy findings on how experienced operators actually behave.
-
Approval-plus-interruption is now the observed pattern, not the exception: Anthropic's coding-agent data shows auto-approval rising with experience while interruption rises alongside it, and Notion's Custom Agents ship the same two-axis model as permissions plus run logs and automatic pausing — cite one external implementation when you publish Draft 2 so the revision reads as market observation rather than personal correction.
-
Suno reopened and nobody has recorded the tier: The direct application page is live and the CPO promoted the search publicly about a month before the last scan, per Jack Brody's post, so the Watch label is stale and the creative-provenance gap remains the one true evidence hole on that board.

