Where the signal came from
No qualifying hiring-manager commentary from any target company in the past thirty days. An OpenAI design-systems hire was announced without evaluative criteria. Older statements from Gusto, Amplitude, and Stripe hold as validated pattern but not current temperature.
Practitioner publications — Designer Fund, Lenny's Newsletter, NNG, Maggie Appleton, Amelia Wattenberger — yielded nothing new on agentic design patterns this cycle. Same for the AI company design blog set (Anthropic, Figma, Notion, Vercel, Linear). The posts cited in Cycle 1 remain standing evidence. Nothing additional appeared.
That leaves postings and component libraries as the live channels. Five fresh postings — OpenAI Payments, OpenAI Engineering Acceleration, Giga, Fieldguide, and Vanta — create the cross-reference. All five name shared systems, patterns, and quality standards. Three of five require AI fluency inside the designer's own practice.
On the library side, Agentic Office UI moved provenance below citations into resolvable object references — document revision, semantic locator, reliability metadata — but shipped none of the approval, diff review, or rollback surfaces that would complete the picture. No mature component exists for agent budgets or authorship attribution.
Seven patterns follow. Each is framed as a claim the market is forming. The test for each: if an interviewer presses this claim, can you answer with a specific production decision and its outcome?
Trust
1. Confidence and authority are independent controls
The market is splitting what used to be one slider into two separate design problems. How much the user trusts the agent's output is a different question from how much the agent is allowed to do unsupervised. OpenAI Engineering Acceleration names permissions, provenance, and rollback as distinct surfaces from output quality. Gusto separates ready-to-send review from output suppression. Last cycle's analysis retired the one-way Watch-Verify-Delegate ladder and named these as autonomy scope and intervention sensitivity. This cycle, the split is showing up in posting vocabulary. High confidence — multi-source convergence across postings, library contracts, and prior cycle analysis.
Lead with Carrier IQ's confidence rule: confidence rendered only from observed input, never from the model's self-report. That addresses the first control with a production decision. The Trust essay's five handoffs address the second.
If your published Trust essay still presents a single progression visually, that's a vulnerability an interviewer familiar with the two-control model could press on. Verify the visual framing separates the two controls.
Routed to: OpenAI Engineering Acceleration, Gusto, Vanta. They'll ask: "How do you decide what an agent is allowed to do versus how you communicate what it actually did?"
2. Approval defines the agent's capability boundary
CopilotKit's interrupt and checkpoint contracts pause an agent at a graph-enforced boundary and resume only after a custom UI returns a decision. The agent cannot proceed past this point without authorization. The design decision is where to place the gate. OpenAI Engineering Acceleration names permissions as a design surface. Giga describes self-improving agents, which makes the boundary question acute — a self-improving agent that can widen its own capability boundary without a gate is a fundamentally different product than one that cannot. Moderate confidence — posting vocabulary and library evidence converge, but only one posting names permissions explicitly.
Your answer here is Carrier IQ's control record: wrong output, visibility, correction, changed system behavior, next-run verification, revised release criterion. Six links where the gate is the first link and the system holds until it clears. That's production evidence of designing the boundary.
Routed to: OpenAI Engineering Acceleration, Giga, Fieldguide. They'll ask: "Show me a time you decided where to put the approval gate and what happened when you moved it."
3. Provenance is moving below citations
Agentic Office UI's July releases introduced document surfaces where a human selection carries a revision, semantic locator, content evidence, and reliability metadata — so the reference survives after the agent changes the file. This goes beyond "we cited our sources." The system can point to a specific object, at a specific revision, with a reliability score, and resolve that reference again later. No library has shipped the full surface connecting evidence, authorization, history, and undo. The components exist in fragments. Moderate confidence — library evidence exists and one posting names the vocabulary, but the pattern hasn't converged across multiple postings.
Carrier IQ's trace evidence is partial proof. You can show how the system surfaces what it observed and why it reached a conclusion. You cannot yet show resolvable object-level provenance with revision tracking. Claim gap. If an interviewer asks about provenance at the object level, name the gap and describe what you'd build, grounded in what Carrier IQ already does. Do not improvise a framework.
Routed to: OpenAI Engineering Acceleration, Vanta, Fieldguide. They'll ask: "How does the user verify what the agent changed and undo a specific action without undoing everything?"
Patterns
4. Quality systems are the design-leadership test
The strongest signal this cycle. Five of five fresh postings name shared systems, patterns, quality standards, or definitions of done. Vanta names design systems, definitions of done, launch review, and quality frameworks. Fieldguide asks for design-system contribution and a craft standard others follow. Giga asks the designer to evolve the design system and uphold quality standards. Both OpenAI roles name design systems and shared patterns. Five-of-five convergence is as strong as posting signal gets.
When every posting in a cycle converges on the same vocabulary, expect that vocabulary in the interview. The question will be how you maintain quality when the output is non-deterministic and the system changes its own behavior between releases.
Lead with Carrier IQ's release criterion — the specific, revisable standard that determines whether a new behavior ships. That's a quality system operating on non-deterministic output. Agentic Labs design system work is relevant but secondary. TinyFish strengthens your position here: you're maintaining quality standards on AI product output in production right now, which means you can speak to this as a current operational reality. That carries differently in the room than recounting a past project.
Routed to: All seven companies. Highest urgency at Vanta and Fieldguide, where the posting language is most explicit. They'll ask: "What's your quality bar when the product's output varies every time?"
5. Encoded design judgment
Amplitude built a Design Agent after 300+ AI-generated applications exposed visual and system drift. The agent embeds Amplitude's design context so generated output stays closer to the product language. The pattern: design decisions encoded into the system's behavior rather than applied by a human reviewer after the fact. The Confidence Rule is the same pattern at a different altitude — one encoded judgment, its known boundaries, and the empirical condition that would retire it.
Your strongest pattern. Carrier IQ's confidence rule is a named, bounded, falsifiable design judgment encoded into production. The prior cycle established the standard: one rule, its limits, its retirement condition. I have high confidence that no other candidate in these searches presents encoded judgment at that level of specificity.
Routed to: Giga (self-improving agents need encoded judgment), Amplitude, Gusto. They'll ask: "How do you encode a design standard into the product so it holds without manual review?"
Iteration
6. Evaluation loops are design scope
OpenAI Engineering Acceleration names the full sequence: instrument, launch, observe, investigate, evaluate, decide, iterate. The only posting that spells out the loop this explicitly, but significant — it frames the designer's scope as including the feedback cycle after launch. Moderate confidence — one posting names the full loop, and the concept underlies several others without the explicit vocabulary.
Lead with Carrier IQ's control record, one full cycle. Wrong output, what made it visible, how you corrected it, how the system's behavior changed, how you verified the fix, how the release criterion was revised. Walk through one specific instance where the loop changed a design decision in the next release. TinyFish gives you current-role context here — you're running evaluation loops on AI product output now, so you can describe the operational texture of this work as something you did last week.
Routed to: OpenAI Engineering Acceleration (primary), Giga, Amplitude. They'll ask: "Walk me through how a shipped agent feature's performance changed what you designed next."
Execution
7. Agent spending is a design surface
OpenAI Payments asks candidates to design how customers understand, track, manage, and pay for new forms of AI usage. Stripe Link says agents will transact on users' behalf. No component library has shipped budget controls, spending limits, or cost-attribution surfaces for agent actions. The pattern is forming at exactly two companies. Speculative — thin sourcing, no library convergence, no production evidence.
Claim gap. No Agentic Labs project addresses cost attribution or spending controls for agent actions. If OpenAI Payments is a live target, this gap needs portfolio work or a credible design walkthrough grounded in adjacent production decisions — Carrier IQ's control record applied to a cost boundary. Do not pad with frameworks.
Routed to: OpenAI Payments (primary), Stripe Link. They'll ask: "How would you design the experience of an agent that spends money on the user's behalf?"
Pattern-to-company mapping
| Pattern | Trust | Iteration | Patterns | Execution |
|---|---|---|---|---|
| 1. Independent trust controls | OpenAI EA, Gusto, Vanta | |||
| 2. Approval as capability boundary | OpenAI EA, Giga, Fieldguide | |||
| 3. Provenance below citations | OpenAI EA, Vanta, Fieldguide | |||
| 4. Quality systems | All seven; Vanta, Fieldguide highest | |||
| 5. Encoded design judgment | Giga, Amplitude, Gusto | |||
| 6. Evaluation loops | OpenAI EA, Giga, Amplitude | |||
| 7. Agent spending | OpenAI Payments, Stripe Link |
Top three anticipated interview questions this cycle
1. "What's your quality bar when the product's output varies every time?" (Pattern 4. Five-posting convergence makes this the most likely question across all targets.) Your answer: Carrier IQ's release criterion. A specific, revisable standard applied to non-deterministic output. Walk through one instance where the criterion caught a regression and what changed in the next release. Your current work at TinyFish means you can ground this in what you're solving now.
2. "How do you decide what an agent is allowed to do versus how you communicate what it actually did?" (Pattern 1. Tests whether you've separated the two controls.) Your answer: Carrier IQ confidence rule handles the communication side. The Trust essay's five handoffs handle the authority side. Name both, then describe a production moment where they pulled in opposite directions.
3. "Walk me through how a shipped agent feature's performance changed what you designed next." (Pattern 6. OpenAI Engineering Acceleration's explicit loop.) Your answer: Carrier IQ control record, one full cycle. Wrong output, what made it visible, how you corrected it, how the system's behavior changed, how you verified the fix, how the release criterion was revised.
Claim gaps
Provenance at the object level (Pattern 3). Carrier IQ has trace evidence but not resolvable object-level provenance with revision tracking. Priority if Vanta or Fieldguide advance to portfolio review.
Agent spending as a design surface (Pattern 7). No production evidence of designing cost attribution, budget controls, or spending authorization for agent actions. Priority if OpenAI Payments is a live target.
Non-deterministic system design (isolated vocabulary). Giga is the only posting that names it explicitly, but the concept underlies Patterns 4, 5, and 6. If Giga advances, prepare a specific walkthrough of designing for output variance — a production decision where variance broke something and you changed the design.
- Gusto's builder mandate: Katie Kovalcin's designers-to-builders post describes designers creating agent skills for research and component auditing, which may preview the evaluation criteria behind Gusto's live Head of Design posting.
- Amplitude's Design Agent: Amplitude published a detailed account of building an internal Design Agent after 300+ AI-generated applications exposed visual drift, making it the closest public precedent for encoded design judgment at a target company.
- CopilotKit's interrupt contract: The human-in-the-loop documentation specifies graph-enforced checkpoints and custom approval UIs that map directly to the capability-boundary pattern forming across OpenAI EA, Giga, and Fieldguide postings.
- Agentic Office UI's reference model: The July
0.5.xreleases introduced resolvable object references carrying revision, semantic locator, and reliability metadata, the first library-level implementation of provenance below citations.

