Six companies clear Act threshold (company rubric 10+ AND role rubric 12+). Four named: OpenAI Payments, OpenAI Engineering Acceleration, Gusto, Giga ML. Two surfaced during research. Fieldguide clears cleanly: posted July 10, inside the 42-day window, and while comp at $180K–$210K drags the role score, scope and regulated-AI ownership compensate. Stripe Link clears on rubric (role 13+, comp $210K–$316K), but I cannot confirm the original posting date. Verify the window before you spend time on outreach.
Ordered by action urgency. OpenAI Payments leads because your Alibaba evidence answers the problem they wrote into the posting almost line for line, and both OpenAI roles posted August 6, the freshest windows on the list. Gusto and Giga follow on how tightly your evidence sits against the problem each named. Fieldguide trails on comp, Stripe on the unresolved window.
OpenAI Payments
What the posting tells you the room is solving: "Support and monetize emerging forms of product usage" across "pay-as-you-go models, Codex usage, and agentic workflows." The research mandate covers how users "evaluate pricing, track usage, manage billing, and make purchasing decisions." They name "complex platform dependencies, financial logic, and technical constraints."
Claim: Juno designed transaction trust into a platform where buyer and seller shared no language, currency, or legal framework, and she measured the result.
Evidence chain: Alibaba.com, B2B trust redesign at $50B+ GMV. Three cross-functional sprints. +20% transactions, +7% DAU, 47% fewer security concerns. You designed the system that made a buyer in Ohio commit money to a supplier in Shenzhen. OpenAI Payments is designing the system that makes a user comfortable letting an agent spend on their behalf. Different transaction, same problem underneath. TinyFish establishes where you work now and why agent platforms are familiar territory; keep it as context and do not present it as portfolio proof.
Boundary: Alibaba proves cross-border transaction trust at scale. It does not prove agent-initiated spending, usage-based billing design, or autonomous budget management. You have no published evidence of designing spending controls where the user isn't the one clicking buy.
Hardest question they'll ask: "How would you design spending limits and authorization for an agent making purchases on a user's behalf?" Reason it through live. Start from the specific signals you built at Alibaba, whatever made a buyer willing to commit funds to goods she couldn't inspect, and extend that logic to delegation. They want to watch you think.
Do not lead with: TinyFish as portfolio proof. OpenAI has more agent expertise in the building than you can bring; they are not buying yours. Leave the Allē numbers out entirely. This room is reading for trust architecture and loyalty metrics will read as off-topic scale.
Differentiator: Published, measured transaction-trust outcomes at marketplace scale. Candidates coming out of payments backgrounds have mostly optimized checkout flows that already worked. You built the trust layer that made the transaction possible at all.
OpenAI Engineering Acceleration
What the posting tells you the room is solving: A named loop, "instrument, launch, observe, investigate, evaluate, decide, and iterate," spanning "data freshness, coverage, reliability, uncertainty, experiment assignment, metric validity, rollout state, and regressions." It separately names "provenance, permissions, and edge cases," and says outright that this is more than dashboard design.
Claim: Juno has designed an information surface that carries a person from an ambiguous signal to a verified conclusion and a recorded decision, with the evidence chain intact.
Evidence chain: Carrier IQ. The live product shows session, navigate, fill, extract, verify progression across carrier portals, with side-by-side comparison, confidence scoring based on completeness, and an approval flow that keeps the evidence attached: proof attachment, case note, approve bind. It moves an operator from "I have quotes from multiple carriers" to "I can verify this one is correct and approve it with a record of why." That is the investigate-evaluate-decide stretch of the loop they named. The Trust essay supplies the Watch, Verify, Delegate philosophy behind it and explains when human judgment stays in.
The lab page says eight carrier portals; the homepage says five carriers; the live app defaults to five of eight. Reconcile before citing a specific count in conversation.
Boundary: Carrier IQ proves an inspectable decision surface and an approval flow that preserves evidence. It does not prove that marking a result wrong changed the system's behavior on the next run. As covered in Build the Control That Says No, the correction-to-next-run record is absent from the public portfolio. You can describe verification and approval states. You cannot claim a closed feedback loop.
Hardest question they'll ask: "Walk me through a time your design of a data surface led someone to a wrong conclusion, and what you changed." They want the whole loop: wrong output, diagnosis, design change, measurable improvement. If TinyFish production work gives you this, tell it verbally. If it doesn't, say so and describe what you would build instead. Naming the gap yourself costs less than having them find it.
Do not lead with: The correction loop you cannot prove. Say "closed-loop feedback" without a specific instance behind it and the debrief has two versions of you in it, one interviewer who heard a capability and another who heard a hole.
Differentiator: A live, clickable agent-interface product with inspectable verification states. Candidates from observability and developer-tools backgrounds have designed dashboards. Very few have shipped an agent decision surface where a human inspects, verifies, and approves with evidence attached. Moderate confidence this carries a first-round screen. Lower confidence it survives a final round without a verbal production example of the full loop.
Gusto
What the posting tells you the room is solving: Telling apart outputs that are "ready to ship," cases where "a human is needed," and material that "should not be sent out at all." Standards for "trust, uncertainty, graceful failure, and AI-human handoffs." Head of Design title, three reports, strategy authorship, direct prototyping of the hardest flows.
(The ship/review/suppress distinction is named in the posting. That the room is buying judgment authority over where those thresholds sit, rather than interface design around thresholds someone else sets, is my inference from the Head title and the strategy-authorship language.)
Claim: Juno has designed the boundary between automated action and required human review inside a regulated system where sending the wrong thing had real consequences, and kept the human gate functional without turning it into a bottleneck.
Evidence chain: Thermo Fisher, a 0-to-1 pharma supply chain platform launched in twelve months across six partners and nine sites, $20M+ margin opportunity, 100% partner adoption. The Qualified Person release was the human gate. The system automated what it could and pushed exceptions upstream, but the final release decision stayed with a person, because a wrong shipment meant regulatory failure and patient harm. Gusto is solving the same problem in a different domain. The Trust essay's Watch, Verify, Delegate framework gives you the vocabulary for the same architecture: when output ships untouched, when a human verifies first, when the system suppresses entirely. Bring it so the room hears a repeatable model rather than one case you happened to solve.
Boundary: Thermo Fisher proves you have designed the human gate in a regulated automated system. It does not prove you have done it for conversational AI or customer-facing service interactions. The transfer is real but incomplete: pharma exceptions are structured and classifiable, while conversational AI failures read as fluent even when they're wrong.
Hardest question they'll ask: "How do you decide the threshold between 'AI can send this' and 'a human needs to review'?" Answer from the Thermo Fisher decision architecture first, the actual criteria that let a shipment proceed without QP review, then extend to conversational AI and name what breaks in the transfer. They are testing your judgment about boundaries, not your familiarity with their domain.
Do not lead with: GMV or marketplace scale. And do not position as a pure IC. This is a Head role with reports and strategy ownership, and if you open on craft and never reach the org dimension, the recruiter writes "senior IC" on the scorecard. That label travels to the panel ahead of your portfolio and everything you show gets read against it.
Differentiator: Regulated-system work where the human review gate was a design decision you owned. Plenty of AI-product designers have shipped features that use a model. Few have designed the logic that decides when the model's output should never leave the building.
Giga ML
What the posting tells you the room is solving: Enterprise users need "conversational and agentic interfaces that have no established playbook." Output is non-deterministic by design, and the designer owns the patterns that carry that uncertainty, zero-to-one first, then at scale. Staff IC, roughly 20% mentorship, design-system authorship, quality standards.
(The posting says "no established playbook." That the room's real test is whether you have written one, rather than whether you can work without one, is inferred from the Staff-level scope and the emphasis on design-system authorship.)
Claim: Juno has a published design philosophy for human-agent trust, developed by shipping agent products, and it applies directly to interfaces whose output is unpredictable by design.
Evidence chain: The Trust essay sets out the Watch, Verify, Delegate framework and five handoff patterns for human-agent interaction. The live Carrier IQ app shows that philosophy producing a working product: confidence signals, verification states, side-by-side comparison across carrier portals, approval with evidence preserved. TinyFish supplies current production context and nothing more.
Boundary: The Trust essay proves systematic thinking about non-deterministic systems and the live app proves you can ship agent interfaces. But as noted in the Design Context File, the Labs work demonstrates velocity and product judgment while remaining less interrogable at the component level than your traditional cases. If this room's primary test turns out to be visual craft and design-system rigor, the Labs work will not pass it on its own.
Hardest question they'll ask: "Show me something you designed for a non-deterministic system where the output surprised you, and how you handled it in the interface." They want lived experience, so bring a specific moment from Carrier IQ or TinyFish where an agent returned something you didn't expect and you had to choose: surface the uncertainty, suppress the output, or pass it through with a confidence signal.
Do not lead with: Management or team-building. This is a Staff IC role at a growth-stage company, and a directing title at another startup already signals delegation habits they can't afford. Opening with headcount confirms the concern. Lead with what you built and how you think about building it.
Differentiator: A published trust framework for agent interaction with a live shipped product behind it. 91% of designers report weekly AI use in Designer Fund's 2026 survey, so tool fluency buys you nothing here. A public design philosophy for non-deterministic systems, worked out through production, is scarce.
Fieldguide
What the posting tells you the room is solving: Auditors working in AI-assisted workflows have to trust AI-generated findings enough to attach their professional judgment, and eventually their signature, to the result. The designer owns the interaction patterns that make that trust warranted rather than performed. Design-system contribution, peer mentoring. $75M Series C at roughly $700M valuation, 160 employees planning to double.
(The posting names AI-assisted audit workflows and the regulated context. That the core design problem is professional trust sufficient for sign-off is my inference from the domain, not something the posting states.)
Claim: Juno has designed regulated workflows where catching an exception early cost a fraction of catching it late, and she built the surface that moved detection upstream.
Evidence chain: Thermo Fisher again, but the warrant differs from Gusto. For Gusto the case is about the human gate. Here it is about upstream detection: the platform surfaced discrepancies before they reached Qualified Person review, where late resolution reportedly ran three to five times more expensive. Audit runs on the same economics. A discrepancy that surfaces during fieldwork is worth several times one caught at final review, when rework costs compound and the deadline has stopped moving.
Boundary: Thermo Fisher proves upstream exception detection in a regulated workflow. It does not prove audit-domain knowledge, accounting-specific compliance understanding, or experience designing where a professional signature carries legal liability.
Hardest question they'll ask: "How do you design for auditors who need to trust AI-assisted findings enough to sign off on them professionally?" Answer from how Qualified Persons actually made release decisions at Thermo Fisher, specifically what the system put in front of them, and then concede plainly that audit sign-off carries a different kind of exposure.
Do not lead with: The TinyFish directing title. This is a Staff IC at a 160-person company about to double; they need hands on the work and quality shaped through mentorship. The worry here is operational altitude, so lead with the regulated-workflow design itself. Comp is $180K–$210K, materially below the rest of the Act tier. Weight your prioritization accordingly.
Differentiator: Regulated-workflow design where a missed exception carried real downstream cost. Most B2B SaaS designers have made workflows faster. You designed the detection layer that keeps a compliance-governed process from paying for a miss later.
Stripe Link
Clears the rubric (company 10, role 13+, comp $210K–$316K) but original posting date is unconfirmed. Verify the window before investing outreach time.
What the posting tells you the room is solving: "AI-partner surfaces" and "agents transacting for users," alongside zero-to-one product work, 300M+ consumers, and distinctive visual systems. It explicitly requires strong visual and mobile foundations. The role represents Link's vision in senior reviews.
(The posting returns to "distinctive visual systems" and "strong visual/mobile foundations" repeatedly. That visual craft is the primary gate, ahead of systems thinking or strategic scope, is my reading of that emphasis pattern rather than a stated requirement.)
Claim: Juno can design an economic system across two surfaces, consumer and provider, with measurable outcomes on both.
Evidence chain: Allē, a dual-surface loyalty redesign across 30M members and 40K practices. 3.2x redemption lift, 47% reactivation, $42 CAC. Allē meant designing for two audiences whose goals pulled against each other, consumers redeeming rewards and providers driving visits, inside one economic system. Link has the same structural shape with a third actor added: consumers want frictionless payment, merchants want conversion, and agents now transact on both their behalf.
Boundary: Allē proves dual-surface economic design with results. It does not prove the visual-system craft this posting keeps returning to. Equinox+ (zero-to-one in 90 days, 4.8-star App Store, five brands) adds consumer craft evidence, though it may still not satisfy a room testing specifically for a distinctive visual point of view. If the portfolio review turns out to be primarily a craft test, prepare on that assumption.
Hardest question they'll ask: "Show me the visual system you'd build for Link. What makes it distinctive?" This is a making test, so arrive with a position on how Link should look and feel, tied to the product's actual difficulty: establishing trust in the time it takes to complete one click.
Do not lead with: B2B enterprise work, and specifically not Thermo Fisher. This room is buying consumer craft and visual distinction. Open on Allē's consumer outcomes and Equinox+'s speed to a shipped, well-rated product.
Differentiator: Dual-surface economic system design with published outcomes on both sides. Consumer fintech designers usually own one surface. You designed the logic connecting both and have the metrics showing it worked.
Cross-Reference: Exhibits by Company
| Exhibit | Companies | How the warrant differs |
|---|---|---|
| Thermo Fisher | Gusto, Fieldguide | Gusto: the human gate, which outputs ship, which need review, which get suppressed. Fieldguide: upstream exception detection, catching discrepancies before late-stage cost multiplies. |
| Carrier IQ | Eng. Acceleration (primary), Giga ML (secondary) | Eng. Accel: trace from ambiguous signal to verified conclusion with evidence chain intact. Giga: proof that the Trust framework produced a shipped product. |
| Trust essay | Giga ML (primary), Gusto (secondary), Eng. Acceleration (supporting) | Giga: published design philosophy for non-deterministic systems. Gusto: systematic vocabulary for the ship/review/suppress architecture. Eng. Accel: when human judgment stays in the loop. |
| Alibaba | OpenAI Payments | Marketplace transaction trust at scale. Not used elsewhere. |
| Allē | Stripe Link | Dual-surface economic system with outcomes on both sides. Not used elsewhere. |
| Equinox+ | Stripe Link (supporting) | Consumer craft speed. Not used elsewhere. |
- Carrier IQ number inconsistency: The lab page, homepage, and live app each report different carrier counts and quote-return times, and prior published work already flagged this — reconcile before any of these rooms ask you to walk through it.
- CHI transparency tension: A peer-reviewed CHI 2026 study found that eight of twelve participants preferred progressive or on-demand transparency over maximal process visibility, which directly challenges the Trust essay's emphasis on showing work — prepare a position on how much agent process to surface, because Gusto and Engineering Acceleration will both ask.
- Anthropic's permission model: Anthropic's agent research on trustworthy agents frames permission as action-specific (always allow, require approval, or block) rather than trust-level-based, which strengthens the two-axis confidence/authority revision and gives you a cited industry reference if Gusto or Giga asks how you'd structure autonomy boundaries.
- Stripe window verification: The former Agentic Commerce posting now returns a 404, so treat cached search results as dead signal — and confirm the Link posting's publication date before outreach, since the rubric clears but the fresh-trigger evidence does not.

