Every tier's evaluation process is built around one place where design rationale tends to stop traveling in that domain. For AI-native panels, it's the gap between what you intended and what the model actually did. Growth-stage panels want to know what happened after you shipped. Enterprise panels test whether a decision you made on one product still holds on the three products next to it. Regulated panels test whether you know which decisions can't be taken back.
Nobody uses the word "seam." They say "product thinking," or "craft excellence," or "impact." But reconstruct what they follow up on after the portfolio walkthrough, what they constrain in the whiteboard prompt, what scenarios they reach for in behavioral rounds, and the pattern holds.
This is the evaluator's side of the loop: what gets probed, in what order, and where your profile buys you credit or invites doubt. It also corrects one claim I made about regulated-tier exercises that the evidence doesn't support.
AI-Native — The Strategy-to-Model-Behavior Seam
What they're buying: A designer who already treats the model as material with variable behavior. Output varies, latency varies, capability varies by task. OpenAI, Anthropic, and Stripe's agentic commerce team are hiring people who have absorbed this, not people who will absorb it on the job.
The break they're trained to find
A candidate who can articulate a strategy for an AI product but can't account for what happens when the model doesn't perform as intended. Ian Silber, OpenAI's Head of Product Design, said it plainly in April 2026: they want designers close to the model, trying it, seeing where it breaks, understanding how the behavior can be changed.
Stage by stage
Recruiter screen. Pro forma if your background signals AI adjacency. TinyFish as current role context (Head of Product at an enterprise web agent platform) clears this gate before the question gets asked. The recruiter is sorting you into one of two categories: AI-native practitioner, or traditional designer who's interested in AI. TinyFish answers that on the first line.
Portfolio review. The real test starts here. Reported OpenAI practice includes a 30-45 minute portfolio review, then a portfolio presentation in the onsite. The follow-up questions probe whether your prior design choices survived real constraints: latency, scalability, actual model performance against what you assumed it would be.
Lead with Agentic Labs. Brand Pulse, Retail Velocity, and Carrier IQ are live, working systems you built alone, which answers Silber's question directly. Have you experimented enough to know where the technology breaks? Follow immediately with the Trust essay's five-handoff framework, which names the design problem this tier is hiring for. This tier weights product/design evidence and metric/outcome evidence ahead of strategic or organizational framing. Show what you built and what it did before you describe how you organized to build it.
Whiteboard. Reported format: an ambiguous prompt for a product incorporating AI. The evaluator watches whether you define the problem before reaching for a solution, and whether your solution accounts for the model's failure modes.
This is where TinyFish does its heaviest work verbally. You've shipped three products in three months working daily with agent traces, auditability, and governance. You know what happens when an agent hallucinates inside a live enterprise deployment. That's what this room values most, and it stays spoken. TinyFish work isn't in your published portfolio and can't be presented as design proof.
First conversation with the hiring manager is still the decisive event at this tier, consistent with Issue #5. The hiring manager is forming a theory about whether you design model behavior or use AI tools. That theory forms in the first five minutes. Everything after confirms or contradicts it. The altitude is 50-feet: your relationship to the model as material, demonstrated through specific decisions you made when it misbehaved. Strategic vision gets evaluated, but only after the evaluator believes you've worked close enough to the material to have earned one.
Confidence: Moderate that OpenAI tests model-behavior literacy through whiteboard exercises and portfolio constraint probing. Low-to-moderate that the format is stable across teams. One Glassdoor report describes rejection before any hiring-manager conversation, which conflicts with the reported sequence. No comparably specific interview accounts exist for Anthropic or Stripe design roles.
What gets you killed
"I designed an AI-powered feature that increased engagement by 30%." The word "feature" says you treat AI as a component. The evaluator hears a designer who needs retraining on how to work with models.
Polished mockups of AI interfaces with no degraded states. The immediate question is what happens when the model is slow, wrong, or uncertain. If the case doesn't show that, you've broken at the exact place they're watching.
"I'm passionate about AI and excited to learn." You're competing against people who already know. Enthusiasm about learning says you haven't started.
Leading with org-building or team-scaling. The first filter is model-adjacent craft. An org-building lead buries the signal they're scanning for.
Describing Agentic Labs as clean successes. The Labs prove you built working systems. If you don't describe the model behaviors you had to design around, you've shown tool use rather than material understanding.
Tier deviations
Stripe (Agentic Commerce): The posting sets strategic and craft expectations but publishes no interview mechanics. Stripe's design culture has historically weighted craft depth more heavily than other AI-adjacent companies. Recognition cue: if the first interviewer asks about design systems or component-level decisions, you're in a craft-first evaluation rather than a model-behavior-first one.
Ambience sits on your list as both AI-native and regulated. If the first question is about clinical workflows or patient safety rather than model behavior, treat it as a regulated evaluation regardless of the company's AI identity.
Two-directional red flags
Check whether design reports to a design leader or to an engineering or research lead. If no design executive appears on the leadership page, design may be functioning as a production layer for research decisions. Ask in the first call: "Who does the design team report to, and how does design influence what gets built versus how it gets built?"
If the whiteboard exercise is purely technical — prompt engineering rather than experience design around model behavior — the role is closer to a technical IC than a design leader, whatever the title says.
Growth-Stage — The Artifact-to-Outcome Seam
What they're buying: A player-coach who ships personally, measures what shipped, and iterates on what the measurement shows, at a pace that matches a company burning runway. Ramp, Gusto, Headway, Babylist, and Amplitude are hiring someone to produce results and build the practice that produces them, in that order.
The break they're trained to find
A candidate who shows impressive artifacts but can't connect them to observed outcomes. Or who cites outcomes but can't name the design decision that caused them.
This is the best-documented tier on your list. Babylist, Ramp, and Headway each publish explicit exclusion criteria, and the three triangulate the same underlying test. High confidence on the advertised screen.
Stage by stage
Profile scan. Growth-stage recruiters check two things fast: can this person make things, and have they operated at this pace. Your TinyFish title creates a speed bump. "Head of Product" at a Series A signals velocity and ownership, which helps. Without visible design output attached to it, it signals delegation, which is the primary thing this tier screens against. The recruiter's note carries that ambiguity forward into the ATS.
Preempt it. In written outreach, bridge from TinyFish to design evidence inside the same sentence: "I moved to product to build AI-natively from zero, shipped three products in three months, and I'm returning to design to apply that depth to [specific domain]." Then point at portfolio proof. If the bridge doesn't arrive in the same breath as the title, the title writes the note for you.
Portfolio review, the decisive event. Growth-stage panels spend most of their evaluative energy here, scanning for one sequence: problem, design decision, shipped artifact, measured outcome, next decision informed by the outcome. Any break in that chain raises doubt. The altitude is 50-feet first. Did you make the thing, can you show decisions at the component and interaction level, before the evaluator accepts 10,000-feet claims about strategy or organizational impact. Metric and outcome evidence needs to arrive in the first breath of each case rather than on a closing slide. This tier weights what happened after launch ahead of how you organized to get to launch.
Lead with Allē or Red Cross. Allē gives you the whole chain: dual-surface redesign, 30M members, 3.2× redemption lift, $42 CAC. Red Cross gives you 0-to-1 velocity: six systems consolidated to one, national deployment, $847K disbursed. Both show you stayed with the work through measurement.
For companies with a provider/patient or employer/employee split — Gusto and Headway among them — Allē's dual-surface architecture (30M members on one side, 40K providers on the other) maps onto their structural design problem. Name the parallel out loud. Don't make the evaluator do the translation.
The making test. This tier will probe whether you personally made the work. Babylist's posting says the Director must "enter the codebase and data, build prototypes, and contribute to production-quality interfaces rather than delegating." Ramp expects the Director to "enter implementation details, unblock the team, and ship meaningful work alongside designers, product managers, and engineers."
Agentic Labs is your strongest answer to this probe. You made it, it's live, it works. Deploy it here as making evidence, not as AI-native vision evidence. Different room, different purpose.
Behavioral rounds. These test velocity under ambiguity: tell me about a time you shipped before you had complete information. Your TinyFish record answers with recent, concrete material. Verbally.
What gets you killed
"I set the vision and trust my team to handle execution." Babylist and Ramp both explicitly reject delegation-only leadership.
"We shipped on time, so the project succeeded." Shipment is the midpoint. If the case study ends at launch, you've broken at the seam.
"My first priority is installing the design operating model." Babylist asks for deliberate distinctions between enabling process and burdensome process. Lead with what you shipped, then how you organized.
"I use AI to generate concepts faster, then hand polished Figma files to Engineering." Headway says its designers use Claude, Cursor, and MagicPatterns, and stay involved through launch and iteration directly in code. Static handoff stops before the stages these companies care about.
"My job is to create alignment and approve everything before it ships." Babylist says not to confuse collaboration with consensus.
Tier deviations
Gusto deviates on title and scope. Its "Head of Design" governs three designers inside an 80-plus-person design organization, not the whole function. Recognition cue: a posting that names a small team inside a large existing design org is likely to weight the enterprise test — cross-functional influence and coherence across a broader system — ahead of raw shipping velocity. Ask early: "How much of this role is building the product versus building the organization?"
Amplitude publishes an explicit anti-pattern list, the most transparent exclusion criteria on your target list. If Amplitude is live in your pipeline, read the "this role is not" section before your first call.
Two-directional red flags
If the loop includes no design review — no moment where anyone evaluates your actual craft decisions — the role may be project management with a design title. Ask: "Will I present work in the interview, and who evaluates it?"
If every interviewer asks about process and team-building and nobody asks about a specific design decision you made, they're hiring a manager rather than a player-coach.
Enterprise — The Local-Decision-to-System-Coherence Seam
What they're buying: Judgment that scales without your presence. Rubrik, Salesforce, Atlassian, Adobe, Samsara, and PointClickCare operate at a scale where one design choice — a pattern, a component, an interaction model — gets replicated hundreds of times by dozens of people. They need decisions that hold up when multiplied across products, teams, geographies, and years.
The break they're trained to find
A candidate with excellent judgment on a single product who can't show that judgment holding coherence across a system. "I made a great decision here," with no evidence the decision still works over there.
Stage by stage
Profile scan. Enterprise recruiters check for organizational scale. Alibaba clears it: Head of Design and Research for North America, a $50B+ GMV platform, cross-functional sprints across homepage, search, and PDP. Your recent trajectory (consulting, then a Series A) may raise a scale-recency question. Don't wait for it. In the portfolio presentation, connect Alibaba to a current view of enterprise design problems. The bridge isn't "I used to do this." It's "here's what I learned at Alibaba that I'd apply differently now, given what I've built since."
Portfolio review, the setup. Enterprise reviews are more formally structured. Atlassian's handbook splits the follow-up into two rounds: Product Thinking (understanding products, services, and their business value) and Craft Excellence (technical proficiency in design).
That split means one narrative won't carry you. Alibaba has to work at both altitudes: the strategic rationale (why this redesign, why these three surfaces, what business problem it solved) and the craft execution (specific interaction decisions, specific visual choices, specific trade-offs at the component level). Enterprise panels weight strategic/organizational, product/design, and metric/outcome evidence roughly equally. No single impact type earns disproportionate credit here. Alibaba carries all three (mandate-building, cross-surface redesign, +20% daily transactions), so distribute them across both rounds instead of front-loading one.
The triad, the decisive event. For senior IC and management candidates, Atlassian introduces Product and Engineering counterparts into the loop. The triad discusses how you approach trade-offs, how you see design's role in the decision, and how you collaborate, with mutual working compatibility named as an explicit purpose.
The altitude spans both levels at once — craft rationale and strategic framing — tested under cross-functional pressure. The Product counterpart is checking whether you understand business value trade-offs: can you explain why a design decision serves the business goal, and can you adjust when the goal moves. The Engineering counterpart is checking whether you understand feasibility constraints and can negotiate scope without either capitulating on quality or getting rigid about implementation. Both are watching how you handle disagreement.
Two candidate types show up in this room. The first describes cross-functional work as harmony: great partnership, everyone aligned, collaboration was smooth. The evaluator hears someone who either hasn't faced real trade-offs at scale or smooths over conflict in the retelling, and both predict badly for a role where Product, Engineering, and Design disagree weekly about what ships. The second describes productive tension with clear ownership boundaries: "Product wanted to ship the simplified version to hit the quarterly target. Engineering flagged that the simplified version created technical debt in the component library that would cost us three sprints later. I proposed a middle path that preserved the interaction pattern we'd need for the next two surfaces while reducing build scope by 40%. We gave up the animation layer. I still think that was the right trade-off given the timeline."
That answer shows you know what Design owned in the decision (the interaction pattern), what it didn't own (the timeline, the technical debt assessment), and how to describe a concession without framing it as a loss.
Lead with Alibaba's cross-functional sprint structure. Three sprints across homepage, search, and PDP required coordinating teams with different priorities and constraints. Pick the sprint where the trade-off was hardest. Name what Design owned, what it conceded, and what the concession cost or preserved downstream. If the triad format is unfamiliar — and it may be, since few companies formalize it this way — prepare for the fact that your evaluators are your future peers rather than your future manager. They're asking themselves whether they'd want to resolve a hard trade-off with you next Tuesday.
Confidence: High for Atlassian's documented process. Moderate for connecting the triad to system-coherence judgment. Atlassian's Principal Product Designer mandate for Jira Service Management includes aligning senior leaders around a multi-team AI-native vision and creating scalable frameworks across organizational boundaries, but the interview handbook never says the triad scores "system coherence." Low for generalizing the format to other enterprise platforms. Samsara's Director posting sets cross-product strategy expectations and discloses no interview format at all.
What gets you killed
Presenting a single-product case with no connection to the broader system. The question comes back: that's a great solution for that product — how does it affect the three next to it. Without an answer, you've confirmed the gap they were probing.
Describing design decisions without the trade-off with Product or Engineering. Enterprise evaluators assume every significant decision at this scale involved a negotiation. A case study where design is sole author makes them doubt either your honesty or your experience.
Leading with velocity or 0-to-1. This tier values durability over speed. Lead with decisions that held up over time and across contexts.
Craft excellence without product thinking, or the reverse. Atlassian separates these into different rounds. Alibaba has to serve both.
"I built five design systems" for the Allē work. Apply the custody rule: claim the mechanism (you identified the fragmentation, proposed the consolidation approach), attribute the scope to the program, same breath.
Tier deviations
Atlassian (Developer Tooling AI) breaks the tier by accepting working concepts made with front-end code or AI-assisted development. Recognition cue: a posting that mentions prototyping in code alongside enterprise-scale strategy is telling you there's a making test most enterprise loops don't include.
PointClickCare is healthcare, so the regulated test may run alongside the enterprise one. If the first call opens on patient outcomes or clinical workflows, switch to regulated positioning regardless of the company's scale.
Two-directional red flags
If design reports to VP Product or Engineering with no design executive on the leadership team, design is a service function whatever the posting claims. Check the leadership page before the first call. If no design title appears there, ask: "Who does the Head of Design report to?"
If the loop contains no triad or cross-functional round — if you only talk to designers — the role probably lacks the cross-functional authority the posting implies.
Regulated — The Decision-to-Consequence Seam
What they're buying: A designer who understands that some decisions can't be reversed. Ambience, Maven Clinic, Vanta, and Nourish operate where a mis-displayed dosage or a misrepresented compliance status doesn't cost you engagement metrics. There's no next sprint that undoes it.
The break they should be testing for
A candidate who can design a clean workflow but doesn't instinctively identify which transitions inside it are irreversible: the points where a mistake survives the next screen.
"Should be" is deliberate. This is the tier where the evidence is thinnest, and where I need to correct something.
What the evidence actually shows
In the prior tier playbook, I wrote that a Vanta exercise tests whether a candidate instinctively begins with the failure case. The underlying evidence — Glassdoor reports describing generic consumer scenarios for a different role level — does not support that claim. No attributable product-design candidate account identifies a domain-specific design exercise at Ambience, Maven, or Vanta.
Glassdoor reports from Vanta candidates between 2022 and 2024 describe generic consumer scenarios: a group-travel app, a party-planning wireframe exercise. Those reports cover a different role level than the current Head of Design posting, and they're old. They're also the only concrete named-company exercise evidence available, and they contradict the assumption that a compliance company runs compliance-specific design exercises.
For your preparation: don't assume the exercise will be domain-specific. Assume the behavioral probing will be.
Stage by stage, reconstructed with appropriate uncertainty
Profile scan. Thermo Fisher (pharma, $20M+ margin, exception-first design) and Red Cross (mission-critical, national deployment) signal that you've worked where errors have consequences. Strong currency.
Portfolio review. The consequence test runs even when the format is conventional. The evaluator is listening for how you talk about failure states, edge cases, and error handling. Do you describe what happens when the system works, or what happens when it doesn't?
Lead with Thermo Fisher. The design problem was what happens when the system is wrong: exception-first design in a pharma supply chain where errors carry direct financial and operational consequence. That demonstrates consequence literacy without requiring you to claim clinical expertise you don't have. This tier weights strategic/organizational and metric/outcome evidence ahead of product/design craft. How you structured the decision process, and what the measured consequences were, matters more here than the interaction design. Frame the $20M+ margin and 100% partner adoption as evidence that a consequence-aware approach produced results, not just careful process.
Follow with Red Cross: six systems into one, mission-critical disbursement, regulatory compliance from day one.
Behavioral rounds. The decisive event at this tier can't be identified from public evidence. My inference, moderate confidence, is that it sits in the behavioral rounds, where consequence literacy shows through stories rather than exercises, and where the evaluation runs at 10,000-feet: how you think about consequences systemically rather than how you execute a specific interaction. Expect questions along the lines of a design decision with consequences you didn't anticipate, how you decide when a design needs human review before going live, a time you slowed a shipping timeline because of a risk you'd identified.
Your TinyFish work on auditability, attribution, and governance in enterprise agent deployments gives you recent, concrete answers. An agent taking an action inside a production enterprise environment is a decision with consequences, some of which no redeploy will undo. Draw that parallel out loud.
Domain-stakeholder interaction. Maven's posting requires coordination across Clinical, Design, Product, Engineering, and Data. Ambience's posting includes embedded clinician research. Whether a clinical or compliance stakeholder actually appears in your loop is unknown from public evidence, but both postings make working alongside domain experts a core expectation.
If one does appear, they're testing whether you can work with someone whose expertise constrains your decisions, and whether you treat those constraints as design material or as obstacles.
Confidence: High that the public record does not support claims about domain-specific interview exercises. High that the postings establish consequence literacy as a requirement. Low on the mechanics of how these companies evaluate it in practice.
What gets you killed
Leading with speed. "We shipped in three months" is a virtue at growth-stage companies. Here the first thought is: what did you skip. Red Cross — 0-to-1 in six months with compliance from day one — is disciplined urgency. Frame it that way.
Treating compliance as a constraint you worked around. The evaluator hears someone who'll fight the compliance team instead of designing with them. Compliance is a design material at these companies the same way model behavior is at AI-native companies.
Assuming the interview tests domain knowledge. Studying clinical terminology or compliance frameworks is preparing for the wrong test. The test is whether you think about consequences, irreversibility, and human oversight by reflex, whatever the prompt is about.
Presenting AI work without uncertainty and human deferral. Maven's posting specifically asks for designing how AI expresses uncertainty and defers to human care. The Trust essay's Watch, Verify, Delegate ladder maps onto it directly, but present it inside a clinical or high-stakes context, or they'll file you as AI-native rather than regulated-capable.
Tier deviations
Ambience sits at the intersection of AI-native and regulated. The posting covers high-stakes clinical workflows and AI research collaboration both. The first question in your call tells you which frame they're watching through. Model behavior means AI-native evaluation with a regulated overlay; clinical workflow means regulated evaluation where AI fluency buys attention. Pick the frame the question implies and commit to it. Trying to serve both equally in one conversation dilutes both.
Vanta is a compliance company whose reported design exercises were generic consumer scenarios. Don't assume the Head of Design loop mirrors the Senior Product Designer loop, but prepare for the possibility that the exercise tests general product thinking. If it does, your strength is the behavioral rounds, where Thermo Fisher and Red Cross carry consequence literacy through story.
Two-directional red flags
If the loop includes no interaction with a clinical, compliance, or domain stakeholder, the design team may be walled off from the expertise that should be constraining its work. Ask: "How does the design team interact with clinical or compliance stakeholders in the regular development process?"
If every interviewer emphasizes speed and shipping velocity and nobody mentions quality gates, review processes, or consequence management, the company is operating like a growth-stage startup that happens to sit in a regulated domain. Different risk profile, and a different role than the posting describes.
Quick-Take Cards
AI-Native
Hiring posture: Buying a designer who already treats model behavior as a design material. Lead pillar: Agentic Labs (live AI systems you built) + Trust essay (names the design problem they're hiring for) + TinyFish verbal context (production agent deployment). Top landmine: Presenting polished AI interfaces without addressing what happens when the model fails, slows, or hallucinates. Opening question to ask: "How does the design team's work change when model capabilities shift between releases?" Pattern break cue: First interviewer asks about org-building or team scaling before asking about your relationship to the model — the role may be enterprise-systems in AI-native clothing.
Growth-Stage
Hiring posture: Buying a player-coach who ships personally, measures what shipped, and iterates based on the measurement — at startup speed. Lead pillar: Allē (full artifact-to-outcome chain, dual-surface) or Red Cross (0-to-1 velocity with measurement). Bridge TinyFish title to design evidence in the same sentence. Top landmine: Case study that ends at launch. Shipment is the midpoint. The evaluator wants what happened after. Opening question to ask: "What's the ratio of building new product to iterating on existing product right now?" Pattern break cue: Every interviewer asks about process and team-building, nobody asks about a specific design decision — the role is a manager position, not a player-coach.
Enterprise
Hiring posture: Buying judgment that scales without personal presence — design decisions that hold when multiplied across products, teams, and years. Lead pillar: Alibaba (cross-product coherence, cross-functional sprints, business-scale impact). Prepare the case at both Product Thinking and Craft Excellence altitudes. Top landmine: Presenting a single-product case without connecting it to the system around it. The evaluator's question is always "how does this affect the products next to it?" Opening question to ask: "How does the design team maintain coherence across product lines today, and where does that break down?" Pattern break cue: Posting requires code or technical prototyping — a technical-builder test layered on enterprise-systems, which changes what you prepare for the making round.
Regulated
Hiring posture: Buying consequence literacy — a designer who instinctively identifies which decisions can't be undone. Lead pillar: Thermo Fisher (exception-first design, pharma) + Red Cross (mission-critical, regulatory compliance from day one). Trust essay bridges to AI-in-healthcare contexts. Top landmine: Leading with speed as a value. "We shipped in three months" triggers "what did you skip?" Frame speed as disciplined urgency with compliance from day one. Opening question to ask: "How does the design team interact with clinical or compliance stakeholders in the regular development process?" Pattern break cue: The design exercise is a generic consumer scenario (Vanta precedent) — shift your energy to behavioral rounds, where your consequence literacy shows through stories.
- OpenAI's design hiring signal: Ian Silber described wanting designers who are close to the model, trying it, seeing where it breaks — the most specific first-person account of what AI-native evaluators actually scan for.
- Atlassian's triad format: Their design interview handbook is the only employer-published guide that documents a cross-functional evaluation round with Product and Engineering counterparts for senior candidates.
- Growth-stage exclusion criteria: Babylist's Director posting includes an unusually explicit anti-delegation, anti-consensus stance that triangulates with Ramp and Headway to form the best-documented tier-wide screen.
- NASA on oversight versus theater: Their crew-interface standards require automation to expose system state, override capability, and notification when a decision aid exceeds its competence — a useful framework for evaluating whether a regulated role's "human oversight" claim has substance.

