"AI-product fluency" now appears in enough senior design postings that it reads like a single requirement. It is five different requirements sharing a label. If you prepare one version of your story and carry it into all five rooms, you will spend your strongest evidence in the room least equipped to score it.
The difference between the five is not tier and not interview format, though those are the two things most candidates use to calibrate. What actually separates them is what the role is accountable for when the model does something nobody designed for. That accountability splits five ways, and each way listens for different proof, penalizes different gaps, and surfaces in the first ten minutes if you know what to listen for.
This revises the five archetypes from Issue #6. Player-coach and enterprise-systems fold into product-area and function-builder; governance now stands on its own. The revised five: model behavior, technical builder, product area, function builder, governance.
Thermo Fisher, Red Cross, Equinox+, Allē, and Cummins/ZED Connect belong to your Product Design Director role at BCG Digital Ventures. When any of these appear as lead evidence below, present them under that attribution.
Model behavior
Accountable for: How the model responds, improves, and degrades. OpenAI's Model Designer sits inside Core Models and works alongside researchers, shaping behavior through user feedback, quantitative signals, and the research roadmap. The output is changed model behavior. There is often no conventional interface at the end of it. This test lives almost entirely at AI-native companies — OpenAI, Anthropic, Midjourney — where the model is the product.
How to spot it: The first concrete work object named in the conversation is an output, a behavioral variance, an eval, or the way a model change alters what users experience. When the opening minutes are about response quality or data strategy rather than a screen or a team, you are in this room.
What the test looks like: This is the least documented format across your target list. No public source confirms a Model Designer-specific exercise at OpenAI. Reported OpenAI product-design loops include a portfolio follow-up that pushes your decisions against latency, scalability, and actual model performance — whether the design still holds once real model behavior replaces the assumptions in your deck. The whiteboard reportedly asks you to define an ambiguous product incorporating AI, which forces your assumptions about model capability into the open. Expect both pressures to run harder in a model-behavior role.
Confidence on the specific format: low-to-moderate. Confidence that the evaluation centers on what happens when the model behaves outside your design: high.
What to do: Lead with your Trust essay and any cleared forward-looking artifact that exposes model thresholds, degraded states, and the transfer of control. Follow with Alibaba or Thermo Fisher to carry the shipped-at-scale claim. Hold the management narrative until they ask for it. TinyFish gives you current AI production context; use it for that and not as portfolio proof.
Technical builder
Accountable for: A working, shippable system. The evaluation object is executable — code, a prototype, component behavior, implementation decisions. AI-native companies produce these roles (Anthropic's design-engineering inventory), and so do growth-stage companies threading AI into existing products (Babylist, Ashby).
Ashby's Design Engineer loop runs either a one-hour technical screen or a roughly three-hour design take-home with live follow-up, conducted with design-engineering peers. Anthropic draws a hard line between technical and nontechnical hiring: technical roles use live coding environments, nontechnical roles get conversational interviews. Babylist's Director, Product Design (AI Builder) expects the hire to work inside the codebase, stand up prototypes, and contribute production-quality interfaces.
How to spot it: The interviewer asks what you can personally ship, names a repository or a component, or hands you something to build during the call. Questions about your tools and your build speed arriving before questions about your team or your product strategy is the tell.
What the test looks like: You build something, explain it, run it, and defend the trade-offs — and the artifact either works or it doesn't. Ashby's loop simulates the actual work. Anthropic expects candidates to write, run, and debug while narrating decisions.
One distinction here is worth getting right, because misreading it wastes your preparation in both directions. OpenAI's Product Designer, Engineering Acceleration role owns observability and experimentation workflows for technical users — querying, filtering, rollout state, regressions. It demands technical fluency and high interaction craft. It does not explicitly require production coding. That is a technically demanding product-design role rather than a technical-builder role, and the evaluation object is a designed system that makes complexity legible, not a system you personally built and can be asked to modify live. Misread it one way and you spend a week building code samples for a room that wants design judgment. Misread it the other way and you bring design judgment to a room that wants to watch you build.
What to do: Lead with Agentic Labs as your solo-builder claim, and know where it stops — your approved public inventory has no public repository and no code history, so it remains partial evidence anywhere executable code is explicitly tested. Hold the function-building narrative entirely until the room has accepted the making claim. If the role is the Engineering Acceleration variant, lead with Thermo Fisher or Alibaba for complex-system legibility and bridge to Agentic Labs for AI production context.
Product area
Accountable for: A customer segment, a product surface, or a business outcome, end to end. Headway's Staff Product Designer, Group Practices names one customer segment and owns its lifecycle across onboarding, supervisory billing, permissions, packaging, compliance, and analytics. The posting names owners, administrators, supervisors, and supervisees repeatedly, along with the workflows connecting them. It never mentions design headcount or function architecture.
This test appears across every tier. AI-native companies have product areas; so do growth-stage, enterprise, and healthcare companies. Tier won't predict it; read the posting's unit of accountability instead.
How to spot it: The opening problem arrives through a customer, a surface, or a business outcome. Listen for a named segment, a lifecycle, a KPI, a workflow, a product boundary. The interviewer is describing a product group you would join rather than a design organization you would change.
What the test looks like: Portfolio review anchored to a domain, followed by probes about what happened after launch — metrics, customer behavior, what you changed on the second pass. Hightouch's AI Creative role runs a 60-minute live design exercise described as a problem similar to the team's daily work, where the point is showing how you begin breaking it down. It follows a team portfolio review. What is being scored is collaborative decomposition of an unfinished problem, not the polish of a finished one.
What to do: Lead with the shipped case whose product structure most closely matches the area you would own. Allē for paired consumer and provider surfaces. Alibaba for enterprise and marketplace coherence. Thermo Fisher for operational exceptions. Red Cross for consequential multi-role workflow. Add an approved AI artifact when the product itself carries probabilistic behavior or delegated action. Keep the broad function-building claims secondary until the interviewer moves the conversation to team or operating model.
Function builder
Accountable for: The design organization itself — quality, coherence, talent, and the mechanisms that keep multiple product teams aligned. Amplitude's Head of Product Design leads fifteen designers, creates structure inside a messy org, preserves coherence across pods, develops talent, and holds the craft bar. The unit of accountability is the function. Growth-stage and enterprise companies generate this test most often, for the obvious reason: they have enough design headcount that the org has become the problem.
How to spot it: The opening problem arrives through the team or the design system. Listen for an inherited team, a hiring plan, a management layer, inconsistent quality, critique cadence, standards, career development, coherence across autonomous pods. The interviewer is describing an organization to be changed rather than a product group to be joined.
What the test looks like: Atlassian's design interview handbook is the clearest public evidence of how this evaluation diverges from a product-area test. IC candidates get detailed questioning about their design decisions. Management candidates get asked how they led the team and shaped the outcome. Same portfolio format, different follow-up, different scoring surface, and Atlassian states the distinction outright. Expect scenario-based leadership interviews and probes about raising quality without becoming the bottleneck, developing designers, and installing mechanisms that survive your absence.
What to do: Lead with mandate creation, team leadership, standards, and quality that outlasted a single release. Alibaba is your strongest approved evidence for mandate creation and cross-surface coherence. Equinox+ and Allē carry hands-on team leadership and multi-surface work. Use Agentic Labs for AI currency, not as proof you have run a design organization.
Governance
Accountable for: Permissions, escalation, reversal, and evidence in domains where AI acts on behalf of humans and the consequences are real. This is where three-party delegation shows up — provider, agent, patient — and per-action human review becomes structurally impossible, which turns the trust boundary itself into the design problem. Headway's Group Practices role touches this. Brex and Stripe live in financial consequence domains. Ambience and Maven Clinic live in clinical ones. Healthcare, regulated, and financial-infrastructure companies produce this test most reliably, though any company shipping delegated AI action may surface it.
How to spot it: A clinical, compliance, risk, or safety participant appears on the panel. Or the interviewer's framing starts from what happens when the system is wrong rather than what happens when it works. A first question about failure, recovery, or authorization rather than craft or growth puts you here.
What the test looks like: The least documented of the five. No target-company source I reviewed confirms a design exercise built specifically around agent permissions, rollback, or escalation paths. What I expect, at moderate confidence: a portfolio or behavioral probe about consequential failure — what went wrong, who had authority to stop it, how the system recovered, what changed for the next run. The correction-loop artifact, a failure-to-correction-to-next-run sequence, remains the strongest piece of proof you do not yet have. Until you build it, narrate the pattern instead. Describe a specific failure, the gate that caught it, the change you made, and what that change did to the next cycle. Say plainly what the artifact would show. Naming an in-progress piece of work verbally costs you nothing and holds the ground until it exists.
What to do: Lead with Thermo Fisher and Red Cross for exception visibility, human gates, role-specific action, auditability, and consequential workflows. Add the Trust essay to expose permissions, evidence, escalation, and reversal. Note that you are using it differently here than in a model-behavior room: there it demonstrates you understand thresholds and degraded states; here it demonstrates you understand where human authority has to interrupt automated action and how recovery works afterward. Do not use Allē as proof of healthcare compliance. Do not use autonomy language without a visible human review or recovery path attached to it.
When two tests share one role
Babylist and Ramp combine product-area and function-building accountability in a single role. Babylist's Director owns one product area, leads three to four designers, changes the design operating model, and contributes to production interfaces. You may get both tests inside one conversation.
When that happens, listen for which unit the interviewer returns to after a tangent — the product area and its KPIs, or the design organization and how multiple teams work. The one they come back to is the primary test. The other is a qualification check.
One thing from the earlier recognition-cue table holds across all five: interviewer seniority is weak evidence compared to the object they frame the role around. A founder will interview an area owner. A design VP will interview a senior IC. Titles will not tell you which room you are in.
Everyone in these final rounds has a strong background. What separates candidates at this level happens in the first ten minutes, before you have committed to an opening story, while you are still figuring out what the room is accountable for. Diagnose the room first, then choose which story to tell.
- Atlassian's IC-versus-management split: Their design interview handbook is the only public source I found that explicitly documents different portfolio follow-up questions for IC and management candidates using the same format.
- Anthropic's technical routing question: Their careers guidance draws a hard line between technical and nontechnical interview tracks, but does not disclose whether Design Engineer roles enter the technical track — worth confirming with the recruiter in the first call.
- Hightouch's live exercise format: This AI Creative role listing is one of the few that publishes a step-by-step interview sequence including a 60-minute collaborative design exercise simulating daily work, which is useful structural precedent even if the company is not on your target list.
- FDA's decision-support distinction: The January 2026 Clinical Decision Support guidance separates software that supports professional judgment from software that replaces it, which gives you concrete regulatory language for governance-room conversations at Ambience or Maven Clinic.

