Every company on your list says it wants AI fluency. Two different exams run under that label, and the one you prepare for determines whether your strongest evidence lands or sits unused.
AI-tool fluency: Can you prototype with Claude, ship in Cursor, compress your design workflow with AI, and show what changed? AI-product fluency: Can you design the trust boundaries, recovery paths, delegation thresholds, and evaluation loops that govern an AI system acting on behalf of a user?
Nobody separates these by name. Postings say "AI experience," recruiters say "comfort with AI," hiring managers say they need someone who gets AI. Read which exam you're sitting for before you start answering: from the posting language, from the first ten minutes of conversation, from the tier pattern.
What each test looks like in the room
The tool-fluency test has published formats. You can prepare for it mechanically.
Automattic runs the most explicit version. Candidates answer an application-stage question about AI experience, and those showing none don't advance. Interview questions cover specific tools and workflow changes. The paid trial weights working prototypes built with Claude Code or Cursor over static mockups, and candidates document the augmented workflow alongside the design decisions. That's four places the same criterion gets graded: written evidence, verbal specifics, a working artifact, a process record.
Cursor runs a two-day on-site trial. Candidates pick among real product projects, work independently in the codebase, and present what they built. Design judgment and implementation get evaluated together.
Amplitude says every designer ships code and names Cursor, Claude Code, and V0 among the team's tools. But its interview loop (case study, Design Jam, partnership interviews) has no standalone AI build exercise. The Head of Product Design posting asks what you built and what you learned. AI fluency lives inside the conventional loop as a discussion prompt.
Coinbase does the same. The portfolio introduction asks how you've used recent AI advances to augment your workflow, embedded in an otherwise standard design evaluation.
So the range runs from a conversational prompt to a full work simulation in the company's own codebase, and the interview format itself tells you which version you're facing. A paid trial or take-home with a prototyping expectation is a build test. A portfolio review with an AI question is a discussion test. Both want specifics: tools, workflow changes, decisions, outcomes. Naming the tools without the judgment behind them scores as nothing.
The product-fluency test has no published format. Across the public materials for OpenAI, Anthropic, and Stripe, no employer names a distinct product-fluency round or publishes a separate scorecard for uncertainty, delegation, recovery, or evaluation loops.
What exists instead is role language that describes the product responsibilities without calling them a test. OpenAI's Engineering Acceleration posting names data freshness, coverage, reliability, uncertainty, experiment validity, rollout state, regressions, provenance, permissions, and edge cases. Stripe's Senior Staff Product Designer, Data and AI emphasizes trust, ambiguous zero-to-one work, policies, and edge cases. Stripe's agentic-commerce materials describe permissions and controls as product requirements for agents acting in commercial systems.
Design-leader statements fill in the picture. Ian Silber at OpenAI described staying close to the model, testing where it breaks, and sometimes asking whether the model or the token budget can solve a problem before moving to interface design. Joel Lewenstein at Anthropic said the team looks for candidates who have reconsidered their process, hold strong opinions about what AI software may become across several time horizons, and can move between vision, production code, and user feedback.
The OpenAI interview guide describes a four-to-six-hour final loop covering communication, collaboration, expertise, and work outside the candidate's comfort zone. No product-fluency scorecard appears in it anywhere.
Which means the product-fluency test is implicit. It surfaces through how you talk about your work, what you ask about theirs, and whether your portfolio shows judgment about systems whose behavior is partly probabilistic. You detect it from role language and conversation cues. If the posting names reliability, provenance, permissions, or recovery, or the hiring manager starts describing problems where the system's output is uncertain, you are in a product-fluency evaluation whether anyone announces it or not.
One practitioner account documents receiving a craft-based rejection from a company that, through internal contacts, she later learned had wanted an "AI approach" it never disclosed during the process. The product-fluency test can be the real filter while the stated rejection cites something else.
How each tier weights the two tests
Issue 6 separated model-behavior roles from technical-builder roles and argued that the role archetype determines what the panel grades. This extends that into how the two AI tests distribute across tiers.
AI-native (Anthropic, OpenAI, Stripe, Suno). These companies split internally by role type. A design-engineer or technical-builder posting, one that names Cursor, Claude Code, or shipping code, is primarily a tool-fluency test. A product-design posting that names uncertainty, trust, permissions, or model behavior is primarily a product-fluency test. Both Silber and Lewenstein describe wanting senior candidates who combine the two, but the weighting shifts with the role. Read the posting. Implementation language up front means prepare for a build test. System-behavior language up front means prepare to demonstrate judgment about delegation, recovery, and evaluation. Where a posting leads with both, product-fluency judgment is the differentiator, because tool fluency is assumed at that level. At Director+ the pattern sharpens: the panel takes for granted that you can build, and interview time goes to whether you can reason about what the system should do and why. At Staff IC the weighting is more balanced, and you may face a build artifact and a product-judgment conversation in the same loop. High confidence on this one, corroborated across multiple public statements and posting sequences.
Growth-stage platform (Ramp, Gusto, Headway, Amplitude, Babylist). Tool fluency is becoming table stakes for the operating model. Amplitude saying every designer ships code is an efficiency expectation, not a statement about what the product does. The tool-fluency test at this tier is a discussion prompt inside a conventional loop. Product fluency becomes the primary test when the company's own product involves AI acting for users. Headway's therapist matching and Ramp's expense automation put AI in a delegated-action position where the design problem is supervision, recovery, and trust boundaries. Gusto's payroll automation carries similar consequence weight. Amplitude and Babylist are using AI to accelerate internal workflows and product features, so tool fluency just clears a bar and the real evaluation is whether you can build and lead a design function at growth-stage speed. Sort by the company's product, not the tier label.
Enterprise platform (Salesforce, Atlassian, Adobe, Rubrik). The primary evaluation stays organizational leverage, cross-functional influence, and systems coherence at scale. Neither AI test is the gate. But the tier has shifted enough that a candidate who can't speak to AI credibly now reads as behind. Salesforce has Einstein, Atlassian has Intelligence, Adobe has Firefly. These are shipping AI products, and a design leader who treats AI as someone else's problem signals they haven't kept pace with where the organization is heading. Lead with Alibaba-scale evidence. Bring AI fluency in as proof you can lead design through the transition these companies are actively navigating. At this tier AI fluency is a credibility threshold rather than the scoring surface.
Healthcare/Regulated (Ambience, Maven Clinic, Vanta, Nourish). The product-fluency test here is the consequence-design test in different vocabulary. Decision gates, recovery paths, trust boundaries in domains where errors cost something real — these companies have been evaluating that judgment since long before anyone called it AI-product fluency. A therapist-matching algorithm that fails silently is a patient-safety problem; a compliance automation that hallucinates a policy interpretation is a regulatory one. When you hear "clinical decision support" or "compliance workflow," you're in the product-fluency exam. Thermo Fisher and Red Cross are your lead evidence. The Trust essay's five-handoffs framework maps directly onto how these companies think about system behavior, and making that connection explicit in the room is worth the time: the graduated-trust model, Watch to Verify to Delegate, applies whether the system manages pharmaceutical supply chains or triages patient intake.
Where each test surfaces
Profile scan and portfolio review
Tool fluency surfaces here only if you've made it visible. Agentic Labs shows you built live AI systems independently, which is strong for a discussion-format tool-fluency test. What it doesn't show, as Issue 7 flagged, is a repository, implementation history, or versioned prompt artifacts. For a Cursor-style build trial, that gap matters.
Product fluency surfaces through the Trust essay and the forward-looking artifacts, and this is where the four artifact domains do their most important work. Be precise about what kind of proof they are and where they currently stand. Inference-Aware UX, Intent-Based Interaction, Agent Infrastructure as UX, Human-Agent System Design: these are entering your published inventory, and the first inference-aware artifact is live, but not all four are public yet. Until a domain is published at junochen.com, you can present it in a conversation or a portfolio walkthrough, but you cannot cite it as inspectable proof the way you cite Agentic Labs. A hiring manager at Anthropic or Stripe who lands on a published interaction model for intent-based interaction, or a state diagram for human-agent workflow boundaries, is looking at product-fluency evidence rather than tool-fluency evidence. That distinction determines which artifacts you surface for which test. Tool-fluency evaluation: lead with Agentic Labs. Product-fluency evaluation: lead with published forward-looking artifacts and the Trust essay. Where the relevant artifact isn't public yet, lead with the Trust essay and the Agentic Labs control patterns, then bridge to the shipped cases.
Recruiter screen
At AI-native companies, expect "what AI tools do you use in your workflow?" Lead with TinyFish production context: you shipped an agentic platform from zero to one in three months, working daily with agent traces, auditability, and governance. Past tense. Then name Agentic Labs as the published, inspectable proof.
At enterprise-platform companies, the recruiter is more likely to surface the bridge question first (why you left a Head of Product role, what you're looking for now) before AI fluency comes up at all. Have the TinyFish-to-design bridge ready as your opening frame. At healthcare and regulated companies, the recruiter may screen for domain familiarity before AI enters the conversation; lead with Thermo Fisher and Red Cross to clear that filter.
Product fluency rarely surfaces at the recruiter stage in any tier. If it does — if a recruiter asks about designing for uncertainty or delegation — the hiring manager briefed them specifically on that criterion. Treat it as the primary test.
First hiring-manager conversation
This is where you identify which exam you're sitting for, and you have roughly ten minutes to make the read. Listen to the first problem the hiring manager describes.
Tool-fluency cues: Claude Code, Cursor, code generation, working prototypes, augmented workflow, "what did you build?"
Product-fluency cues: reliability, uncertainty, provenance, permissions, edge cases, rollout state, "testing the model," "finding where it breaks."
A calibration on vocabulary. The same word signals different tests depending on the tier. "Uncertainty" from a hiring manager at OpenAI means model behavior: confidence scores, hallucination rates, output variability. "Uncertainty" from a hiring manager at Maven Clinic means clinical consequence — what happens to a patient when the system is wrong. Both are product-fluency evaluations. The evidence differs. OpenAI wants the Trust essay framework and the forward-looking artifacts. Maven wants Thermo Fisher's exception-first design and Red Cross's mission-critical deployment.
Lead with the matching evidence. Tool fluency: TinyFish production context, then Agentic Labs as inspectable proof. Product fluency: Trust essay framework, then published forward-looking artifacts, then Agentic Labs control patterns, then Thermo Fisher and Red Cross for consequential delivery at scale.
Portfolio deep dive
For tool fluency, annotate how AI contributed to decisions and execution in your existing cases. Automattic weights this explicitly; candidates who document the augmented workflow alongside the design decisions score above those who present finished artifacts with no process visibility. For product fluency, Alibaba, Thermo Fisher, and Red Cross carry the weight: shipped systems where design decisions had consequences, where you navigated multi-stakeholder complexity, where recovery and trust were operational requirements. Connect those patterns forward. The judgment that governed a six-system consolidation at Red Cross is the judgment that governs recovery paths for an agentic system.
Work sample or take-home
A build test (Cursor-style trial, Automattic-style paid trial) is a pure tool-fluency evaluation. Agentic Labs and the TinyFish production work are your preparation foundation, but the assessment is on what you produce in their environment. A design exercise framed around AI product behavior — "design the experience when the model is uncertain," "design the authorization flow for an agent acting on behalf of a user" — is a product-fluency evaluation. Use the Trust essay framework as the analytical scaffold, then design from it.
What gets you killed
Collapsing the two tests into one. Lead with your Cursor prototyping speed in an Anthropic conversation and you've answered a question nobody asked. The reverse fails just as hard: opening with a trust-boundary framework at Amplitude over-indexes on a criterion they weight below shipping velocity.
Presenting the Trust essay as delivery proof. The essay is a framework. It shows how you think about AI-product problems. It does not show that you implemented the interaction states or governed a shipped system. Issue 7 already bounded this. Lead with it for product-fluency positioning, then bridge immediately to Agentic Labs or the shipped cases for implementation evidence.
Citing TinyFish as portfolio proof. TinyFish isn't in your published portfolio and can't be verified by a hiring manager who visits junochen.com. Use it in its four roles — recent role context, technical AI currency, technical grounding for the forward-looking artifacts, bridge narrative — always past tense. A hiring manager who hears you reference TinyFish work and then can't find it will start wondering what else is unverifiable.
Leading with Alibaba at AI-native companies. The Alibaba case is powerful evidence of organizational leadership at scale, but at an AI-native company running a product-fluency evaluation it reads as proof you've operated in large organizations, which isn't the test in the room. Subordinate it to the forward-looking artifacts and the Trust essay. Bring it in when they ask about scale.
Assuming the product-fluency test will announce itself. The practitioner account above is the cautionary case. If the role language names model behavior, uncertainty, or permissions, treat product fluency as the primary test regardless of what anyone says out loud.
Tier deviations
Stripe sits in the AI-native tier, but the Senior Staff Product Designer, Data and AI role reads as a product-fluency test wrapped in enterprise-platform expectations. The posting emphasizes trust, policies, edge cases, and movement between strategy and detailed interaction. If the conversation leads with platform coherence and policy design rather than model exploration, you're in the hybrid. Lead with Thermo Fisher's multi-stakeholder complexity and the Trust essay.
Headway sits in the growth-stage tier, but its product puts AI into a three-party relationship among provider, agent, and patient. That's a product-fluency evaluation wearing growth-stage clothes, and the consequence-design burden is closer to healthcare than to Amplitude or Babylist. Lead with Red Cross and Thermo Fisher, bridge to the Trust essay's delegation framework.
Amplitude says every designer ships code, which sounds tool-fluency-heavy, but the Head of Product Design role is a function-building mandate. The primary test is whether you can build and lead a design organization. The AI-tool expectation is table stakes. Lead with organizational evidence.
Two-directional red flags
The company can't tell you which test they're running. The hiring committee hasn't aligned on what AI fluency means for this role, which means you'll be evaluated against private criteria that differ by interviewer. Ask in the first call: "When you say AI experience, are you looking for someone who builds with AI tools, or someone who designs the product behavior of AI systems?" A vague answer means the committee hasn't done its work.
Tool fluency is the only test and the role is Director+. The company may be hiring an elevated IC under a leadership title. A Director of Design whose primary evaluation surface is prototyping speed isn't being hired to build a function or influence product strategy. Fine if you want an IC role with a senior title. A problem if you want the mandate the title implies.
Product fluency is tested but design doesn't own the trust layer. You'll be graded on judgment you won't have authority to apply. Ask who owns the decision when the model's confidence falls below threshold: design, product, or engineering? A good answer names design as a participant with real influence over that call, even when final authority is shared. "Engineering sets the threshold" or "product decides" means the test is measuring something the role can't deliver on.
Quick-Take Card
The two tests: AI-tool fluency (build with AI, show the workflow) vs. AI-product fluency (design trust, delegation, recovery for AI systems). No employer names them separately. Identify which one from posting language and the first ten minutes of conversation.
Lead evidence by test: Tool fluency: TinyFish context (past tense) + Agentic Labs (inspectable). Product fluency: Trust essay + published forward-looking artifacts (check current inventory) + Agentic Labs + Thermo Fisher/Red Cross.
Top landmine: Collapsing the two tests. Leading with prototyping speed at a company testing product judgment, or trust frameworks at a company testing build velocity.
Opening question: "When you say AI experience for this role, are you looking for someone who builds with AI tools, or someone who designs the product behavior of AI systems?"
Pattern break cue: If the hiring manager describes model-behavior problems and the loop also expects a build artifact, the company is testing both. Lead with product fluency; tool fluency is assumed at that level.
- Automattic's paid trial format: Their design hiring FAQ is the most transparent AI-assessment process any employer has published — worth studying as a template for what build-test preparation looks like when the company actually tells you the rules.
- Anthropic's design-and-code integration: Anthropic practitioners describe using Claude Code to make substantial state-management changes as part of design work, which means their product-fluency and tool-fluency tests may be less separable than at other AI-native companies.
- Recruiter-proposed assessment formats: Designer Fund's 2026 AI in Design report quotes recruiter Garrett Fowler proposing a "fishbowl" build observed by an interviewer and a "rework" exercise revisiting a prior project with AI at its center — formats that haven't been confirmed as adopted by named employers but signal where assessment design is heading.
- The undisclosed-criteria rejection pattern: Xiaofang's firsthand account of learning through internal contacts that a company rejected her for an "AI approach" they never surfaced during the process is the only documented case of its kind in the public record — worth tracking whether more accounts like this emerge.

