
Issue Frame

Recognition cue. Headway, Babylist, Maven Clinic, and Ramp now name Claude Code, Cursor, and MagicPatterns as core practice expectations. Tool fluency is baseline across all four tiers. When a posting says "prototype and iterate with AI," you are being tested on workflow speed. High confidence — verified across current postings in each tier.
Positioning anchor. Establish tool fluency in one sentence — Agentic Labs solo-built, TinyFish production deployment — and move past it. At AI-native companies, tool fluency is assumed. The evaluation is whether you can design review states, recovery flows, and permission models for agent output. Lead with the Trust essay's five handoffs framework and Agentic Labs design decisions.
Landmine. Conflating the two layers. Babylist's posting separates them explicitly: "AI generates the pixels. You define the patterns, behaviors, and quality bar." An interviewer who draws that line will hear conflation as a misread of scope.
Lead question. "When you say AI-first design, are you describing how the team works or what the product does?" The answer tells you which evidence layer to surface next.
AI-Tool Fluency vs. AI-Product Fluency — How to Read Which Test You're Sitting For
Every company on your target list says it wants "AI experience." Two different evaluations run under that label. One tests whether you build with AI tools; the other tests whether you can design the product behavior of AI systems. The posting won't distinguish them and neither will the recruiter, but the evidence you lead with differs sharply, and preparing for the wrong exam turns your strongest proof into background noise. Here's how each tier weights the two tests, where each one surfaces in the interview sequence, and which pillar to lead with once you've made the read.

Five Tests Wearing One Name
"AI-product fluency" now appears in enough senior postings that it reads like one requirement. It is five different requirements sharing a label. A model-behavior room at OpenAI listens for different proof than a governance room at Headway, and both penalize different gaps than a technical-builder loop at Ashby. This piece maps all five: what each is actually scoring, how to identify which one you are in within the first ten minutes, and which evidence layer to lead with. Preparing one story for all five is a positioning error that wastes your strongest evidence.

Your AI Evidence Stack — What You Have, What Each Proves, What to Deploy Now
Every company posting for a senior design leader says "AI experience." They're testing for one of two things: whether you've built and shipped with AI in production, or whether you can design what AI products need to become. Your evidence stack has assets for both, but they live in different places and prove different things. Lead with the wrong one and the evaluator slots you into a category you can't easily reverse. This is the full inventory — what's ready, what's in progress, and which asset to lead with based on who you're talking to this week.

Lead With Builds — How Growth-Stage Companies Evaluate Design Leadership
At a growth-stage company, your BCG DV builds lead and your forward-looking artifacts support. The evaluation sequence tests function-building speed, product judgment under constraint, and cross-functional credibility — with AI-tool fluency as a working-model signal, not a portfolio category. This piece maps that evaluation across Ramp, Gusto, Amplitude, Headway, and Babylist: what each screens for, which of your evidence layers matches, which framings fail despite sounding credible, and three companies that break the growth-stage pattern in ways you need to recognize before the first call. Priority ranking included.