You've assigned a tier label to a company — AI-native, growth-stage platform, enterprise platform, healthcare/regulated — and that label is now controlling which version of your background you lead with. If the label is wrong, your positioning is wrong, and you will spend thirty minutes selling the solution to a problem the company does not have.
The posting alone cannot confirm or break the label. The checksum is three readings you extract across your first conversations, each testing a different layer of the company's actual design mandate. Together they produce a verdict: label holds, label has partially failed, label has fully failed.
This piece covers detection. What you do after a deviation shows up — which evidence to pivot to, how to reposition — is a separate problem for a separate piece.
The three readings and their diagnostic weight
Reading 1 — Posting against tier prediction. Run before the call. Compare the posting's accountability language to what the tier pattern predicts. Moderate confidence for identifying the advertised territory. Weak confidence for predicting actual authority.
Reading 2 — What the assessment evaluates. Surface during the recruiter screen. Compare what the interview process tests against what the posting advertises. Moderate confidence when you learn specific tasks and scoring dimensions. Low confidence when all you get is format names.
Reading 3 — A recent real decision. Surface during the hiring manager call. Get a specific story about a product decision where design either shaped the outcome, was consulted after the fact, or was absent entirely.
Reading 3 carries the most diagnostic weight across all four tiers. A posting states an aspiration. An assessment can be miscalibrated by accident — someone reused last year's exercise. But a real decision has participants, a sequence, and an outcome, and those are hard to describe inaccurately under a follow-up question. When Reading 3 contradicts the other two, trust Reading 3.
Reading 1. Posting Against Tier Prediction
Run this before you pick up the phone. You already have the posting. The only question is whether its language matches what the tier pattern predicts, or whether the company has labeled itself one thing and then described a different job.
AI-native — what should be there: Responsibility for model behavior, evaluation loops, uncertainty handling, delegation boundaries, supervision, correction, recovery. The unit of accountability is what the AI does and whether that outcome is acceptable.
AI-native — signals the label may have failed: The posting centers on design systems, growth metrics, or visual craft. A company building AI products can still hire designers for work that has nothing to do with model behavior. That's a real job. It isn't the archetype you labeled.
Growth-stage — what should be there: Outcome ownership, 0-to-1 building, direct customer learning, rapid iteration, coherence across a product surface that's expanding faster than the team, senior hands-on execution.
Growth-stage — signals the label may have failed: Compliance language dominates. Cross-product governance or formal people leadership is the primary scope. The posting reads like an enterprise platform role wearing a startup's brand.
Enterprise — what should be there: Cross-product coherence, platform strategy, standards, leadership through managers or senior ICs, impact defined at organization scale.
Enterprise — signals the label may have failed: The role owns a narrow product area with no platform reach, requires direct production work, or reads like a startup builder embedded inside a large company.
Healthcare/regulated — what should be there: Workflow consequence, traceability, evidence requirements, oversight, escalation, safe failure modes, recurring access to clinicians, operators, or compliance specialists.
Healthcare/regulated — signals the label may have failed: The company is regulated but the role owns a growth surface, internal tooling, or an AI product whose primary accountability is model behavior rather than regulated workflow. Operating in healthcare doesn't make every design role a regulated-workflow role.
A posting is a high-reliability source for what a company says it wants and a low-reliability source for what it will actually hire. It's HR language layered over a draft that was often written by someone who has never done the job. Issue #7 covered why "design-led" self-descriptions are structurally unreliable as evidence. Reading 1 tells you which territory the company is advertising. It tells you nothing about whether design has authority inside that territory, or how disputes resolve when design's recommendation costs money.
Reading 2. What the Assessment Evaluates
Format tells you almost nothing. A portfolio review can probe craft, organizational influence, or outcomes. A whiteboard can test model-behavior judgment or a generic consumer workflow. What's diagnostic is whether the content of the assessment matches what the posting advertises.
Ask this in the recruiter screen:
"Can you walk me through the interview stages and what each one is designed to evaluate?"
Listen for content, not format names. Then compare against the tier:
- AI-native: Does any stage evaluate judgment about model behavior, delegation, evaluation, or recovery? If the whole loop tests visual execution and collaboration style, the assessment is testing a different role than the posting describes.
- Growth-stage: Does it test building, outcomes, and speed? Or organizational influence and stakeholder management — enterprise competencies in startup branding?
- Enterprise: Does the loop include a cross-functional panel with product, engineering, and design in the same room? Atlassian publishes the clearest example: a PED triad that discusses tradeoffs, design's role, and collaboration. A cross-functional panel confirms partnership is being evaluated. It does not confirm the three functions carry equal authority after you're hired.
- Healthcare/regulated: Does any stage test consequence literacy — what happens when the system fails, how errors get caught, what changes before the next release? If every stage could belong to a consumer SaaS loop, the company may be regulated but the role isn't being evaluated as a regulated-workflow role.
Calibration note. There's no stable relationship between company tier and interview format. AI-native companies are not demonstrably more likely to use system-behavior exercises. Vanta's only attributable design exercises were generic consumer scenarios. Ramp, Headway, and Render don't publicly disclose stable design loops. Salesforce runs different processes by region inside the same employer. OPM assessment guidance is blunt about it: a work sample has validity only when its tasks correspond to the actual work. What matters is whether the exercise content corresponds to the actual work, not what the exercise is called.
A limit on what recruiters can give you. They can usually clarify stages, participants, and stated evaluation criteria. They often don't know the prompt content or the scoring rubric. Take what's available here and reserve the deeper probing for the hiring manager.
Issue #9 drew the line between AI-tool fluency — building with AI — and AI-product fluency, meaning the design of trust boundaries and delegation, and found no named, distinct product-fluency round across OpenAI, Anthropic, or Stripe. If the companies actually building AI products haven't standardized how to test for AI-product design judgment, you can't assume any exercise format will tell you what you need. Reading 2 is moderate confidence at best, and only when you get specifics.
Reading 3. A Recent Real Decision
You can't run this one from a posting or a process description. You need a human to tell you a story about something that happened.
What you're after: get the person across from you to reconstruct a specific recent event where design changed a product direction, was consulted after the direction was already set, or wasn't involved. Whichever it is, the answer places design in the company's real authority structure, independent of what the posting says or the assessment tests.
Three phrasings, each surfacing a different angle:
Process reconstruction:
"Walk me through how a recent product decision got made, from idea to shipped feature. Who was involved at each stage?"
This avoids asking whether the company values design, which produces a rehearsed answer every time. It requires the speaker to place design in an actual sequence, and the sequence shows whether design entered early enough to shape the decision or late enough to execute somebody else's.
Trajectory change:
"When was the last time design or research findings changed the direction of a project? What happened?"
Listen for a decision not to build something, an explicit tradeoff, a scope change driven by design evidence. If the person can't name an example, that absence is the finding.
Conflict resolution:
"If design and engineering disagree on an approach, how does that typically resolve?"
Then: "Could you walk me through the most recent time that happened?"
The first question gets the aspirational answer. The second converts it into a critical incident — a real event, with real participants and a real outcome.
Tier-specific variants:
- AI-native: "Walk me through the last model or feature release. Where did design first enter the behavior-evaluation loop, and what changed because of that input?"
- Growth-stage: "Take one recent workflow from commitment to launch. When did design enter, and did its evidence change the scope or the business rules?"
- Enterprise: "What's the most recent platform or design-standard decision that caused a product team to change its plan — not accept feedback, but change direction?"
- Healthcare/regulated: "After the last meaningful error or near miss, what changed before the next release, and what part did design play?"
Record six things from the answer:
- The decision being described
- When design entered
- What evidence design supplied
- Whether that evidence changed the trajectory
- Who made the final call
- Whether the result became a reusable standard, gate, or evaluation, or was a one-time accommodation
The last item is the most diagnostic of the six. Issue #8 separated momentary influence from authority that becomes repeatable through routines. A design leader who changed one decision is not the same as a design leader whose judgment is now built into how the company releases product.
Sequencing. Ask the process question in the recruiter screen if the recruiter has product context. Save trajectory-change and conflict-resolution for the hiring manager, or for a product or engineering partner who was in the work. Recruiters can clarify reporting lines, assessment stages, and stated success criteria. They usually can't reconstruct a contested product decision.
Confidence level. Moderate to high when you get a specific incident with named participants and observable outcomes. Low when the answer stays abstract — "we collaborate closely" — or slides into process description instead of events. The abstract answer is still data. It may simply mean the person doesn't have a concrete example available, which tells you how often design actually moves outcomes there.
What Label Failure Looks Like
An executive recruiter described a 2026 Series B AI startup that posted a VP of Design role. Asked what the hire would own in the first ninety days, the founder listed capabilities rather than decisions. The title existed before the mandate did.
Peter Merholz described a company that brought in a senior design leader to fix product quality. When the leader's changes created friction, executives told them to back off — while supporting comparable disruption from newly hired Product and Product Marketing leaders.
Expedia publicly introduced Doug Powell as VP of Design Practice Management with a transformation mission. A reorganization dissolved both the mission and the role. When IBM's General Manager of Design described her scope, she explained that her budget covered only a fraction of the organization; other general managers funded the rest. Both positions were later eliminated.
Those are mandate collapses: authority advertised, then withdrawn or never funded in the first place. The checksum also catches a quieter deviation, where the tier label is correct about the company and wrong about the job. An AI-native company hiring for what is functionally growth-stage building. An enterprise platform hiring for a narrow product area with no platform reach. Both kinds of failure surface the same way, in the recent-decision story, which puts design somewhere other than where the label promised.
Verdict Calibration
After the call, score the label against your three readings. Your Reading 3 notes are the primary input: when design entered, what evidence it supplied, whether that evidence changed the trajectory, who made the final call, whether the result became repeatable.
Label holds. All three readings name the same unit of accountability — the posting advertises model-behavior authority, the assessment tests model-behavior judgment, the hiring manager's example places design inside the behavior-evaluation loop. Two aligned readings plus one genuinely silent reading, where the recruiter didn't know or the assessment details weren't shared, supports a provisional hold at moderate confidence.
Label has partially failed. One reading contradicts the other two. The posting says AI-native, the assessment tests visual craft, but the hiring manager's story places design in the model-evaluation loop. Or the posting and assessment align cleanly and the recent-decision story reveals that design's evidence was heard and then ignored. The label describes some of the role and not the operating reality.
Label has fully failed. The recent decision places design in a different role than the posting advertises, and the assessment reinforces that different role. The company is AI-native and the design mandate is growth-stage building. The company is enterprise and the role is a narrow product area with no platform reach. The label you assigned doesn't describe the job you'd actually be doing.
If you couldn't surface a reading, you have missing data, not a negative finding. Preserve the unknown and design your next conversation to close it.
Run the checksum on every company you've labeled, and run it early enough that the verdict can still change what you say in the next round.
- Render's candid authority language: Render's Agent Experience posting says the designer "owns" agent interfaces, trust, safety, and governance while also stating they must align the organization "without formal authority" — a posting that names the gap before you have to diagnose it.
- Atlassian's PED triad structure: Atlassian's design interview handbook is the only target-company process that documents a cross-functional panel with product and engineering at the senior level, making it the clearest public example of partnership as an explicit evaluation object.
- McKinsey on org charts: A McKinsey analysis drawing on millions of professional profiles and 250+ design-leader surveys found little correlation between design-team size and financial performance, reinforcing why Reading 3's decision-level evidence outweighs structural signals like headcount or reporting line.
- The airline case study: A peer-reviewed embedded case study of a legacy airline with a CX mandate on the management board found that the digital department doing the work had no hierarchical link to operational divisions, and three of four studied projects stalled or transferred — the strongest-quality example of a mandate that existed on paper but not in practice.

