Five evaluation archetypes cut across every company tier. Read them off the posting and the recruiter screen. The panel is scoring a role-level scorecard, and the tier tag will not tell you which one.
Picking the wrong company is recoverable. Walking into a first conversation with the wrong evaluation model loaded is not, because you get graded against criteria you never saw and you never find out. You brought scale; they were testing whether you can specify a component out loud. You brought judgment under uncertainty; they wanted a headcount number.
I made the narrow version of this argument in Where Design Authority Actually Lives: the splits inside a tier matter more than the tier does. This piece is the instrument for reading those splits. Not which company evaluates you how. Which of five evaluation archetypes a specific role belongs to, and how to settle that from public evidence before you draft a line of outreach.
Sixteen current senior design postings across twelve companies, five published employer interview guides, and one multi-year set of candidate-side interview reports. Where a pattern rests on a single company, I flag it.
The five: model-behavior, technical builder, player-coach, enterprise systems, regulated decision-workflow.
None of this retires the tier tag Signal Watch hands you. Tier tells you what the company is buying and which rubric dimensions are live: AI centrality, equity window, influence ceiling. Archetype tells you what the panel is grading in the room. Both survive. When they disagree, take the archetype, because the archetype is the person sitting across the table. Classification never moves the role score. It changes what you have to prove to earn it.
A single company will run several at once. Adobe currently has a Staff designer whose work centers on adoption of AI tools and direct engagement with model capabilities, alongside a Senior Staff designer charged with uniting identity, asset, brand, and enterprise administration. Both senior IC. Nothing else in common. Salesforce fields a VP who owns long-term platform cohesion next to a Director expected to build code-backed prototypes and run technical spikes.
Two companies, four roles, four scorecards.
Three instruments, run in this order
Ninety seconds per posting. Cheapest signal first.
1. Adjacency. What is this role housed next to?
Not the team name. The named collaborators, and the function the reporting line runs through. OpenAI's Model Designer sits inside Core Models and works with researchers on predicting model behavior. Ashby's Design Engineer reports through Engineering. Ambience's Staff Product Designer works alongside clinicians and clinical AI researchers. Adobe's platform designer works across foundational systems rather than products.
Adjacency is the most reliable of the three, for two reasons: employers state it plainly, and it classifies more roles than either other test. A product designer posted beside prompt and evaluation owners is doing a categorically different job from the same title posted beside a design systems team, and no amount of title-reading gets you there. High confidence.
2. Consequence vocabulary. Does the posting name what breaks, and who absorbs it?
Regulated roles name clinical outcomes, uncertainty, escalation, human judgment. Enterprise roles name fragmentation, cohesion, technical debt. Rubrik's Director posting ties design decisions to technical debt and business outcomes inside one clause. Model-behavior roles name ambiguity and trust in machine output. Technical builder roles name reliability, production quality, adoption.
Now read the absence. A posting with no consequence vocabulary anywhere in it is defining the role by output volume rather than by what the output decides. Downgrade it before you score anything else. High confidence. This language appears verbatim across the corpus, which is what separates it from HR boilerplate.
3. Hands-on distribution. Where does the posting locate the hands-on expectation?
Every senior posting claims hands-on. Location is the differentiator. Clustered at first-of-kind problems, irreversible decisions, and exemplar-setting work, it signals a judgment mandate: they want you in the room when the hard call gets made. Spread evenly across every sprint and every surface, it signals missing capacity. They are one designer short and they dressed the gap in leadership language.
Amplitude's Head of Product Design is the clean version of the first. Fifteen reports, pairs with designers on the hard problems, explicitly not an executive role and explicitly not a pure IC role. Read that as a distribution statement. The title is incidental.
Moderate confidence. Employers never name this distinction, so you are inferring intent from where the verbs cluster. It is the sharpest single test separating technical builder from player-coach, and it will occasionally mislead you.
Model-behavior
Confidence: low. Thinnest evidence base of the five. Hold it loosely.
Recognition. Adjacency to research, model quality, evaluations, or prompts. Consequence vocabulary about ambiguity and trust in machine output. Hands-on clustered at prototyping. The strongest tell: a posting asking a designer for quantitative evidence and data-collection strategy alongside interaction craft.
What the panel needs to say in the debrief. Some version of she treats probabilistic behavior as design material. Not that you have used AI tools. That you can reason about a system whose output varies and still design something a person can trust twice in a row.
What they open first. Inferred, not documented. In a room seeded with researchers, the artifact under review is the system's behavior rather than the screen, so expect them to move past visual work looking for a case where the output was wrong and the interface still held. Lead with the failure state. They will hunt for it otherwise, and finding it themselves costs you the framing.
Loop. Unknown. Admit that to yourself before you assume otherwise. OpenAI publishes that its process may include pair coding, take-homes, or technical tests, followed by four to six hours of final interviews with four to six people, with format varying by team. It does not publish what a Model Designer assessment contains. Anthropic separates technical interviews using Colab or CodeSignal from conversational nontechnical ones, without saying which track design enters. No public source confirms a researcher or evaluation owner sits on a design panel at any of these companies. Ask the recruiter outright. If they name one, that answer is worth more than the entire posting.
Where the screening event falls. The first substantive conversation, almost certainly, and on framing rather than portfolio. I argued this at tier level in Issue #1. It holds harder one level down.
Your coverage: strongest of the five. The published trust essay and the three live Agentic Labs systems are directly inspectable evidence of agent handoffs and human oversight of machine action. The gap is outcomes. The Labs publish no adoption, revenue, or retention numbers, so an evaluator hunting for shipped consequence finds craft and framing and stops there. TinyFish genuinely amplifies here: you are working agent traces, attribution, and governance in production right now, which is current-context credibility no portfolio-only candidate can claim.
Technical builder
Confidence: high on mechanics, moderate on how widely they generalize.
Recognition. Adjacency to Engineering, sometimes reporting into it. Named stacks. Production code, code review, experimentation infrastructure. Hands-on distribution total by design.
What the panel needs to say. She would ship this herself and we would merge it. The organizational failure being repaired is design output stalling at handoff, and someone concluded the fix is a person who produces the artifact rather than specifying it.
What they open first. The system. Not the screens. Ashby runs a dedicated design-system interview with Design Engineering peers, which tells you the artifact under inspection is your component logic, naming, states, and edge cases. A polished case study is not the entry point. The entry point is whether you can defend a specification out loud while three people who maintain one pick at it.
Loop. The most transparent process of any archetype. Ashby's Design Engineer loop opens with a screen-share conversation with the co-founder and VP of Engineering, then either a one-hour technical screen or a roughly three-hour design take-home with live follow-up. Final round: past-project deep dive, design-system interview, and whichever assessment was not used earlier, with three to four Design Engineering peers. Ashby describes the intent as simulating the work. Pair programming, collaborative specification, decision discussion.
DoorDash is the counterexample, and it matters. Its UX Design Engineer posting is unambiguously this archetype. React, backend services, production code, code review. But DoorDash's published design process describes two end-to-end case studies and a panel portfolio review, with no coding assessment disclosed. A build exercise confirms the archetype. Its absence does not reclassify the role.
Where the screening event falls. The build assessment, where one exists. Ashby tests it twice.
Your coverage: this is your structural weak point, and not on ability. On public evidence. The Agentic Labs systems are described as solo-built and live, but your published inventory contains no inspectable repository, no implementation history, nothing a peer can open. Postings in this archetype test code artifacts live. Babylist's Staff AI Builder role wants production front-end code alongside design-system evolution and mentorship. Enter this archetype only with a compensating artifact you can hand over. Otherwise skip it and put the hours somewhere they compound. TinyFish lets you talk about this credibly. It does not substitute for the thing itself.
Player-coach
Confidence: high.
Recognition. Explicit team size. Negation language: not an executive role, not a pure IC role. Hands-on clustered at the hardest problems rather than spread across the backlog. Consequence vocabulary about coherence, velocity, and cross-pod consistency rather than about the product's effect on users.
What the panel needs to say. She raises the bar without becoming the bottleneck. Amplitude states the underlying failure outright: coherence across autonomous pods, without creating an approval gate. Adobe's GenStudio Director and Maven's VP of Design describe the same shape in different vocabulary. Lead the team, stay in the work, give line-by-line feedback. Three postings, so treat it as a supported inference rather than an established rule.
What they open first. Atlassian's guide says that for manager candidates the portfolio questioning shifts away from your individual choices and toward how you led the team and shaped the outcome. The reviewer opens on your role in the story, not the artifact. Then separate Product Thinking and Craft Excellence interviewers go back underneath it and audit whether the craft claim survives with the team narrative stripped off. Interviewers submit feedback independently before calibrating, so each one is building a case in isolation and a bad read in round two cannot be repaired in round four.
Where the screening event falls. The portfolio review and its two follow-ups. You have to sustain a leadership claim while two separate people audit product thinking and craft underneath it. Candidates at your level fail here by holding altitude too long.
Your coverage: competitive, with one live liability. Alibaba, Thermo Fisher, and Red Cross carry mandate-creation, cross-functional, and adoption evidence. The Agentic Labs work does not establish team leadership and nobody will read it that way. The liability is TinyFish. A Head of Product returning to design raises an unspoken question about whether you will stay in the seat, and in this archetype it is the question the hiring manager holds while nodding along. Bring the answer before the first call instead of waiting for it: vertical over horizontal, creative impact, wanting to build rather than write requirements.
Enterprise systems
Confidence: moderate to high on scope, moderate on loop.
Recognition. Adjacency to foundational systems rather than products: identity, assets, administration, shared platforms. Consequence vocabulary about cohesion, fragmentation, technical debt. Verbs of alignment and influence in place of verbs of production. An executive audience named explicitly.
What the panel needs to say. She can hold a system view across teams that do not report to her. The failure being repaired is that independent product teams built independently coherent things that no longer resemble each other, and the cost has finally become visible to someone with a budget.
What they open first. This review is a room, not a reader. The Chan Zuckerberg Initiative guide describes a roughly five-person portfolio audience that may include designers, PMs, engineers, researchers, or a subject-matter expert, followed by four or five individual interviews including an engineer, a PM, and a hypothetical whiteboard case with a designer. Rooms scan differently than individuals. The engineer is checking feasibility. The PM wants to hear trade-offs named. The specialist is watching one thing the other four are ignoring. Structure the walkthrough so each of them finds their answer without asking, because in a five-person audience nobody gets enough airtime to ask twice.
Loop. Atlassian's altitude split shows up here too: the portfolio review interrogates success metrics, individual contribution, leadership, and craft, while separate rounds isolate product thinking from craft. Atlassian confirms Product and Engineering in the loop. No current employer guide confirms a design systems lead or platform engineer as a standard interviewer. If the recruiter names one, treat it as role-specific confirmation rather than something you should have anticipated.
Where the screening event falls. Portfolio artifact quality and altitude control, frequently before a senior human engages with you at all. There is no systems-design exercise here comparable to what engineering hiring runs.
Your coverage: partial. Alibaba supplies real cross-surface coherence at scale. The exposure is that your evidence is stronger on naming a structural gap and shipping change than on years spent scaling a mature internal design organization, and the second is what a Senior Staff platform role at Adobe or a VP role at Salesforce is actually buying. Expect the scrutiny to land there. TinyFish is close to neutral in this archetype and will not offset it.
Regulated decision-workflow
Confidence: high on recognition, low on loop.
Recognition. The consequence vocabulary gives it away and is unmistakable. Clinical outcomes. Uncertainty. Human judgment. Escalation. Fallback behavior. Maven's maternity role asks the designer to define how the system handles uncertainty and how human-AI handoffs work. Ambience's Staff role includes clinician shadowing and postmortems. Adjacency runs to Clinical, Legal, Compliance, or Data.
What the panel needs to say. She has designed the moment someone qualified accepts or rejects the machine's output, and she has been wrong before and handled it.
What they open first. Speculative, given how little of this loop is documented. But the postings point at the exception path themselves. Ambience names postmortems as part of the job. Maven names fallback behavior. Assume the deep dive opens on what happened when your design was wrong in production, and assume the happy path is table stakes rather than evidence.
Loop. Genuinely underdocumented, and the gap should change your behavior. No public source confirms that a clinician, medical director, or compliance owner routinely sits on a product-design panel. The best candidate-side evidence available comes from a company not on your list. Vanta, a compliance-domain company, where Senior Product Designer reports span several years, which is more than any named target offers. Those reports describe a conventional sequence: design leadership conversation, portfolio, a design or problem-solving exercise, a collaboration exercise, PM and engineering rounds. The exercises reported were generic consumer scenarios. One was planning group travel, at a compliance company.
Take that as a diagnostic rather than a grievance. When the exercise has nothing to do with the domain, the domain is not yet the design problem at that company, which tells you what the job will actually be.
Your coverage: strong. Thermo Fisher and Red Cross carry published exception handling, oversight, and consequence evidence, which is exactly the material this archetype evaluates. Allē supplies two-sided workflow structure but is weak as compliance proof and will not survive a push.
Where to spend the hours
| Priority | Archetype | Why |
|---|---|---|
| 1 | Model-behavior | Best coverage-to-competition ratio you have, and the only archetype where the published essay does classification work before you open your mouth. |
| 2 | Regulated decision-workflow | Second-strongest coverage, evaluated on precisely the material Thermo Fisher and Red Cross document. |
| 3 | Player-coach | Competitive on evidence, but part of the loop goes to answering the return-to-design question instead of advancing. |
| 4 | Enterprise systems | Alibaba gets you in the room. The org-scaling gap is real and you will be tested on it. |
| 5 | Technical builder | Skip unless a compensating artifact exists. |
Your strongest archetype has the thinnest documented loop. That makes a single recruiter question, is a researcher or evaluation owner in this panel, disproportionately valuable in model-behavior. It is the only lever you have on a process that is otherwise unreadable from outside.
What gets you killed
- Model-behavior: presenting AI fluency instead of model behavior. Describing what you built using AI reads as tooling literacy, and they are testing whether you can reason about probabilistic output as a material with properties.
- Player-coach: holding altitude in the portfolio review. Atlassian runs separate craft interviews specifically to catch leaders who cannot descend, and answering "how did the team decide" to a question about a specific affordance reads as distance from the work.
- Technical builder: leading with scale. GMV and transaction lift are the wrong currency in front of three Design Engineering peers about to spec a component with you, and scale is no defense against a missing artifact.
- Every archetype: saying "hands-on" without locating it. The posting already told them where they want it, so name where yours sits or the phrase costs you credibility rather than earning any.
- Every archetype: treating "cross-functional panel" as information. It is not. Named functions and what each evaluates are information, and as I covered in Issue #4, the same panel shape can mean peer authority, product oversight, or the absence of a design function entirely.
- Every archetype: positioning TinyFish as design proof. It is not in your published portfolio and cannot be verified by anyone who checks, so it works as current-role context, technical currency, and the bridge story. Nothing past that survives contact with a hiring manager who opens junochen.com.
Deviations on your target list
Stripe. Carried from prior issues and read off posting language rather than re-verified here, so treat it as inference. The Link and Risk roles run two scorecards at once: a zero-to-one interaction mandate and a policy, operations, and trust mandate. Classify each independently. The split shows up in the consequence vocabulary of each posting's first paragraph.
Anthropic. The design-adjacent cluster spans product design, design engineering, prompts, evaluations, model quality, and UI infrastructure, while published hiring guidance splits technical from nontechnical tracks and never says which one design enters. Ask the recruiter which track applies. That one question resolves model-behavior versus technical builder.
Ambience. The Staff Product Designer role is clinical, research-adjacent, enterprise B2B, hands-on, systems-oriented, and mentorship-heavy simultaneously. This is the rule working rather than an exception to it. The tier tag tells you what Ambience is buying and still cannot tell you which of three scorecards is in the room. Make the recruiter walk you through every interviewer before you prepare a thing.
Amplitude. The Head title carries no executive altitude, by the posting's own admission. Read it as player-coach and price the role accordingly.
Red flags, both directions
Classification tells you how you will be evaluated. It says nothing about whether the job is worth winning.
Model-behavior with no design input to model specifications. If nobody who owns evaluations or model quality appears anywhere in the loop, the role probably shapes the interface around behavior it has no ability to influence.
Technical builder reporting through Engineering with no design leader above. Ask who arbitrates when you and an engineering manager disagree on a craft decision. If the answer is the engineering manager, you are staffing a team rather than leading a function.
Player-coach with explicit executive negation. Amplitude is honest about this, which is to its credit. Ask who represents design in planning and headcount conversations, and whether that person has ever reversed a product decision.
Enterprise systems where the portfolio review never asks about metrics ownership. Alignment authority without measurement authority is an advisory role wearing a leadership title.
Regulated in name only. If the design exercise is a generic consumer scenario at a company whose posting is dense with compliance vocabulary, the domain is not yet the design problem there. Expect surface work under a high-stakes title, and price the equity story accordingly.
Quick-take cards
Model-behavior
- Buying: someone to own how machine behavior presents itself to a human.
- Recognize: research or evaluation adjacency; ambiguity and trust in output; quantitative evidence asked of a designer.
- Loop: undocumented. OpenAI confirms only that assessments vary by team.
- Ask: is a researcher or evaluation owner in the loop?
- Screening event: the first substantive conversation.
- They open on: the failure state, not the screens.
- Landmine: describing AI fluency in place of treating model behavior as design material.
- Your coverage: strongest of the five. The Labs publish no adoption outcomes.
- Breaks pattern when: the posting names screens and flows before it names behavior.
- Confidence: low.
Technical builder
- Buying: someone who ships the artifact instead of specifying it.
- Recognize: Engineering adjacency or reporting line; named stacks; production code and code review; hands-on total.
- Loop: Ashby runs pair programming, a technical screen or three-hour take-home, and a design-system interview with three to four peers.
- Ask: is there a build or design-system exercise?
- Screening event: the build assessment.
- They open on: your component system, defended live.
- Landmine: scale metrics in front of a peer build panel.
- Your coverage: public-evidence gap, no inspectable implementation history. Enter with a compensating artifact or skip.
- Breaks pattern when: technical posting, portfolio-only published process (DoorDash).
- Confidence: high on mechanics.
Player-coach
- Buying: coherence across teams that shipped independently, without an approval bottleneck.
- Recognize: explicit team size; negation language; hands-on clustered at the hardest problems.
- Loop: manager-framed portfolio review, then separate Product Thinking and Craft Excellence rounds. Feedback submitted independently before calibration.
- Ask: which interviewer evaluates craft?
- Screening event: the portfolio review and its follow-ups.
- They open on: your role in the story, then audit the craft underneath it.
- Landmine: holding altitude when the question is a craft question.
- Your coverage: competitive. Live liability is whether a Head of Product returning to design stays in the seat. Answer it before the first call.
- Breaks pattern when: team size is unstated.
- Confidence: high.
Enterprise systems
- Buying: someone to unify what independent teams built independently.
- Recognize: adjacency to foundational systems; cohesion and technical debt vocabulary; named executive audience; verbs of influence over production.
- Loop: portfolio review split across altitudes. Product and Engineering confirmed. Design-systems or platform-engineering interviewers not confirmed as typical.
- Ask: who owns the system today, and what did they already try?
- Screening event: portfolio artifact quality and altitude control, often before a senior human engages.
- They open on: a five-person room where each seat wants a different answer.
- Landmine: telling mandate-creation stories when they want organization-scaling stories.
- Your coverage: Alibaba carries the scale. Exposure sits on scaling a mature internal design org.
- Breaks pattern when: the posting names a single product.
- Confidence: moderate to high.
Regulated decision-workflow
- Buying: someone to design the moment a qualified human accepts or rejects machine output.
- Recognize: uncertainty, fallback, escalation, clinical outcomes, human judgment; adjacency to Clinical, Legal, or Compliance.
- Loop: underdocumented. No public source confirms clinicians or compliance owners as routine design interviewers.
- Ask: is a domain expert in the loop?
- Screening event: whichever round contains one. Absent that, the deep dive on how you handled being wrong.
- They open on: the exception path.
- Landmine: treating the domain as context rather than as the binding design constraint.
- Your coverage: strong on published exception and oversight evidence.
- Breaks pattern when: the design exercise is a generic consumer scenario, which means the domain is not yet the design problem there.
- Confidence: high on recognition, low on loop.
-
Explanations can backfire: A 2023 CSCW study found that feature-based explanations did not improve outcomes and increased overreliance when the AI was wrong, which is worth reading before you frame any trust case around explainability — the full paper is open and gives you a citable line for a model-behavior conversation.
-
Friction that works but annoys: A CHI experiment with 199 participants showed cognitive forcing interventions reduced overreliance more effectively than standard explainable-AI approaches, while participants rated those same interventions least favorably — the exact trade-off a regulated decision-workflow panel will ask you to defend.
-
How scorecards get built: Greenhouse's documentation on converting role outcomes into scorecard attributes before interviewing explains the mechanics behind structured hiring, which is the machinery your archetype classification is trying to reverse-engineer from the outside.
-
Unusually explicit mandate evidence: Maven's VP of Design posting publishes its reporting line, team size, executive access, clinical outcome accountability, and compensation band in one place, making it a useful control case for what a posting looks like when a company already knows what it is buying.

