"Head of Design" at an AI-native company carries at least three structurally different mandates under the same title. The difference between them determines whether you spend two years shaping how intelligence behaves or two years polishing the surface above decisions you were never invited to make.
This matters now because the product surface is separating from the product. Agent capabilities are expanding beyond chat into autonomous action, and the layer that governs what the human sees, trusts, and delegates is becoming a product requirement on its own. That layer needs an owner. Whether the company knows it, and whether the title they posted reflects the job they actually need filled, are separate questions with separate answers.
Where does design enter the model or agent behavior conversation at this company? Everything downstream — positioning, lead evidence, risk assessment, and whether to pursue at all — follows from the answer.
The Fork
Pull up any two AI-native job postings with "design" in the title. Read the verbs.
If the verbs are shape, define, evaluate, determine, design enters the behavior conversation. The company is hiring someone to influence what the model does, when it acts, how it recovers, what it refuses. Behavior-shaping mandate.
If the verbs are craft, build, ship, scale, design enters after behavior decisions are made. The company is hiring someone to make the product usable, beautiful, and coherent. Product-surface mandate. Can be a great role. Different role.
A third category is harder to read from postings because most companies lack the vocabulary for it. Some companies need someone to own the architecture of human-agent trust. Trust-architecture mandate. The model team decides what the system can do. The product-surface team decides how it looks. The trust-architecture mandate decides what the human needs to see, and when, to act on what the system produced.
Your published work maps to this third mandate more precisely than to either of the other two. Most companies hiring for it don't yet know they're hiring for it.
What This Tier Is Actually Hiring to Solve
The trust-architecture mandate in practice: An agent acts on your behalf. It books, cancels, summarizes, diagnoses, recommends, or decides. You weren't watching when it happened. Now you're looking at the result and you need to know: was this right? Should I accept it, override it, or undo it? And if I can't tell, what do I do next?
That sequence has at least four design surfaces. The moment of delegation, when the user hands control to the system. The period of autonomous action, when the system is working and the user can't observe every step. Completion, when the result arrives and must be evaluated. And failure, because the model will be wrong and the recovery pattern determines whether the user ever delegates again. Your trust essay's Watch-Verify-Delegate framework names this ladder. The companies on your list are building products that require it. Most of them haven't yet connected the design hire to this problem.
Spotting the Deviations
Not every company on your list fits the pure AI-native pattern where the trust object is an autonomous software agent. Some wear the AI-native label but the trust problem is grounded in physical, regulated, or identity-verification contexts. Recognition cue: if the job posting references hardware, compliance, biometric data, or clinical workflows alongside AI, the trust architecture is hybrid. Don't treat these as disqualifiers. Fewer candidates have designed across that span, and the consequence stakes are higher.
Reading the Eight
Company by company, here is how the fork applies. Confidence labels are explicit because this tier moves fast and public signals decay within weeks.
OpenAI
The clearest evidence that "AI design" is splintering into distinct functions. Current postings show Product Design Manager, Model Designer, and Content Designer as separate tracks. "Model Designer" is the signal: OpenAI has created a role where design enters the behavior conversation by title, without requiring negotiation. Jony Ive and LoveFrom sit above this at the creative-direction layer. The design org is several functions with different reporting lines and different proximity to model decisions.
Confidence: high on the splintering. What you cannot see from public signal is whether these functions coordinate through a single design leader or operate as parallel tracks under product and research leadership. Do not approach with a generic "I'm interested in design at OpenAI" message. Wait for the specific posting that matches your mandate.
Anthropic
Mike Krieger as CPO oversees product engineering, product management, and product design. But the postings carrying the most behavior-shaping language are not titled "design." Model behaviors, computer use, policy design manager. That last one is telling: a "policy design manager" at Anthropic carries language that would be called "design" at most other companies. The behavior-shaping work is happening, but it's distributed across titles that don't register on a design-leader job search. Strongest evidence on your list that reading titles is insufficient at this tier.
Confidence: medium. Anthropic publishes less about its design org than about its research. The absence of a public design blog or design-practice statement is itself a signal. Design may be deeply embedded in product and research rather than operating as a visible, autonomous function. Or it may simply be too early and too small to have a public presence. You can't distinguish these from outside.
Suno
The Head of Product Design posting is a clean product-surface mandate with function-building authority. The role owns Product Design, Design Engineering, and UX Research. The verbs are craft, scale, ship. Meanwhile, model-trust language (detection systems, evaluation models, harmful-content moderation) lives in a separate Trust & Safety Engineering Manager posting. Design and trust are structurally separated in the org as currently posted. No named senior design leader is visible in public sources, which means this is likely a function-building hire from zero.
Confidence: high on mandate type, medium on permanence. Fast-moving orgs restructure. The separation between design and trust may be a snapshot, not a commitment.
Cartesia
One Design Engineer role. No Head of Design. No Product Designer. The design engineer owns web surfaces, UI implementation, React design systems, accessibility, and polish. Model behavior and evaluation authority sits in Research and Evals Operations. The PM role leads 0-to-1 products and lists designers as peer inputs, not co-owners.
Do not pursue at current posting level. Design enters last, owns the least consequential layer, and has no visible path to behavior-shaping authority. Monitor for a future Head of Design posting.
Google DeepMind
Design roles route through Google Careers. The most visible design-adjacent role is UX Research Operations Lead for GeminiApp. No standalone Head of Design or Product Design leadership posting surfaced. The public role taxonomy foregrounds Research, Product Management, and Responsibility.
Confidence: low. DeepMind's design function may exist robustly inside Google's broader design org and simply not surface as a DeepMind-specific track. The disqualifying question is structural: would you report into DeepMind leadership or into Google's design org? These are different jobs at different companies. Clarify before investing any outreach energy.
Tools for Humanity
The Head of Product Design posting is strong. Strategic design vision, team building, craft quality, product partnership. But the design problem is unusual. It spans World App (consumer), World ID (identity protocol), the Orb (biometric hardware), and operator-mediated verification experiences.
This is a trust-architecture mandate where the trust object is identity, not agent behavior. The user question here is "can I prove I'm human?" Your trust essay's Watch-Verify-Delegate framework applies to the Orb enrollment flow more directly than it might first appear. Enrollment is a trust moment with biometric stakes, zero tolerance for ambiguity, and a user who needs to understand what is happening to their data in real time.
Confidence: high on mandate type. Your Red Cross work, consolidating six legacy disaster-assistance systems into one crisis workflow with 1,689 cases opened in the first two weeks, is the closest analog for high-stakes, low-tolerance-for-error interaction design spanning physical and digital contexts.
Plaud
The most interesting signal set on your list. The Head of Hardware Product Design posting is hardware-focused. But the Staff Product Designer posting explicitly names AI-native interaction paradigms, agents, intent-based workflows, and multi-turn dialogues. The Senior Product Designer posting says "translate AI capabilities into clear, usable, trustworthy UX." And a separate ML Engineer, Model Evaluations role owns evaluation harnesses for voice quality.
Plaud is one of the few companies on your list where design and evals are both visible as distinct functions and the design postings carry agent-behavior language. The trust problem is grounded: recorded conversations, transcription accuracy, summary completeness, action-item reliability across clinical, legal, and executive contexts. Errors in a clinical note carry liability.
Confidence: high. This is a trust-architecture mandate with real consequence stakes, and your Thermo Fisher pharma-partner workflow maps directly to their professional-evidence domain.
Heidi Health
The Product & Design department exists, but visible roles are PM-heavy. Deployed Product Manager owns clinician workflow observation, deployment, compliance, and LLM evaluation. Product Manager, Revenue Cycle owns coding accuracy, human-in-loop decisions, and evaluation infrastructure. The design-system work sits in Engineering.
No Head of Design is visible on the public leadership page. The authority you need lives in PM titles here. This is a regulated-consequence domain where your Thermo Fisher and Red Cross evidence would land hard. But confirm that a design leadership role exists or is planned before investing outreach energy. Ask directly. If the answer is vague, the role doesn't exist yet and you'd be selling them on creating it, which is a different conversation with a different timeline.
Positioning by Mandate Type
The company-level guidance above tells you what to say to each specific audience. What follows is how to frame your overall candidacy depending on which mandate type you're walking into.
Behavior-Shaping Rooms
OpenAI Model Designer track, potentially Anthropic under Krieger.
Your candidacy rests on demonstrating that you think in terms of model behavior. Interface decisions follow from behavior decisions, and the room is testing whether you see it that way. The trust essay opens this door because the Watch-Verify-Delegate framework is a design argument about system behavior: when to act, how to verify, how to recover. But frameworks without applied evidence feel academic in these rooms. Always pair it: "Here's the model. Here's Brand Pulse, where 76 live signals become one decision surface. Here's what we learned about verification latency when the agent sees more than the human can."
The risk: they may test for frontier-model fluency you'll need to demonstrate verbally. Know what RLHF and constitutional AI mean as design constraints that shape what you can and cannot build. Know where human evaluation still outperforms automated metrics for user-experience-relevant quality.
Trust-Architecture Rooms
Tools for Humanity, Plaud, Heidi Health if a design role materializes.
Describe their reality back to them. Their own job postings describe the operating environment you've already navigated. Tools for Humanity's Orb enrollment is a trust moment with biometric stakes; lead with Red Cross. Plaud's professional evidence trail maps to Thermo Fisher's pharma-partner workflow where errors carry regulatory cost; lead with Thermo Fisher. Heidi's clinical domain is the same consequence frame; lead with Thermo Fisher, support with Red Cross.
Product-Surface Rooms
Suno at current posting level.
Lead with Alibaba for enterprise-scale consumer product coherence, then Agentic Labs to show AI-era thinking. The trust essay becomes supporting evidence rather than the lead. The risk: you may be overqualified for the behavior conversation they're not having yet, and the role may not give you the authority to start it. That can be fine if the function-building scope is real.
What Gets You Killed
Seven framings that fail with AI-native evaluators.
-
Treating AI as a feature. "I've designed AI-powered products" sounds like you bolted a recommendation engine onto an existing flow. The reframe: you design the operating contract between the human and the system. Brand Pulse is the proof.
-
Leading with visual craft. Everyone in the room assumes you can make things beautiful. What they're evaluating is whether you understand what happens when the model is wrong and the user doesn't know it.
-
Portfolio cases without failure states. Every AI-native evaluator has watched a model fail in production. If your case studies only show the happy path, you look like you've never shipped under uncertainty. Show the failure, the recovery pattern, what the human sees when the system breaks.
-
"Human-centered design" as a philosophy statement. At AI-native companies, this phrase has been drained of meaning by overuse. What lands: specific examples of how you designed the moment where the human overrides, corrects, or rejects the system's output.
-
Conflating speed with recklessness. AI-native companies move fast. The ones building consequential products know that speed without verification is liability. Your speed-under-consequence framing (Red Cross crisis-workflow consolidation, Thermo Fisher across six pharma partners) is the right counter. Do not apologize for caring about getting it right. Speed where errors carry real consequences is harder, rarer, and more valuable than speed where you can roll back.
-
Asking about "design maturity" in the first conversation. At a 50-person AI lab, this question sounds like you need organizational infrastructure to function. Ask instead about where design enters the model conversation. The answer tells you everything "design maturity" would, without triggering the concern.
-
Presenting the trust essay as theory. The essay opens the door. The cases walk through it. Always pair framework with applied evidence, or the room codes you as academic.
Disqualifying Questions
These are for you, not for them. Ask versions of these in the first substantive conversation. If the answers fail, the role lacks the mandate you need.
"When the model produces an output that affects the user, who decides what the user sees?" If the answer is PM writes the spec, engineering implements, design makes it look right: design enters last. Do not continue.
"Does design participate in eval design or review?" If evals are entirely owned by research and ML with no design input on user-experience-relevant evaluation criteria, the behavior layer is closed to you.
"Who does this role report to?" CPO or CEO: real authority is possible. VP Engineering or VP Product with no design peer at the leadership table: you will spend your tenure fighting for a seat instead of using one.
"Is there a budget for design research, or does design borrow research capacity from product?" Borrowed capacity means borrowed authority.
"When the model fails in production, who decides the recovery pattern?" If nobody has thought about this yet, that's actually a good sign. It means the problem is unsolved and the role can own it. If someone has thought about it and the answer is "engineering handles it," the trust-architecture mandate does not exist here regardless of what the job posting says.
Priority Ranking
Act this week:
- Plaud. Staff Product Designer posting carries the strongest agent-behavior design language on your list. The trust problem is grounded in real professional consequence. Your Thermo Fisher evidence maps directly.
- Tools for Humanity. Head of Product Design with clear function-building authority and a trust problem your essay directly addresses.
Prepare for next cycle:
- Suno. Function-building opportunity. Confirm whether the role will have any authority over model-trust decisions before investing heavily.
- Heidi Health. Your consequence-domain evidence is strong. Confirm a design leadership role exists or is planned.
Monitor:
- OpenAI and Anthropic. The design functions are splintering into specialized tracks. Wait for the specific posting that matches your mandate.
- DeepMind. Structural ambiguity between DeepMind-specific and Google-consolidated design. Clarify before engaging.
Do not pursue at current posting level:
- Cartesia. Design enters last, owns the least, no visible path to the authority you need.
The Test
Read the verbs. Ask where design enters. If the company can tell you clearly, evaluate the answer against your disqualifying questions. If the company can't tell you, they haven't decided yet. That ambiguity could be your greatest opportunity or your most expensive mistake. It depends on whether the role gives you the authority to answer the question for them.
- OpenAI's Content Designer role explicitly asks candidates to define user mental models, terminology, and moments where trust or comprehension may break, which makes it a closer trust-architecture match than the Product Design Manager posting.
- Anthropic's computer-use team is building the agent harness that teaches Claude to operate computer interfaces, and their Product Engineer posting collapses product, engineering, design, and model-behavior boundaries into a single role.
- NIST's Generative AI Profile now frames trustworthiness as a lifecycle concern spanning design through decommissioning, and its focus areas include content provenance, pre-deployment testing, and incident disclosure, which gives you governance vocabulary that lands in regulated trust-architecture rooms.
- Headway's Provider Experience director posting is worth watching as a cross-tier comparison because it explicitly asks the design leader to own AI-native features, contribute prompts, and prototype in code, showing how growth-stage regulated companies are absorbing AI-native expectations into consequence-domain mandates.

