The fear predicts the screen
Every senior design hire exists because something is failing. The posting tells you what the company wants built. The failure underneath it, the thing that cannot keep going wrong, is what actually governs how they evaluate you.
Take a posting that says "define the interaction model for agent-assisted workflows." The fear underneath could be that the model's behavior is unreliable and users can't tell when to trust the output. Or it could be that users trust it fine and never come back. The surface language is identical. The portfolio evidence that wins is not.
Feared failures cut across tiers, which is why tier alone is a bad sort. Anthropic Evals and Abridge sit in different tiers and share a fear: unreliable model behavior. Anthropic Evals and OpenAI Codex Growth sit in the same tier and fear nothing alike. Sort by tier and the evidence you lead with will be wrong a meaningful share of the time.
Five fears are visible across current postings. Each has recognition language you can spot on a first read, evidence from your background that addresses it, and evidence that sounds relevant but confirms the concern instead. Diagnose the fear before you choose the lead. The other pieces in this section — tier positioning, evidence grading, process diagnostics — build from what you identify here.
Executive design recruiters and senior hiring specialists are consistent on one point: portfolio evidence gets weighed against the specific organizational problem the role exists to solve. That's high confidence. The five-category split below is my own coding of current posting language, not a validated framework anyone else uses. I trust it as a diagnostic. Actual evaluations are messier than any taxonomy.
1. Unreliable model behavior
Recognition language: Prompts, graders, transcripts, comparison sets, regression diagnosis, model releases, "aligned with what users expect," ground truth, verification, auditability.
Current postings: Anthropic Evals — "keep the model and the product aligned with what users expect" and "what safety requires." Abridge — "maps AI-generated summaries to ground truth, helping providers quickly trust and verify the output."
What the evaluator screens for: Evidence that you've designed for outputs that change, degrade, or contradict themselves across contexts. The model doesn't crash. It drifts. They want to see that you treat the failure as probabilistic and have designed verification, confidence surfacing, and recovery paths around it.
Lead with: Forward-looking artifacts showing the full arc — claim, assumptions, failure signal, changed interaction, recovery path, revision condition. The Trust essay gives you the conceptual foundation with the five handoffs framework. Agentic Labs is the production proof: Carrier IQ's provenance tracking, Brand Pulse's visibility states. TinyFish gives you practitioner context on traces, auditability, and governance; use it verbally, as bridge narrative.
What aggravates it: Generic responsible-AI framing that never touches prompts, evals, regressions, or recovery. A polished happy path with no failed run and no changed behavior. As I covered in "Five Tests Wearing One Name", model-behavior accountability listens for temporal proof: a correction-and-improvement loop. Static artifacts don't answer that, however sophisticated. Your Trust essay is strong conceptual architecture, but presented alone it reads as a position paper to someone screening for this fear. Put Agentic Labs' production states immediately adjacent to it so the framework reads as practiced judgment. That pairing matters most right now, while the forward-looking artifacts aren't yet published and separately linkable.
2. Adoption failure
Recognition language: Discovery, onboarding, activation, engagement, conversion, retention, experiments, "lasting customer value," "everyday use," growth metrics, funnel language.
Current postings: OpenAI Codex Growth — "making powerful AI intuitive, useful, and trustworthy," but the accountability sequence it names runs discovery through retention. Adobe GenStudio — "measure adoption, value, and trust."
What the evaluator screens for: Evidence that you've moved a product from built to used. The fear is that the thing works and nobody cares, or that early interest decays. They want shipped behavior with outcome chains attached: what you changed, what moved, whether it held.
Lead with: Allē for a complete adoption arc — 30M members, dual-surface consumer and provider funnel, 3.2× redemption, $42 CAC. Alibaba for transaction and NPS movement at enterprise scale. Equinox+ for fast consumer product creation where experience quality drove uptake. Your BCG DV portfolio is strongest against this fear. Use it.
What aggravates it: Leading with trust architecture when the posting is asking about funnels. Trust is a condition of sustained adoption, not a substitute for activation and retention evidence. The Agentic Labs apps are working systems, but they publish no verified adoption, retention, or revenue data. Don't position them as adoption proof.
3. Craft and brand incoherence
Recognition language: Visual direction, taste, new visual languages, motion, emotional tone, pixel-level quality, creative bar, "delight," identity, component patterns.
Current postings: Stripe Link — "define visual direction, new visual languages, component patterns, and Link's identity." Babylist AI Builder — "functional and emotional quality." Amplitude — "strong visual design builds trust, clarity, and delight."
What the evaluator screens for: Evidence that you make things that look and feel right, and that you can hold that standard across surfaces, states, and other people's work. The fear is a product that feels generic, inconsistent, or emotionally flat. They will look at your work before they read about it.
Lead with: Equinox+ for consumer craft across five brands. Allē for range across dual surfaces and multiple design systems. Agentic Labs and any inspectable forward-looking artifacts to show current AI interaction and visual judgment. The work has to be visible early. This fear is screened in the portfolio scan, before anyone talks to you.
What aggravates it: Strategy diagrams without close interaction or visual evidence. As I covered in "Four Rooms Doubt Your Title", the Head of Product title already raises the question of whether you still make things. Your portfolio's current presentation leads with strategic framing and outcome metrics, which is the right order for adoption or platform fears and the wrong order here. A craft evaluator landing on junochen.com meets Alibaba's business impact and the Trust essay's framework before seeing any close interaction work. Agentic Labs demonstrates AI judgment; it doesn't necessarily demonstrate the visual range or emotional tone a Stripe Link or Babylist posting is screening for. For this fear, the first thing on screen should be a made thing. If the evaluator scrolls past three frameworks before reaching a screen, you've confirmed the suspicion they arrived with.
4. Platform and system incoherence
Recognition language: "Hold together as a system," coherent model, design system, cross-product consistency, decision rights, governance, standardized patterns, "reduce rework," portfolio-wide, multi-surface.
Current postings: Airwallex Spend — "solutions hold together as a system." PointClickCare — "consistent interaction patterns," "clarified Product/UX decision rights." Adobe Assets — "coherent asset-management, brand-governance, and collaborative workflows across Adobe's portfolio." Render Agent Experience — "consistency across CLI, API, MCP, SDKs, and dashboard."
What the evaluator screens for: Evidence that you've created coherence across products, surfaces, or teams that were previously fragmented. The fear is a product that feels like it was built by twelve teams who never spoke, because it was. What they're checking is whether you understand the problem as organizational rather than visual. Systems break because decision rights are unclear, not because nobody drew the components.
Lead with: Alibaba's redesign across homepage, search, and product-detail experiences, and the operating mandate that made coherence possible in the first place. Support with Thermo Fisher's multi-party operational platform. Allē's five distinct design systems show range. Where the system spans human and agent surfaces (Render, Vanta), the forward-looking Agent Infrastructure work is relevant. Use Cummins/ZED Connect only when connected operations or IoT are directly in play.
What aggravates it: A screen-by-screen tour that never establishes coherence across products or teams. "Built a design system" with no account of how it changed repeated decisions across the organization. And be precise about Alibaba: a full redesign of an outdated, non-localized desktop site, not a multi-market unification. Overstating the scope undercuts a story that's already strong enough.
5. Consequential harm
Recognition language: Uncertainty, errors, incorrect recommendations, user agency, fraud, financial exposure, clinical workflow, verification, patient consequence, "high-confidence decision-making," exception handling, compliance, disbursement, mission-critical.
Current postings: Gusto Finance — "build trust across consequential automated decisions," in payments, identity, fraud, and financial exposure. Abridge — auditable clinical AI affecting patients and care teams; this one carries unreliable model behavior and consequential harm together. Vanta — security monitoring and continuous verification inside compliance workflows.
What the evaluator screens for: Evidence that you've designed where getting it wrong causes real damage — financial, clinical, operational. The fear is that someone loses money, gets hurt, or decides badly because the interface didn't surface the right thing at the right moment. They want to hear that you think in exception paths, human review gates, and what it takes to preserve user agency when automation is pushing for speed.
Lead with: Thermo Fisher mySupply for regulated exception handling with a retained human gate, $20M+ margin, 100% partner adoption. Red Cross for $847K disbursed, six systems consolidated to one, national deployment. Attribute both to your Product Design Director role at BCG Digital Ventures. The Trust essay and the Human-Agent System Design work add depth when the workflow includes delegated or probabilistic action.
What aggravates it: "Human approves" offered as a design pattern without specifying what the reviewer knows, how much time they have, what authority they carry, what evidence they see, and what recovery looks like when they miss something. Treating compliance as a final review step rather than a structural constraint. And calling Red Cross "healthcare." It was national disaster relief, including Service to the Armed Forces. Misstating your own case in a conversation about consequential harm costs you more than it would anywhere else.
"Trust" is five different fears
"Trust" appears in nearly every posting on your list, which makes it the easiest word to misread. The language around it tells you which fear it belongs to:
- Anthropic: "aligned with what users expect" plus "what safety requires" — model behavior reliability
- OpenAI Codex Growth: "trustworthy," but the accountability chain runs discovery through retention through "lasting customer value" — adoption failure
- Gusto: "trust across consequential automated decisions," plus payments, fraud, financial exposure — consequential harm
- Render: "trust, safety, and governance — how access is granted, scoped, and revoked" — platform/system incoherence, specifically control architecture
- Amplitude: "strong visual design builds trust, clarity, and delight" — craft/brand incoherence
Read the words around "trust," then lead with the evidence that addresses that fear specifically. A general trust narrative — "I believe trust is the foundation of great AI products" — sounds relevant to all five and specific to none.
Function-building is a different kind of fear
Several current postings carry an explicit organizational-capability mandate alongside the product one. Amplitude names a 15-person team, hiring, retention, development. PointClickCare names team development, accountability models, operating rhythms. Both Adobe Director roles name organization design, culture, and team scaling. The fear is that the design function itself lacks the talent, structure, or repeatable quality mechanism to produce good work consistently.
I'm keeping this outside the five-category split because it's different in kind. The five fears above are product and customer failures. Function-building failure is the organizational cause underneath several of them, and it's evaluated on different evidence: hiring track record, development mechanisms, operating models, decision-rights architecture. Alibaba shows mandate formation and cross-functional leadership. What your published portfolio doesn't establish is multi-year function leadership with durable reporting lines and resource authority. Know that before you meet it in a conversation. "What Survives the Second Question" went into this: a project-recovery story doesn't substitute for sustained function-building evidence.
Composite mandates
Several postings carry two fears at comparable weight. Abridge combines unreliable model behavior with consequential harm: an auditable clinical summary checked against ground truth inside a workflow that affects patients. Adobe GenStudio combines adoption failure with platform incoherence. PointClickCare layers platform incoherence over consequential harm and puts function-building on top.
When you hit a composite, find which fear the posting's language treats first. The one carrying more operational detail and more named failure conditions is usually the one screened earliest. Lead with evidence for that. Hold the second fear's evidence for the conversation, where you can read which way the evaluator leans.
Signals that the mandate won't deliver
Watch for these in the posting or the first call. They tell you whether the role carries the authority you need.
The fear is real but the authority isn't. Heavy feared-failure language with no mention of decision rights, shipping authority, or a reporting line to someone who can enforce a design decision. The company knows what it's afraid of and hasn't given the role the structural power to address it. This shows up most with consequential harm: the posting describes the stakes in detail and the design role has no gate on what ships. Cross-reference with "The Mandate Checksum" — compare the posting's language against a recent real decision and see whether the advertised authority actually held.
The fear has already been solved. Recognition cue: the posting emphasizes maintaining or scaling an existing system rather than building or fixing one. Most common with platform incoherence at mature enterprise companies, where the design system exists, the governance model works, and the job is custodial. The role may be real. The interesting problem is behind it.
The fear belongs to a different function. If the feared failure is adoption and the design role has no access to growth levers, experimentation infrastructure, or funnel data, the company is hoping interface quality alone will fix a product-management or engineering problem. Adoption is where this mismatch turns up most often: the fear is genuine, the design scope doesn't reach its causes.
Quick-Take Card
Diagnostic move: Read the posting for its feared failure before you read it for its tier. The fear predicts which evidence gets weighted.
Five fears, five leads:
- Unreliable model behavior → forward-looking artifacts + Trust essay + Agentic Labs
- Adoption failure → Allē + Alibaba + Equinox+ (BCG DV portfolio is strongest here)
- Craft/brand incoherence → visual work early, Equinox+ and Allē lead, artifact before strategy
- Platform/system incoherence → Alibaba cross-surface redesign + Thermo Fisher
- Consequential harm → Thermo Fisher + Red Cross (BCG DV), Trust essay for delegated action
Top landmine: A generic trust narrative, when "trust" means five different things depending on the words next to it.
Diagnostic question to ask: "What's the failure mode you're most concerned about in the current product experience?" It surfaces the fear directly.
Composite signal: When a posting carries two fears at equal weight, the one with more operational detail is usually the earlier screen.
- Amplitude's internal design agent: Their Principal Product Designer built and deployed a design agent in two days that generated over 2,219 session snapshots — which means craft-incoherence evaluators at growth-stage companies may now ask how you encode and distribute design judgment through AI tools, not just whether you use them.
- FDA human-factors guidance updated: The August 2026 revision to FDA's usability engineering guidance for medical devices reinforces that intended users, use environments, and hazard analysis must shape design from the start — relevant context for any consequential-harm conversation at Ambience, Abridge, or Maven.
- Portfolio validity research: A revised meta-analysis of selection methods reported .33 operational validity for work samples, with OPM noting that portfolio assessors often cannot verify authorship — which means the mechanism connecting your craft to the outcome matters more than the outcome slide alone.
- Gusto's transitional player-coach signal: Gusto Banking's posting explicitly describes the role as "player-coach, for now" with an expected shift toward people leadership as the business grows — a rare case where the company names the transition instead of leaving you to discover it after you start.

