Last validated: June 30, 2026. Posting activity at AI-native companies typically slows through the July 4th holiday week and picks up mid-July when H2 headcount unlocks. Re-check evaluation criteria if you see a new cluster of design-leadership postings after July 14.
What They Actually Screen For
You need to know what the objection is really about before you can answer it. Read the postings. The actual postings, the job descriptions themselves.
OpenAI's Product Design Manager role asks for "AI-first product behavior," the ability to operate across "multiple altitudes" from pixel-level detail to system architecture and long-term strategy, and experience in both enterprise and consumer contexts. Zero lines about having previously worked at an AI lab.
Anthropic's Design Engineer, Education Labs posting describes the role as "part researcher, part product builder, part interaction designer." It emphasizes inventing new interaction patterns. It measures success by whether users become more capable. It asks for ambiguity tolerance and coalition-building across Product, Design, Research, and Engineering.
Mistral's Product Designer posting asks for designing workflows and interfaces for technical users and enterprise teams deploying AI models securely. Research translation. Cross-functional product definition. Scalable interaction quality across complex product areas.
Runway has open senior product design roles. Suno does not have a citable design-leadership posting as of this date. Watch for a design or creative-direction posting cluster there; that would be the signal that a leadership search is forming, and worth re-checking within 48 hours.
Five evaluation criteria emerge when you read across these postings:
- Product behavior design. Can you design for systems that act, not just display?
- Trust architecture. Can you build the structures that let humans calibrate confidence in AI outputs?
- Ambiguity tolerance. Can you operate when the interaction paradigm does not yet exist?
- Interaction-paradigm invention. Can you create new patterns, not apply established ones?
- System-level judgment. Can you hold the whole system in your head while making decisions at the component level?
Every piece of evidence you present should map to at least one of these. Anything that doesn't is portfolio theater.
The Objection, Stated Plainly
Surface form: "Has she actually built AI products, or is this traditional design with an AI label?"
Underlying fear: They hire you, you get into the room with their models team, and you reach for screen-design patterns that don't apply. You design interfaces for things that should not have interfaces. You treat the AI as a feature inside a product rather than as the product's core behavior. They spend six months teaching you what they need instead of shipping.
You have not shipped from inside a frontier lab. You have not sat in the room where a model's capabilities changed overnight and the design had to change with it. If asked, state this directly. Do not hedge. Do not pretend Agentic Labs is equivalent to building Claude's conversation interface at Anthropic scale.
What you have done is different. The difference is the reframe, and the reframe has three layers of evidence.
Evidence Layer 1: Agentic Labs Live Systems
Brand Pulse, Retail Velocity, and Carrier IQ are solo-built, live agentic systems published at junochen.com. They are running, and an interviewer can open them in a browser tab.
Brand Pulse is real-time brand sentiment monitoring across Reddit and X. Mentions are scored as they land for sentiment, urgency, and source. It produces an hourly live brand-health score and a weekly narrative written by the agent. The design problem here is specific: what does a human need to see when an autonomous agent is continuously interpreting unstructured social data and generating narrative conclusions? There is no established answer to that question. You had to build one.
Satisfies product behavior design (the system acts; the interface communicates that action), trust architecture (sentiment scoring requires visible confidence calibration), and interaction-paradigm invention (no existing pattern covers "agent writes your weekly brand report while you watch the score move").
Retail Velocity is a CPG market-expansion, field-intelligence, and channel-audit system. Field intelligence is the hardest synthesis problem in enterprise AI design: data arrives partial, contradictory, and from sources with different reliability profiles. A channel audit from one region conflicts with sell-through data from another. The design challenge is presenting messy data in a way that supports decisions without false confidence, giving the operator enough signal to act while making the gaps in coverage visible. Satisfies ambiguity tolerance and system-level judgment.
Carrier IQ is InsurTech quote automation, form completion, and coverage selection in a regulated context. When an agent recommends coverage, the human gate is mandatory. The design must make the gate feel like the product itself. Satisfies trust architecture directly.
The shared design question across all three, as you frame it on the site:
"When an agent acts, what does the human watching it need to see?"
Confidence: high this evidence layer lands. These are live, solo-built systems. The interviewer can verify them in a browser tab. Where it requires careful framing: do not overstate scale. These are not products with millions of users. Lead with the design problem each one solves, not with the system's reach.
Evidence Layer 2: "Trust Is the New Interface"
The live homepage may list this essay under a different title ("When Your Users Don't Have Eyes" appears in the May 2026 writing block). Check which title an interviewer would see when visiting the site, so your verbal reference matches the artifact they find.
Your published essay argues that agentic systems move design upstream. You are designing communication of intent, interpretation, and trust across handoffs where the interaction paradigm is still forming.
The essay names five handoffs: Intent-Setting, In-Progress, Output Review, Decision Gate, and Loop Feedback. This is a behavioral model for agent interaction, describing states that each require different design responses. It breaks the agent-human relationship into those states and maps the design work each one demands. Satisfies product behavior design and interaction-paradigm invention directly. This is the kind of framework an Anthropic interviewer would recognize as thinking at the right altitude.
The essay's Watch → Verify → Delegate trust ladder describes graduated permissioning. The user starts by observing the agent. Moves to validating its outputs. Eventually delegates. The design challenge at each stage is different: Watch requires transparency, Verify requires comparison tools, Delegate requires confidence thresholds and rollback. Satisfies trust architecture and aligns with the NIST AI Risk Management Framework and Microsoft's HAX guidelines on making system capabilities clear, supporting correction, and providing global controls. A note on those references: they are for your preparation, not your conversation. Know them so you can recognize when an interviewer is circling the same concepts. Do not cite NIST in a design interview at Anthropic. Speak from your own work; the frameworks are the validation layer you carry silently.
The essay's key claim:
"Accuracy is what the model achieves. Consistency is what the design delivers."
This is a positioning sentence. It draws the line between what the model team owns and what the design leader owns. Use with caution. Whether this lands depends on whether the interviewer sees system-level consistency as a design leadership concern or as a product/engineering concern. At companies where design owns the end-to-end experience layer, this sentence will resonate. At companies where design is scoped to surface-level interaction and product owns system behavior, it may read as overreach. Read the room. If the interviewer has been talking about design's role in shaping product behavior, deploy it. If the conversation has been scoped to craft and interaction quality, hold it.
Confidence: high on the framework, with one caveat. The five-handoff model and trust ladder map cleanly to evaluation criteria. The caveat: essays are theory. Interviewers will want theory connected to shipped decisions. Always bridge from the essay to a specific Agentic Labs system or a specific moment in the Alibaba or Thermo Fisher work where the framework applied. Never let the essay stand alone.
Evidence Layer 3: Agentic Rebuild Visions in Alibaba and Thermo Fisher
These are enterprise platform projects that close with agentic futures grounded in the operational reality you uncovered during the work. That grounding is what separates them from speculative concept work.
Alibaba.com: You led the B2B platform redesign as Head of Design and Research, North America. Core insight from 32 interviews, gaze tracking, and a Baymard audit: 25% of traffic generated 80% of transaction value, but high-value buyers lacked the trust signals and credibility cues needed for procurement decisions. You restructured the experience around B2B trust: credibility above the fold, trust signals at card level, B2B-native filters, inline tier pricing, Trade Assurance in the primary scan zone.
That redesign is context. The case closes with an agentic sourcing loop: natural-language procurement brief, agent traversal across search and PDP, autonomous order completion when buying signals are clear. The trust architecture is what gives the agent permission to commit. Strip it out and the agent is a search tool. You built the infrastructure that makes agent autonomy possible. That progression from trust infrastructure to agent autonomy is what you should walk the interviewer through. It satisfies system-level judgment, trust architecture, and product behavior design (the sourcing loop is a system behavior).
Thermo Fisher mySupply: You led a 12-month 0-to-1 build across six pharma partners and nine sites. Core insight from 32 warehouse-floor interviews: exceptions surfaced at the delivery gate, where they were 3–5x costlier to fix than if caught a week earlier. You converted five failure modes into five modules: exception-first order management, batch kanban, bidirectional KPIs, forecast portal, and an 18-month capacity heatmap.
The case closes with five agents running continuously and one irreplaceable human gate: batch QA release for regulatory signature. This is the most powerful single piece of evidence you have for trust architecture in a regulated context. You designed for full automation everywhere it was safe and drew a hard line at the moment where a human signature carries legal and safety weight. That distinction is exactly what AI-native companies need their design leaders to understand.
Satisfies trust architecture (the human gate as a deliberate design decision), ambiguity tolerance (pharma supply chain with $20M+ annual margin and regulatory clocks), and system-level judgment (83% IRR, 6/6 partner adoption, 42% overhead reduction).
Confidence: high when framed correctly. Frame it as: "I built the trust infrastructure that makes agentic automation possible, and I know where the human gate belongs because I've stood on the warehouse floor and watched what happens when it's missing." Do not frame it as: "I built AI products at Alibaba and Thermo Fisher." You didn't. You built the systems that agentic AI needs to exist on top of, and you designed the transition point. That distinction matters, and getting it wrong will cost you credibility in the room.
Response Language
"So, have you actually built AI products?"
Lead with one sentence: "I've built three live agentic systems from scratch — they're running at junochen.com right now." Then stop. Let the interviewer follow up. Their follow-up tells you which evidence layer to go to next. If they ask what the systems do, go to Brand Pulse and the shared design question. If they push on scale, go to Thermo Fisher. If they ask how you think about the design problem, go to the essay framework. Do not deliver the full case unprompted. High confidence this approach lands. It is concrete, verifiable, and leaves room for a conversation rather than a monologue.
"Your background looks more traditional enterprise. How does that translate to AI-native?"
"The enterprise context is actually the point. Trust architecture for AI is easy to theorize about in consumer chat. Getting it right when someone is committing six months of inventory spend based on what a screen tells them, or when a batch QA release carries regulatory liability — that's where the design work gets real. The trust ladder I use came out of that work." High confidence. This reframes enterprise experience as a strength.
"We need someone who can invent new interaction patterns, not apply existing ones."
"That's what the Agentic Labs work is. There's no established pattern for 'agent writes your weekly brand report while you watch the score move in real time.' There's no template for 'five agents run your pharma supply chain continuously but one human gate is non-negotiable.' I've been inventing those patterns, and I've published the framework that organizes them." Moderate confidence. "Inventing patterns" is a strong claim some interviewers will probe. Be ready with specific design decisions from Brand Pulse or Carrier IQ to back it up. If pressed, go granular: describe a specific interaction state you designed and why no existing pattern applied.
If the IC-versus-manager question surfaces here, that has its own dossier. Do not improvise the answer in this conversation.
Quick Reference — 90-Second Scan
The objection: She hasn't shipped from inside a frontier lab.
The reframe: She's built the thing frontier labs are hiring someone to figure out.
Three evidence layers:
-
Agentic Labs (live at junochen.com). Three solo-built agentic systems. Brand Pulse: live agent monitoring + agent-written narrative. Retail Velocity: field intelligence synthesis under ambiguity. Carrier IQ: regulated automation with human gates. Shared question: what does the human need to see when the agent acts?
-
"Trust Is the New Interface." Five handoffs (Intent-Setting → In-Progress → Output Review → Decision Gate → Loop Feedback). Watch → Verify → Delegate trust ladder. Key line: "Accuracy is what the model achieves. Consistency is what the design delivers" — deploy only when interviewer signals that design owns system-level consistency.
-
Alibaba + Thermo Fisher agentic visions. Alibaba: agentic sourcing loop that works because the trust infrastructure is already in place. Thermo Fisher: five continuous agents, one irreplaceable human gate (batch QA release). Both grounded in 32+ interviews and real operational stakes.
Ready phrases:
- "I've built three live agentic systems. They're running right now."
- "The enterprise context is the point. Trust architecture is easy to theorize about in consumer chat, hard to get right when someone is committing six months of inventory spend."
- "I know where the human gate belongs because I've stood on the warehouse floor and watched what happens when it's missing."
- Suno's careers page silence: No citable design-leadership posting surfaced in this pass, but Suno's Ashby board runs on JavaScript rendering that can hide active roles from standard search, so re-check weekly through mid-July when H2 headcount typically unlocks.
- Anthropic's interaction-pattern language: The Education Labs Design Engineer posting explicitly asks for inventing new interaction patterns rather than optimizing existing ones, which is the closest public validation that Agentic Labs evidence maps to what a frontier lab says it evaluates.
- Runway's senior design openings: Runway's careers page lists Sr./Staff Product Designer and Design Engineer roles, but detailed criteria were not reachable in this pass, so pull the Ashby descriptions directly if Runway moves onto the active target list.
- NIST as silent preparation: The AI Risk Management Framework validates the trust-architecture language in your essay and Agentic Labs framing, but read it for pattern recognition in interviews rather than citation.

