Last validated: July 9, 2026
1. The Objection
"Your AI work is interesting, but have you actually designed for model behavior? Like, inside a company where the AI is the product?"
Surfaces most reliably at AI-native companies (Anthropic, OpenAI, Suno) during screening or portfolio review. Also appears at growth-stage companies with new AI product lines when the hiring manager has a frontier-lab candidate in the pipeline and is using that candidate as the mental benchmark.
2. What They're Actually Asking
Three fears stacked inside one question. First: that you'll treat the model as a black box someone else owns and confine your design contribution to the interface chrome around it. Second: that you lack the vocabulary to collaborate with ML engineers on model behavior, evals, and failure modes, and that this will show up as a communication gap in the first month. Third, and this is the one that does the most damage quietly: that your trust essay and Agentic Labs work are thought leadership rather than operational design. That you've written about the problem but haven't sat in the room where the decision was "should the model refuse this request or comply with a caveat" and shipped the answer. If you feel a flicker of recognition reading that third one, good. That's the fear you need to resolve before you walk into the room, because an interviewer will read hesitation on it instantly.
3. The Strongest Case Against You
- Zero tenure at a frontier AI lab. No Anthropic, no OpenAI, no DeepMind, no company where the model itself is the product.
- Agentic Labs projects (Brand Pulse, Retail Velocity, Carrier IQ) are solo-built systems. They demonstrate architectural thinking but not team-scale production with millions of users and model infrastructure at scale.
- The trust essay uses design-leadership vocabulary, not the evals-and-model-behavior language AI companies use internally. A skeptical reader could file it under "framework" rather than "capability."
- The agentic vision codas in your Alibaba and Thermo Fisher cases are explicitly framed as "if I were building this today." Architectural proposals, not shipped systems.
4. The Grain of Truth
You have not worked inside a frontier lab. That is an actual gap. You won't reframe it away. You have not shipped evals that measure alignment drift across model versions. Acknowledge this cleanly and without hedging. Then ask what the role actually requires — and whether your background covers the harder part of it more deeply than most lab-internal candidates can.
5. The Reframe
Read what AI companies actually post for. OpenAI's Agentic Risk Analyst role describes mapping material risks to mitigations, dependencies, decisions, residual gaps, escalation paths, and launch readiness across multi-step autonomous workflows. Anthropic's Model Behaviors role asks for taxonomies of behavior issues and reinforcement signals. Their Agent Prompts & Evals team owns eval frameworks, dashboards, and regression detection before ship. The common thread across all of these postings is human oversight of automated systems that act under consequence. The control layer between the system's action and the human who is accountable for the outcome.
Your published work maps to that control layer with a specificity most design candidates cannot match. The trust essay's five handoffs name the exact interaction points where oversight operates: Intent-Setting (structured intake before execution), In-Progress (visibility into reasoning as it accumulates), Output Review (what sources, what confidence, what the system didn't find), Decision Gate (automate what you can verify, keep a human where the outcome is irreversible), Loop Feedback (one output becomes the next input, small errors compound). The labels differ from what these companies use internally, but the architectural problem is the same, described from the user-facing side. And the HCI research makes this technically harder to dismiss than most interviewers realize: Chen et al. found that feature-based explanations, the standard "explainable AI" approach, can actually increase overreliance on AI outputs. Bucinca et al.'s cognitive-forcing work confirms the implication: the design problem worth solving is changing how users evaluate and act. Your trust ladder (Watch → Verify → Delegate) is calibration architecture with behavioral teeth. Most design candidates don't draw that line at all.
Your portfolio carries evidence that lab-internal candidates rarely bring. The Thermo Fisher coda turns five supply-chain modules into five continuous agents with exception routing, confidence scoring, and one irreplaceable human gate at batch QA release because a regulatory signature cannot be automated. The Alibaba coda designs the full loop from natural-language procurement brief to autonomous transaction with trust-signal evaluation at every surface. These are proposals, not shipped systems. But they demonstrate agentic oversight design for domains where errors carry regulatory, financial, or operational consequences that don't roll back. The timeline strengthens the case: the Thermo Fisher platform shipped in 2021, the Alibaba redesign and Red Cross national deployment before that. You have been designing for consequential automation longer than most AI-native companies have existed in their current form. The lab gap is real. So is the oversight-design credential, and that one takes longer to build. High confidence this reframe lands across all three evidence layers.
Tier Notes:
- AI-native (Anthropic, OpenAI): Lead with the trust essay's framework mapped to their posting language. They know they need this. They may not expect a design candidate to have already named it.
- Growth-stage (Ramp, Headway): Lead with Agentic Labs as proof you build, not just theorize. The live systems answer "can she ship AI product" before the trust framework answers "does she think about it deeply."
- Enterprise (Salesforce, Atlassian): Lead with the Thermo Fisher coda's regulatory gate. Enterprise buyers understand that speed and regulatory integrity must coexist.
6. What to Say
"I haven't worked inside a frontier lab. I've spent years designing the human control layer for consequential automation — regulated domains, high-stakes operations — since before most AI companies existed in their current form. That's the design problem your postings describe."
7. What Not to Say
- "I'm basically doing the same thing as someone at a lab." The work is complementary. Claiming equivalence invites a technical quiz you'll lose, and it signals you don't understand what lab-internal work actually involves.
- "AI is just another tool." Fastest way to confirm the objection. Signals you haven't grappled with non-determinism.
- "I built these AI products with Claude Code." Leading with the build tool trivializes the architectural thinking. The systems carry the argument. Nobody cares what you built them with.
8. Early Signals
- The interviewer asks which models you've worked with, what your eval process looks like, or how you've handled hallucination in production. These are probes for lab-internal experience. Redirect to the trust framework's Output Review handoff and the Thermo Fisher coda's confidence scoring before the question narrows further.
- The interviewer describes the role as "designing for model behavior" rather than "designing the product." This framing privileges lab-internal experience. Reframe early: "The model behavior question and the product design question converge at the control surface. That's where my work lives."
Objection: "Have you actually designed for model behavior inside an AI company?"
Real fear: She'll treat the model as a black box and only design the chrome around it.
Lead with: Trust essay's five handoffs mapped to OpenAI/Anthropic posting language. Same architectural problem, user-facing side.
Say: "I haven't worked inside a lab. I've been designing the human control layer for consequential automation for years. That's the problem your postings describe."
Avoid: "I'm basically doing the same thing as someone at a lab."
-
Anthropic's non-lab precedent: Mike Krieger joined Anthropic as CPO from Instagram and Artifact, not from a research lab, and told The Verge that risks and mitigations were foundational to his reason for joining — a framing that maps directly to your trust-under-consequence positioning.
-
Overreliance research strengthens your case: Chen et al. found that standard feature-based explanations can actually increase overreliance on AI outputs, which means the "just add explainability" approach most design candidates default to is technically wrong — and your calibration architecture is the harder, more defensible position.
-
Backchannel references are intensifying: The Wall Street Journal reported that companies like Zapier now require live reference conversations for senior roles, with up to 10 calls per hire — make sure your "human control layer" framing is something former colleagues would naturally corroborate in their own words.
-
AI skills signal has a design caveat: A 2026 study of 1,725 recruiters found AI skills boosted interview invitations by 8–15 percentage points, but effects were weaker for graphic designers due to recruiter skepticism about AI in creative work — lead with architectural judgment, not AI tool fluency.

