Quick Reference (Scan This First)
Juno, this is the hardest version of the AI credibility question. Everything below is calibrated for OpenAI and Anthropic specifically.
The objection you'll hear:
- "How close have you been to the model itself?"
- "Have you worked directly with researchers to shape how a model behaves?"
- "What's your experience with model behavior versus interface design?"
The fear beneath all three: She designs wrappers around AI. She doesn't shape what the AI actually does.
Your anchor sentence:
"I've built three live agentic systems where the design decisions determine what the agent does, not just what the user sees. The control surfaces, the confidence architecture, the decision gates. Those shape system behavior in production."
At OpenAI, lead with: Carrier IQ. Multiple insurance carriers running in parallel. Your wrapper brief names eight portals; the live app currently runs five agents. Confidence scores based on data completeness. A human evaluation gate before action. The design is the behavior control. Pivot to your published trust framework.
At Anthropic, lead with: The full portfolio. All three labs map to their stated need for turning model capability into product quality. Brand Pulse, Retail Velocity, Carrier IQ together demonstrate what their Claude.ai team describes: build the interactions and moments that turn a capable model into a product people enjoy using. In your case, a product they can trust.
Support with: "Trust Is the New Interface," your published framework for the five handoffs in every agentic workflow. It shows you think about this problem architecturally, not at the screen level.
Use TinyFish for: Verbal currency only. Agent traces, auditability, attribution, governance. Never cite as portfolio proof. Never name product specifics.
Where the reframe is strong: Product behavior, user trust, oversight architecture, interaction paradigms. Confidence: high.
Where the reframe thins: Data collection strategy, research-roadmap collaboration, technical intuition about how training data changes affect model outputs. Confidence: moderate to low. Acknowledge the gap. Turn it into a reason you want the role.
What the Postings Actually Say
Read these the way you'd read a competitor's org chart. What's present matters, but what's absent matters more.
OpenAI's Model Designer role sits in Applied AI. The team "views the model as the product itself." Responsibilities: "collaborating closely with researchers to understand, predict, and design model behavior" and affecting "how models interact and resonate with users, balance model capabilities, interpret user queries, and uphold user trust."
The "you might thrive" criteria: taste, creativity, writing, ambiguity tolerance, experimentation, empathy, philosophical clarity, and "technical intuition about how data changes can affect model behavior."
Now notice what is not there. RLHF. Fine-tuning. Constitutional AI. Reinforcement learning. Model training. The role lives in the space between interface design and ML research, and OpenAI placed it there on purpose.
Anthropic's Claude.ai posting is a different animal entirely. Staff Software Engineer. The company says plainly it is "a product engineering role first and foremost." The team "builds the consumer web app interfaces, interactions, and moments that turn Claude from a capable model into a product people enjoy using." Shipping features. Owning quality. Treating latency and reliability as first-class concerns. Advocating for UX early.
The contrast determines your positioning. Anthropic's Claude.ai team sits closer to your demonstrated territory: turning model capability into trustworthy product behavior through design and engineering decisions. OpenAI's Model Designer is the harder conversation because it explicitly asks for researcher collaboration and model-behavior intuition. That's where the reframe needs to be sharpest.
The Grain of Truth
State it before they can.
You have not worked inside a frontier lab. You have not sat with researchers tuning model outputs. You have not collected training data to shift how a foundation model responds. OpenAI's language about "technical intuition about how data changes can affect model behavior" describes a feedback loop you have not been inside of at that altitude.
That is real. Don't soften it.
But don't let that gap colonize the entire territory of "shaping model behavior." The postings themselves don't define it that way. OpenAI's own language includes product sense, user feedback, taste, and interaction paradigms alongside researcher collaboration. The posting describes a role with two halves. You can credibly claim one. You cannot yet credibly claim the other. The question is whether the half you own is substantial enough to make you serious.
Your Published Agentic Labs
Three published labs at junochen.com. Live systems where design decisions govern agent behavior in production. Not mockups. They run.
Carrier IQ is your lead card at OpenAI. Multiple insurance carrier portals running in parallel. Your wrapper describes eight; the live app currently runs five. Each agent navigates a real carrier site, fills forms, extracts quotes. Every carrier gets a confidence score based on data completeness. The output is a side-by-side comparison. The human evaluates before acting.
That confidence score is a model-behavior design decision. You decided what the system reports about its own reliability. You decided that data completeness, not just price, determines how the output is ranked. You decided the human sees what the agent found and what it didn't find. A different designer making different decisions about the same agent capability would produce a different product with different trust characteristics and different error rates. Same underlying model. Different behavior in production, because of design.
Brand Pulse shows a different facet. An agent reads Reddit and X continuously, scores every mention by sentiment and urgency, produces an hourly brand-health score plus a weekly narrative written by the agent. The design decision that matters here: what you made visible. Evidence feeds. Platform coverage. Narrative synthesis the user can inspect against raw signals. You shaped how the agent's intelligence reaches the human, which determines whether the human calibrates trust correctly or overtrusts a confident-sounding summary.
Retail Velocity completes the set. Daily account audits, compliance-gap scoring, ranked opportunity surfaces. The field rep sees what changed, what's missing, what needs to move. You designed the prioritization logic's output layer. You decided which agent judgments the human sees first and which get buried.
Across all three: you architect the human-AI decision loop for systems where the agent acts and the human must decide whether to trust what the agent did.
The Design Layer Shapes System Behavior
Two external sources worth naming in conversation. Both do real work for you.
Microsoft Research found that feature-based explanations in AI systems did not improve human decision outcomes. They increased overreliance when the AI was wrong. The model's accuracy was constant across conditions. The system's real-world accuracy changed based on what the designer showed the human. The designer shaped the system's actual performance without touching the model. That finding validates your reframe directly: the design layer determines whether model capability translates into correct outcomes.
EU AI Act Article 14 requires high-risk AI systems to be designed so humans can understand capacities and limitations, monitor operation, detect anomalies, remain aware of automation bias, interpret outputs, and override or interrupt the system. Every one of those requirements is a design decision about what humans see, when they see it, and what controls they have.
Your published essay "Trust Is the New Interface" makes this argument explicitly. Five handoffs in every agentic workflow: Intent-Setting, In-Progress, Output Review, Decision Gate, Loop Feedback. A trust ladder: Watch → Verify → Delegate. You write that "accuracy is what the model achieves; consistency is what the design delivers."
That framework is inspectable. An interviewer can read it before your conversation and see that you think about model behavior as a system-level problem, not a screen-level one.
TinyFish as Currency, Not Evidence
Your current role gives you production-agent vocabulary that signals proximity to real problems right now. Agent traces. Auditability. Attribution. Governance. You can speak to what happens when agents run at production scale and the design system has to account for failure modes.
Use this in conversation. Do not point to it as portfolio evidence. Do not name product specifics. The value: you sound like someone working inside the problem today, not someone who built three labs and stopped. Currency, not proof.
Confidence Tiers
| Tier | Territory | Guidance |
|---|---|---|
| High — use freely | Interaction paradigms, product sense, user trust, oversight architecture, how model capability becomes product behavior | Your published work directly demonstrates this. Both postings name these as core concerns. |
| Moderate — use with precision | Collaborating with researchers to shape model behavior | You've designed systems where design decisions constrain and direct agent behavior. You cannot credibly say you've sat in the researcher-collaboration loop at a frontier lab. Frame as adjacent experience with transferable architecture. Not equivalent experience. |
| Low — acknowledge and pivot | Data collection strategy, training-data curation, technical intuition about how specific data changes affect model outputs | Genuinely outside your demonstrated experience. Say so. Turn the gap into a reason for wanting the role. |
Some frontier-lab work requires research collaboration and model-level intuition you haven't built yet. Claiming otherwise crosses from positioning into fabrication. The honest frame: you own the product-behavior half, you'd need to build the researcher-collaboration half, and you'd bring fifteen years of knowing what makes humans trust or distrust a system's output to that collaboration.
Response Phrasing: Frontier-Lab Calibrated
"How close have you been to the model itself?"
"I've built three live agentic systems where my design decisions determine what the agent reports about its own confidence, what evidence the human sees, and where the human gate sits before action. In Carrier IQ, agents run insurance portals in parallel and every carrier gets a confidence score I designed based on data completeness. That score shapes whether the broker trusts the output. The model's capability is one variable. What the human sees and when they intervene is the other. I own the second variable."
- At Anthropic, add: "That's the work your Claude.ai team describes — turning a capable model into a product people enjoy using. My version of that adds the trust layer: confidence scores, evidence trails, decision gates."
"Have you worked directly with researchers?"
"Not inside a frontier lab. I've worked with agent architectures in production where the design layer governs what the agent does with its outputs. I've published a framework for the five handoffs in every agentic workflow. That researcher collaboration is partly why I want this role. What I'd bring to it on day one is the product-sense layer: knowing what makes humans calibrate trust correctly versus overtrust a confident-looking result."
- Stronger at OpenAI, where researcher collaboration is explicit. At Anthropic, this question is less likely. Their Claude.ai posting doesn't ask for it.
"What's your experience with model behavior versus interface design?"
"I don't think that's a clean split. Microsoft Research showed that explanation design increases overreliance when the model is wrong. The model's accuracy didn't change across conditions. The system's real-world accuracy did, based on what the designer showed the human. The EU AI Act requires human oversight as a design property of the system. When I built Carrier IQ, the confidence score is model-behavior design. I decided what the system tells the human about its own reliability. A different design produces a different trust profile, different error rates, different outcomes. That's shaping behavior."
- At OpenAI, follow with the trust-essay framework.
- At Anthropic, follow with the full lab portfolio.
When the conversation genuinely exceeds your territory:
"I haven't curated training data or worked inside the RLHF loop. That's a gap I'd close by being in the room with researchers, which is part of why I want this specific role. What I wouldn't be learning is how to design the control surfaces that determine whether a model's capability actually works for the person using it. I've been doing that for fifteen years, and for the last year specifically with agentic systems in production."
- Wrapper vs. live app discrepancy: Your Carrier IQ wrapper brief says eight carrier portals but the live app currently runs five agents — reconcile the copy before an interviewer opens both and does the reconciliation for you.
- Gusto's published AI principles: Their June 2026 post names three active design questions — showing what an agent is doing, designing intelligent escalation, preserving agency in consequential automation — that map almost exactly to your trust-essay framework and could serve as a bridge conversation if frontier-lab timing stalls.
- Ramp's Director posting language: Ramp now describes design leadership as shipping PRs, designing memory, and rethinking process as new capabilities unlock, which validates the IC-proximate leadership model and may be a stronger near-term fit than either frontier lab.
- Backdoor references are expanding: A June 2026 Wall Street Journal report covered in Becker's says employers are increasingly running informal reference checks beyond the candidate's list, especially at senior levels — every claim in this dossier needs to survive that channel.

