Surface-form question: "You haven't worked inside a frontier lab. Is your AI work deep enough for this?"
Say the gap in one sentence. Claim the decision layer. Never claim the model layer. Full phrasing in "What the published work actually carries," below.
It sounds like credentialism. It almost never is.
There is a real fear in the room and it deserves better than a defensive answer. They are weeks from shipping something where a model returns an output that looks finished and confident and correct, a person acts on it, and the output is wrong. Nothing in the interface catches that. What they are afraid of buying is a designer with fifteen years of making things clear, in a problem space where clarity is the wrong reflex, because a clean presentation of a wrong answer travels farther than a messy one.
Attention heads do not come up in these conversations. What they want is evidence that you have designed the moment a person decides to trust a machine.
That question has an answer. You do not get to it until you hand over the piece you cannot claim.
Stipulate it, then stop
Juno has not worked inside a frontier model organization. Not OpenAI, not Anthropic, not Google DeepMind. Agentic Labs is not the same thing as building Claude-scale interaction systems from inside Anthropic, and nothing in the portfolio should be dressed to suggest otherwise. Issue #1 settled this. Nothing since has moved it.
One sentence. No softening verb, no pivot riding in on the same breath. Each extra clause you spend on the gap turns a stipulation into a confession, and rooms hear that difference instantly.
The stipulation is safe because it is narrow. It rules out one category of role. It leaves untouched the category most of these companies are actually staffing, and I do not have to argue that part. It is written into their postings.
Two design jobs inside one company
Eleven design and design-engineering postings, seven companies, all read in the last week. Start with OpenAI, and do not read OpenAI as one employer. Read the openings against each other.
Model Designer tells you its material outright. The model is the product. You are asked to understand and predict model behavior, devise data-collection strategies, and build technical intuition for how training data changes what comes out the far end. A research-adjacent job with a research-adjacent admission ticket.
Now the rest of the same design surface.
| Posting | What it asks for |
|---|---|
| Product Designer, People Innovation Labs | Four-plus years shipping large-scale software, interaction craft, research, design systems, and enthusiasm for defining how AI-first products should look and behave |
| Product Designer, ChatGPT | End-to-end product design, concepts through high-fidelity prototypes, collaboration with research |
| Product Designer, Growth | Experiences that help people understand, trust, and return to ChatGPT, and translate complex AI capabilities into interfaces people can read |
Read the absence. No data-collection strategy, no intuition about how training data shifts behavior. Every term that defines Model Designer is missing from the other three, and that gap is too clean to be sloppy copywriting. Someone drew a line and then staffed both sides of it.
Anthropic's most recent design-titled, non-research opening sits on the same side of that line as the three above. The Product Designer role for Claude Code, since removed and preserved on a job-board mirror, wanted a designer who stays close to the models, invents patterns native to AI, builds interfaces that build trust, and gives developers tools to confidently build, monitor, and control agentic systems. Proximity to the model, not ownership of it. No training data, no benchmark construction, no evaluation authorship.
Go to the applied end and the line gets drawn in blunter language than either lab uses. Gusto's Head of Design for its service platform and its Senior Product Design Manager for Payroll scope the work around four decisions: when AI output ships on its own, when a human has to intervene, when an output must not be sent at all, and how a user keeps agency once the automation gets it wrong. A payroll company, not a research organization, carrying the most explicit decision-layer language in the sample.
The absence check holds on Gusto too, in both directions. Gusto never asks for model proximity, research collaboration, or anything resembling frontier-lab tenure. And human override, escalation authority, release blocking, graceful failure are all named in the Gusto postings and appear in none of the three OpenAI product-design ones. Director and above: treat Gusto's four decisions as the bar you will be tested against in any posting written like that one. Eleven postings cannot tell you what the market wants. They can tell you that the company furthest from the model wrote override authority into a leadership job description and the labs did not.
Name the first job the model layer: model behavior is the design material. Name the second the decision layer: the human act of trusting, checking, overriding, or delegating is the design material. Issue #5 called the same split a reversal in the direction of influence, Juno designing the human layer around AI systems while a Model Designer designs the AI layer humans interact with. The postings take it one step further. The split lives inside one company, on one design team, in a single recruiting cycle.
What the published work actually carries
The defensible claim is narrower than "I do AI design." It is this: Juno owns the layer where a person decides what to do with what a machine produced. Four published pieces carry it, all inspectable at junochen.com.
Three built agentic systems. Brand Pulse turns live social signal into a decision surface. Retail Velocity audits field intelligence and ranks the gaps. Carrier IQ automates freight carrier workflows and hands a broker confidence-scored comparisons to act on. In none of them is the output the interesting problem. The interesting problem is what the person does next, including how they recognize the moment to do nothing.
A structural account of where trust fails. "Trust Is the New Interface" names five handoffs between a person and an agentic system: Intent-Setting, In-Progress, Output Review, Decision Gate, Loop Feedback. Under it sits a Watch–Verify–Delegate ladder treating autonomy as earned in stages rather than switched on. Lead with this. It outperforms any single case study in a room because it puts a taxonomy of failure on the table instead of a preference for calm interfaces, and everything in the interview turns on that distinction.
Two enterprise cases landing in the same place. Alibaba closes on an agentic sourcing loop. Thermo Fisher closes on continuous agents held behind a human QA release gate. Both codas date the instinct earlier than the current hype cycle: automation runs, and a person still holds the gate.
TinyFish is conversational context and nothing else: agent traces, auditability, reversibility, governance in production. Drop it as a present-tense aside and keep moving. "Reversibility is the thing I'm arguing about in production right now" buys you technical currency in eight words. If they pull the thread, talk about the problem class, not the product. It is not portfolio proof. Never hand it over as a case study.
Recommended phrasing:
"I haven't worked inside a frontier lab, so I won't claim the model layer. What I've built and published is the decision layer: the handoffs where a person accepts, checks, or overrides what the system produced. That's where trust is won or lost, and it's the layer the roles I'm looking at are scoped to."
Then stop talking. If they want the model layer they will tell you, and you will know it in ninety seconds instead of round four.
The next two questions
This objection rarely dies on the first exchange. Two follow-ups are predictable enough to rehearse. The first tests whether the work transfers off her own projects. The second tests how deep the evidence actually runs.
Probe 1: "Those are your own projects. What happens when you're designing against a model you don't own and can't change?"
Fair test, and the honest answer is that not owning the model describes almost everything she has shipped.
Recommended phrasing:
"That's most of my work. In the Thermo Fisher and Alibaba systems the underlying engine wasn't mine to change, so the design question became what the human is allowed to do when the output is wrong. The constraints move into the interface: what shows as provisional, what requires a second look, what can't proceed without a person signing off. If I could retrain the model, half those decisions would belong to someone else."
Turn the perceived limitation into the job description. Do not reach for TinyFish here, whatever the temptation.
Probe 2: "Show me a case where one of your trust patterns failed in production and you fixed it."
Thinnest part of the public record, and the place where a defensive answer costs real ground.
No published artifact runs a single trust handoff end to end: the failure mode, evidence from real usage, the design change, the criterion that catches it earlier next time, the production monitor. The Agentic Labs systems publish the mechanism and stop before the aftermath. They carry no adoption or customer-impact data on par with the enterprise cases. Whoever probes here has found the seam. Inflating a smaller example into the thing they asked for does more damage than the gap.
Recommended phrasing:
"Not in my published work. The lab systems show the mechanism, not a documented failure-and-correction cycle, and I'm not going to inflate a smaller example into one. What I can do is walk you through how I'd instrument one: which handoff I'd expect to break first, what signal would tell me, what I'd change. Or if you have a surface where trust is breaking right now, I'd rather look at that."
Acknowledge, then move to a live problem. A confident senior operator says "not in my published work" without flinching, and the room is measuring the willingness to say it. Spin this one and you fail it twice.
Where this stops being true
Two conditions kill the reframe. Both are checkable in the posting before you apply.
1. Model as material. The role asks you to predict or shape model behavior, devise data-collection strategy, or reason about how training data moves outputs. Model Designer is the archetype: the model is the product. Misfit, not bridge. Name it and step out, or apply to a different opening at the same company.
2. Evaluation authorship in code. Spotify's Senior Conversation Designer is the clean case: build evaluation frameworks measuring conversation quality, accuracy, and task completion; be comfortable writing and running evaluations; define guardrails for safe model behavior. Not infrastructure. Still a genuine technical admission ticket, because writing and running tests means writing code. Expect this one to arrive as a direct question in the interview.
Recommended phrasing:
"I've authored the criteria, not the code that runs them. I can tell you what counts as a failure at a given handoff, what should block a release, and what a monitor has to catch. Writing and running the evaluation suite itself isn't in my record. What's the actual split on this team?"
Decide on their answer. Evaluation authorship under required qualifications, single designer, no engineering partner: decline. Under preferred, or with someone to pair with: stretch, and stretches are fine. The line is whether the bar depends on you personally.
One correction, because I sent readers hunting for the wrong string. I have been calling the disqualifier evaluation-infrastructure ownership, meaning the systems that schedule, scale, and continuously monitor automated model tests for accuracy drift. Screen postings for that language and you get nothing back. That is the problem. Across all eleven postings, at Anthropic, OpenAI, Gusto, Spotify, Abridge, WRITER, and one early-stage AI tooling company, not one hands a designer a test harness, a regression gate, or distributed evaluation infrastructure. The work sits in named research and engineering roles. Anthropic's Evals Infrastructure Tech Lead holds it. OpenAI's Backend Software Engineer for Evals holds it. So the filter I gave you passes every posting, including the Spotify-shaped ones that would break you.
Source caution. Anthropic had no design-titled opening on August 1; the Claude Code language above comes from a removed posting preserved on a job-board mirror. OpenAI's ChatGPT and Growth pages were retrievable in late July and gone from the August 1 index. Treat all three as recent rather than reliably open. Eleven postings document specific role boundaries. They do not prove what the market wants. Use them to sort the posting in front of you. Do not carry them into a room as a claim about the industry.
Two minutes before you walk in
This sorts the posting, not the answer. Run it before you choose which version of the story the room needs.
- Find the noun the role acts on. Model, data, behavior, benchmark points one way. Interface, workflow, experience, product points the other.
- Look for the verbs. Monitor, control, intervene, override, review, approve. If they are there, the decision layer is scoped into the job.
- Check who owns evaluation. A design posting naming harnesses or regression monitoring is a research role that got titled design. A design posting asking you to run evaluations yourself is stating a hard requirement, not a preference.
- Check what is absent. An AI product posting with no mention of trust, error, or human oversight usually means nobody at that company owns those problems yet. Mandate or trap. One question in the first interview tells you which.
-
Anthropic's own admission standard: Roughly half of Anthropic's technical staff arrived without prior machine-learning experience, and its careers guidance tells candidates to foreground direct evidence such as independent research, writing, or open-source work rather than pedigree — which is the strongest available counterweight to the assumption that lab tenure is the ticket.
-
Explanation can make reliance worse: Chen, Liao, Vaughan, and Bansal found that feature-based explanations increased overreliance while example-based ones helped people override wrong predictions, which is the empirical spine under any claim that a trust framework is operational rather than decorative.
-
Verification cost, not transparency: Vasconcelos and coauthors ran five studies with 731 participants and argued that people engage with an explanation based on how costly it is to check the AI — useful vocabulary for turning the five-handoffs framework into questions an interviewer can test.
-
Where the override language is codified: NIST's AI Risk Management Framework calls for defined roles and responsibilities in human-AI oversight configurations, which gives the decision-layer argument a standards vocabulary that enterprise buyers already recognize.

