Classification: Hybrid
Part perception gap, part evidence gap. The split matters because each requires a different fix.
Perception gap: Evaluators treat frontier-lab tenure as a proxy for AI design capability. What they actually screen for — designing for uncertainty, model behavior, trust boundaries, delegation — can be demonstrated without having sat inside Anthropic or OpenAI. When the objection is purely perceptual, positioning closes it.
Evidence gap: The artifact that would make your alternative path fully inspectable — a correction-loop record showing a failed output, your diagnosis, a system-level change, and verified improvement on comparable later cases — is specified but unshipped. Until it exists publicly, part of your reframe rests on verbal claims rather than proof an evaluator can click through.
Classified as hybrid in Issue #8. Nothing in the current evidence inventory changes that.
This dossier covers AI design judgment and frontier credibility. It does not cover code prototyping capability — that objection has its own piece. If someone raises both in the same conversation, separate them explicitly.
The fact you stipulate
You have not worked inside a frontier AI lab. You have not trained foundation models, run RLHF pipelines, or shipped features to millions of users on top of a model you helped build.
Say it plainly if the question comes up. Trying to soften it sounds defensive. Owning it buys you the credibility to contest what comes next.
What comes next is the inference: that you therefore cannot design for model behavior, uncertainty, delegation, or trust boundaries. That part is worth contesting. But only with evidence that actually contests it.
Construct 1 — Technical Depth
What this proves: You have designed and built control surfaces for agentic systems operating in consequential domains — running applications with real decision architecture, not wireframes.
Primary evidence: Carrier IQ, a running agentic insurance quoting system. What an evaluator sees when they click through:
- Structured intent capture before any agent runs — customer profile, driver history, coverage requirements, carrier selection
- Parallel staged execution — five named stages per carrier (Session, Navigate, Fill, Extract, Verify) with running, queued, and completed states visible in real time
- Agent traces — expandable ordered trace messages and extracted quotes per carrier
- Operational routing — four review states (Bindable, Normalize, Referral, Call review) that prevent every returned quote from being presented as equally actionable
- Coverage deltas highlighting gaps between what was requested and what was returned
- Verification and evidence attachment — active verification sessions, proof attachments, re-verification controls, live-view links
- Operator rationale — free-text note recording the reason for the decision
- Approval gate — "Approve Bind" disabled until the case is complete
This is a control surface for a running agentic workflow in a regulated domain.
What it does not show: Carrier IQ does not expose a reject, correct, or rollback control tied to a versioned system change. Its verification rerun refreshes carrier evidence, but the interface does not demonstrate that an operator's correction altered a prompt, extraction rule, or workflow version. Review and case disposition are visible. System learning is not.
Against employer language
I pulled current posting language from four companies to benchmark what evaluators actually screen for.
OpenAI asks for surfaces that make provenance, reliability, uncertainty, permissions, rollout state, and regressions actionable. Carrier IQ covers observation and human decision through traces, evidence attachment, and the approval gate. It does not cover instrumented learning, permission state, rollout or rollback objects, or regression comparison.
Gusto asks for designing around uncertainty, errors, graceful failure, human override, and deciding not to send an output. Carrier IQ's review states and coverage deltas prevent unreviewed outputs from reaching the decision point. But there is no explicit rejection state, and no evidence that a human override changes future system behavior.
Headway asks for agentic high-stakes workflows, behavioral patterns, experiments, and task-completion, error, and satisfaction measures. Carrier IQ is a high-consequence agentic workflow with parallel execution and operator review. No public user study, experiment record, or quantitative outcome data.
Render asks for one coherent agent loop spanning human supervision, recovery, permissions, shared objects, and agent-facing programmatic surfaces (CLI, API, MCP, SDKs, documentation). Carrier IQ covers the human-review half well. It does not cover the agent-facing half — no scoped access grants, no token management, no revocation flow, no agent-facing documentation, no explicit error-recovery sequence.
Confidence: Moderate. Carrier IQ is real, running, and demonstrates genuine design decisions about human oversight in an agentic system. It meets the human-review portions of what these companies describe. It falls short on system-learning and agent-facing portions. Lead with it as proof you have designed consequential control surfaces. Do not claim it as proof of correction loops or agent-facing infrastructure.
Recommended response: "I built Carrier IQ as a solo project — a running agentic insurance quoting system where agents execute in parallel and a human operator reviews evidence, verifies coverage, and controls the bind decision. The control surface is live — you can use it."
Construct 2 — Design Judgment
What this proves: You have a structured, experience-derived framework for where trust breaks in human-AI systems, and you can articulate design responses to each failure mode.
Primary evidence: "Trust Is the New Interface", published and publicly readable from your homepage. No login required.
- A five-handoff framework (Intent-Setting, In-Progress, Output Review, Decision Gate, Loop Feedback) mapping where trust transfers happen and where design must intervene
- The Watch-Verify-Delegate trust ladder, giving vocabulary for how user confidence should progress
- Three first-person failure examples: a confidently wrong initial agent output, a model-provider update that disrupted power-user workflows, and a simplified first interaction that reduced first-task success — each with the design response
These failures come from building agentic products at TinyFish. The framework is what you extracted from them. If pressed on where the examples happened, say so — TinyFish is the experiential origin, past tense. It grounds the essay in production reality. It is not portfolio proof, and you should not present it as such.
Against employer language
The Trust essay maps well to Gusto's emphasis on uncertainty, graceful failure, and human override. The five-handoff framework is a structured answer to "where does the human stay in the loop, and what does the interface owe them at each point?" It maps to Headway's interest in behavioral patterns and to the judgment layer of what OpenAI describes when they say "investigate, evaluate, decide."
Where it falls short: the essay describes design responses to failures but does not document a complete production correction record with incident dates, sample sizes, containment actions, versioned changes, and restoration metrics. The failures are specific enough to be credible, but they are told as narrative, not shown as artifact.
Confidence: Moderate to high. Published, readable, substantive. For companies screening on judgment and thinking — Gusto's "how do you decide not to send an output?" — it is strong. For companies screening on production proof with measurable outcomes, it is necessary but not sufficient.
Recommended response: "I wrote 'Trust Is the New Interface' after building agentic products and watching trust break in specific, repeatable ways. The five-handoff framework came from those failures. It's published on my site — I'd rather you read it than have me summarize it."
Construct 3 — Frontier Thinking
What this proves: You are working on the problems frontier labs are working on, from outside.
The Trust essay's five-handoff framework already addresses frontier-class problems — trust calibration across delegation boundaries, correction loops, human oversight that scales without collapsing into rubber-stamping. That essay is published and inspectable, and it is covered in Construct 2 as design judgment. What it does not provide is production-level proof that you have closed a correction loop end to end.
Evidence status: Two forward-looking artifacts have been specified. Neither is publicly shipped.
The first is a Carrier IQ Correction Lineage model — an annotated interaction showing a disputed field through three claims: correction receipt, incorporation into a versioned system object, and verified improvement on comparable later runs. This is the artifact that would close the evidence gap identified in Issue #8.
The second is an Inference-Aware Decision Surface — an interaction model where a user states cost, delay, and reliability constraints, marks each as hard or movable, and sees satisfiable, tradeoff, or unsatisfiable states.
Neither is on your site, neither can be linked, and an evaluator cannot inspect them.
Confidence: Use with caution. The Trust essay gives this construct a published foundation — you can point to frontier-class thinking that exists today. But the artifacts that would make that thinking inspectable as production proof are unshipped. You can reference them as work in progress. You cannot lead with them. A verbal description of an unshipped artifact is a claim about future evidence, and the objection you are facing is about whether your experience is production-tested.
Recommended response: "I'm extending Carrier IQ to show the correction loop — what happens when an operator's flag becomes a versioned system change and you verify improvement on later runs. In progress, not published yet." Say it once, in response to a direct question. Do not lead with it.
The real gap
The correction-loop artifact remains the highest-leverage missing proof. Until it ships, "technical depth" is strong on review architecture but incomplete on the learning loop, and "frontier thinking" rests on a published framework plus verbal descriptions of unshipped work. Building the artifact is the fix. Positioning cannot substitute for it.
Which question you are actually being asked
When an evaluator doubts your AI credibility because you haven't worked at a frontier lab, they are usually asking one of three things. Getting the match wrong costs you regardless of how strong the evidence is.
"Has she designed for model behavior that changes underneath her?" Point to the Trust essay's failure examples — the model-provider update that broke power-user workflows, the confidently wrong initial output. Specific enough to demonstrate you have lived with model instability.
"Can she build the control surfaces these systems need?" Point to Carrier IQ. Staged execution, review states, coverage deltas, verification, approval gate. Let them use it.
"Can she close the loop — change the system based on what review reveals?" Stipulate the gap. The correction-lineage artifact is specified and in progress. You can describe it. You cannot show it yet.
The first two questions have inspectable answers today. The third does not. Knowing which question is on the table determines whether you lead with strength or acknowledge a gap you are actively closing.
Quick Reference — Pre-Conversation Scan
The objection: She's never worked at a frontier AI lab, so she can't design for AI.
Classification: Hybrid. Perception gap on what AI design credibility requires; evidence gap on correction-loop proof.
Stipulate first: "I haven't worked inside a frontier lab. That's true. What I have done is build agentic systems in consequential domains and design the control surfaces that keep humans in the decision."
Technical depth (moderate confidence): "Carrier IQ is a running agentic insurance quoting system I built solo. Parallel agent execution, staged visibility, four review states, coverage deltas, verification with evidence attachment, operator rationale, and a gated approval. It's live — you can use it."
Design judgment (moderate-to-high confidence): "'Trust Is the New Interface' is published on my site. Five handoffs where trust transfers in human-AI systems, a trust ladder, and three specific failures I designed through. I'd rather you read it than hear me summarize."
Frontier thinking (use with caution): "I'm extending Carrier IQ to show the correction loop — what happens when an operator's flag becomes a versioned system change and you verify improvement on later runs. In progress, not published yet." Only in response to a direct question.
The gap to name if pressed: "The piece I haven't shipped yet is the correction lineage — the full record from failed output through diagnosis, system change, and verified improvement. That's what I'm building now."
Scope boundary: This dossier does not cover code prototyping. If both come up, separate them.
- Render's agent-facing requirements: Render's Staff Agent Experience posting asks for design across CLI, API, MCP, SDKs, Blueprints, Skills, documentation, and web UI as one coherent loop — a surface category Carrier IQ does not currently address.
- Silber's stated screening criteria: OpenAI's Head of Product Design told Lenny's Podcast that curiosity, prototyping, point of view, strategic thinking, and systems thinking are what stands out when the company evaluates design candidates, which maps closer to the Trust essay than to lab tenure.
- Anthropic's designer-as-implementer model: Anthropic reports that its product designers use Claude Code for direct implementation and state-management changes, with Figma and Claude Code open 80% of the time — a workflow claim that will shape the separate code-prototyping dossier.
- Gusto's small-team leadership geometry: Gusto's Head of Design role currently manages three designers while requiring hands-on prototyping and production AI judgment around uncertainty, errors, and human override, making it a closer mandate match than the title alone suggests.

