Juno, at this level the search is rarely lost on evidence quality. Everyone still in contention has shipped something real. It gets lost on assignment — which proof goes to which question — and the recurring error is overloading a strong asset: sending a case study to answer a question it was never in the room for, or asking an essay to prove you move fast.
You have five distinct forms of evidence. Each carries certain propositions well and certain ones not at all. Below is the boundary for each, then a mapping to the four tier archetypes: what leads, what supports, what breaks under weight.
On sourcing. The claims about evaluator behavior here come from pattern recognition across several hundred senior design searches over the past decade, not from anything I can link you to. Where the sample is thin and I am mostly inferring, I say so. Where I state something flat, I have watched it repeat often enough to bet on it.
An earlier version treated the inventory as four assets. This piece expands to five, updates inspectability as of September 20, and corrects two claims: the approved Agentic Labs set is three apps, not five, and Thermo Fisher's current page publishes "$20M margin opportunity," not the "$20M+ margin recovered annually" language used previously.
The inspectability line
Every form sits on one side of a line that determines how you can use it.
Asynchronous-ready. A reviewer inspects it without you present. You send a link. It survives the handoff from your champion to the committee member who was not on the call.
Live-only. You have to be in the room for it to land. It does not travel without you.
Current state:
| Evidence form | Inspectability | Access |
|---|---|---|
| Published case studies | Async-ready | Five cases linked from homepage |
| TinyFish testimony | Live-only | A TinyFish/Coca-Cola page is reachable, but the permission boundary classifies TinyFish product work as context, not approved design proof |
| Forward-looking artifacts | Live-only | No public URLs. All four domains return 404 |
| Trust essay | Async-ready | Linked from homepage |
| Agentic Labs apps | Partially async | Three approved apps (Brand Pulse, Retail Velocity, Carrier IQ) are live but not linked from the homepage, and their case-page routes are broken |
Read the table against the market. The forms most relevant to AI-native roles are the ones nobody can inspect without you. The forms anyone can inspect prove pre-AI track record. That asymmetry is the portfolio's central problem as of today.
Form 1: Published production case studies
Class: Inspectable exhibit. A reviewer works through these at their own pace, without your narration.
Proves: You shipped real products at real scale with measurable business outcomes.
- Alibaba — enterprise platform scale ($50B+ GMV, 25M+ desktop sessions) with published outcome movement: +20% daily transactions, +7% DAU, +2.2pt NPS, −47% buyer-reported security concerns. Your strongest business-outcome case.
- Thermo Fisher mySupply (BCG Digital Ventures) — zero-to-one delivery in a regulated environment: 12-month build across six pharma partners and nine sites, $20M margin opportunity, 100% partner adoption, 83% IRR. Your strongest complexity-and-speed case.
- American Red Cross (BCG Digital Ventures) — consequential constraints: six-month build, national deployment, six legacy systems consolidated, $847K disbursed, 1,689 cases processed in two weeks. Federal oversight, untrained volunteers, disaster-surge conditions.
- Allē (BCG Digital Ventures) — consumer-facing adoption economics at scale: 30M+ members, 40K+ provider practices, 3.2× redemption, 47% reactivation, CAC reduced from $92 industry average to $42. A flag: the current page uses "47%" for both reactivation rate and notification open rate in different sections. Do not present both figures together until the denominator is clarified.
- Equinox+ (BCG Digital Ventures) — launch speed and multi-brand platform architecture: zero to MVP in 90 days, 4.8-star App Store at launch, five brands on one platform. It does not publish a retention, revenue, or adoption-change metric.
Breaks: None of these involves AI-native product design. They prove shipping, leadership, and outcome movement in complex environments. They say nothing about inference-aware interaction, agentic system design, or any frontier AI problem. If an AI-native evaluator asks what you have built with AI and the answer is Alibaba, you have changed the subject, and they will notice.
Deploy: Lead with case studies when the evaluator's primary anxiety is can this person ship at our scale and complexity? That is the enterprise buyer and the regulated buyer. Do not lead with them when the anxiety is does this person understand AI-native product design? In that room, case studies without AI evidence beside them anchor the evaluator's read of you as a traditional designer exploring a pivot. I have watched that anchor hold through a full correction later in the same conversation. At panel level, first impressions are durable.
Form 2: TinyFish practitioner testimony
Class: Testimony. Verbal and resume credibility for recent AI production experience. Not published design proof.
Proves four things. Recent role context: you were building AI products, not sitting out the shift. Technical currency: you speak about model behavior, agent orchestration, and trust calibration from direct experience rather than from reading. Grounding for the forward-looking artifacts: they come from someone who has built this. And the bridge narrative — how you moved from enterprise product design to AI-native work.
Breaks: Testimony cannot fill a portfolio slot. When a design director opens your materials looking for screens, flows, and interaction decisions, "I did this at TinyFish but cannot show it" leaves the slot empty. Testimony works in conversation. It is inert in the async handoff, when your champion forwards materials to the committee and you are not there to narrate.
The TinyFish/Coca-Cola page creates a specific tension. A reviewer can find it and read it, but the standing permission boundary has not been updated to authorize TinyFish as design proof. Treat it as context that may be discovered. Do not deploy it as evidence.
Deploy: Use TinyFish in every live conversation where AI currency matters, in past tense. Do not rely on it to survive the handoff. Pair it with a forward-looking artifact whenever you can, so the testimony has something visible to attach to.
Form 3: Forward-looking visual design artifacts
Class: Demonstrative evidence. Original design deliverables you made to show how you solve frontier AI problems across four domains: Inference-Aware UX, Intent-Based Interaction, Agent Infrastructure as UX, Human-Agent System Design.
Proves: That you can work through AI-specific interaction problems at a level of visual and conceptual specificity most candidates at this level cannot match. Grounded by TinyFish and framed by the Trust essay, these are your strongest proof for AI-native evaluators.
The four domains are not interchangeable. I have not inspected the artifacts themselves, since they are not publicly reachable, so the read that follows comes from the domain framing rather than the work. Inference-Aware UX and Intent-Based Interaction map most directly to consumer-facing AI products where model uncertainty and conversational interaction are the core UX problems; those are your AI-native leads. Agent Infrastructure as UX reads as platform-layer work, strongest with companies building agent tooling or developer-facing AI products. Human-Agent System Design, with its emphasis on the handoff relationship, fits regulated or consequential contexts where trust calibration has operational stakes. In a live presentation, lead with the domain that matches the evaluator's product context, not the one you like best.
Breaks in two places.
First, inspectability. As of today all four domains return 404. They are live-only. You cannot send a link, you cannot leave them behind after a call, and if your champion wants to put them in front of the hiring committee there is nothing to put. This is the highest-priority gap in the portfolio.
Second, outcome. These artifacts show how you think and what you would build. They do not show that you built it, shipped it, and measured it. Evaluators carrying the production-proof burden — usually the hiring manager and the design-leadership peer — will not accept demonstrative work as a substitute, however strong it is. The evaluator who values these most is the one screening for AI fluency, not the one screening for shipping history.
Deploy: Present these live at any conversation stage with an AI-native or growth-stage company. Get at least one domain published at a dedicated URL so it survives the handoff. Until then, every live presentation is the only look this evidence gets.
Form 4: The Trust essay
Class: Expert reasoning. A published conceptual framework: five handoffs, Watch–Verify–Delegate. It establishes how you think about trust in agentic systems.
Proves: Depth of judgment about the central design problem in AI products right now — how humans calibrate trust in systems that act on their behalf. It is async-ready and publicly linked. For an evaluator who reads it closely, it signals a structured, original perspective on the problem their company is trying to solve.
Breaks: The essay proves judgment without proving execution. I wrote that in Issue 2 and have not found a better way to say it. An evaluator who reads it and asks where you applied this needs TinyFish testimony (live), a forward-looking artifact (live), or an Agentic Labs app (partially async). Alone, with no execution evidence nearby, it reads as a position paper. Position papers do not get people hired at Director+.
The essay also cannot speak to speed, team leadership, or business-outcome delivery. For an evaluator whose primary anxiety is can this person build and lead a team that ships? it is at best irrelevant. At worst it fixes their read of you as a thinker rather than a builder, and that label is hard to peel off once it has been said out loud in a debrief. I have seen it stick to other candidates who led with frameworks, through later evidence that contradicted it.
Deploy: Pair the essay with execution evidence. For AI-native roles, send the link and follow with a live walkthrough of forward-looking artifacts. For enterprise or regulated roles it is a supporting signal of intellectual depth, not an opener. Do not send it first to an evaluator whose concern is operational.
Form 5: Agentic Labs apps
Class: Working artifact with bounded impact proof. Live applications you built solo that a reviewer can open and use.
Proves: You can build and ship functional AI-powered products end to end. Brand Pulse, Retail Velocity, and Carrier IQ work. A reviewer can use them and form an independent judgment about the interaction design, the AI integration, and the product thinking. Most candidates at this level cannot offer that.
Breaks: They publish no adoption, retention, or revenue figures. They are solo builds, so they cannot speak to team leadership or organizational influence. And none is linked from your homepage; a reviewer needs the direct URL, and the case pages that once gave them context return 404.
One constraint to keep clean: Retail Velocity identifies its infrastructure as powered by TinyFish services. It is approved Agentic Labs evidence with a disclosed TinyFish dependency.
Deploy: Use the Labs to bridge the Trust essay's reasoning and the forward-looking artifacts' design specificity. They prove you build. Lead with them only when the evaluator's question is can she actually make things? When the question is can she lead a design org? they are a supporting signal that you stay close to the work, nothing more. A solo-built app offered as lead evidence to someone hiring for organizational scale invites the question of whether you prefer building alone to building through teams.
Deployment by tier
"What They're Afraid Of Predicts What They Screen For" mapped evaluator fears to evidence forms from the demand side. This is the supply-side complement.
AI-native companies. Lead with forward-looking artifacts (live) and the Trust essay (link, before or after the call). Support with TinyFish in conversation and Agentic Labs as proof you build. Hold the production case studies until asked. Do not lead with case studies here. An AI-native evaluator who sees Alibaba and Thermo Fisher before any AI evidence will classify you as a traditional product designer exploring a pivot, and exploring is not a competitive position at this level. The operational constraint compounds it: your lead evidence is live-only. Until the artifacts have public URLs, every AI-native opportunity depends on reaching a live conversation before the committee shortlists off async materials. Narrow window.
Growth-stage platform companies. Lead with Thermo Fisher (zero-to-one delivery, multi-partner complexity, $20M margin opportunity) and Alibaba (platform scale, measurable outcomes). Support with TinyFish for currency and the Trust essay for depth. If the company is building AI features — and most growth-stage platforms are — the forward-looking artifacts and Agentic Labs become useful supporting evidence that you can design and build for AI rather than only talk about it. Do not lead with the artifacts. A growth-stage evaluator who sees demonstrative AI work before shipping proof will question whether you can deliver under their timeline and stakeholder pressure. They buy execution confidence first.
Enterprise platform companies. Lead with Alibaba. It speaks their language: $50B GMV, +20% daily transactions, +2.2pt NPS. Support with Thermo Fisher for partner-adoption proof and Allē for consumer-side adoption economics. The Trust essay and forward-looking artifacts are late-stage differentiators, not openers. An enterprise evaluator who meets the AI work before the scale proof will question whether you are too specialized for their breadth; in my experience that reaction is reliable. The worry is not that you lack AI knowledge. It is that AI knowledge presented first implies a narrower scope than the role covers.
Healthcare and regulated companies. Lead with Thermo Fisher (pharma partners, regulated environment, 100% partner adoption) and Red Cross (federal oversight, disaster-surge deployment, six legacy systems consolidated). Those two prove you can build where mistakes have consequences and compliance is not optional. Support with Alibaba for scale credibility. Do not lead with AI evidence unless the role itself involves AI integration under regulation. If it does, the forward-looking artifacts — particularly Human-Agent System Design — separate you from candidates who have the regulated experience and none of the AI fluency. If it does not, AI-native evidence reads as misalignment with the role's actual constraints. Moderate confidence on that last distinction; the sample of regulated-context AI roles is still small enough that the pattern is emerging rather than established.
Where the portfolio is exposed
Each form is strong inside its class. The exposure is the distance between where demand is concentrated right now — AI-native and AI-adjacent roles — and where your inspectable proof is concentrated: pre-AI production work.
The forward-looking artifacts close that distance. They take the production record, the TinyFish experience, and the Trust framework and turn them into something an AI-native evaluator can look at and react to. Right now they exist only while you are in the room presenting them. Every week that holds, you are asking a champion to sell a committee on work the committee cannot see.
Publish one domain at a stable URL. Pick the one that matches the tier you are actively working. It is the highest-return week of work available to you.
- AI-native roles diverge sharply: OpenAI's current postings split across growth/adoption, identity/authorization, and core-product invention, while Anthropic's Evals & Prompts role requires production Python and evaluation pipelines — a qualification boundary, not an evidence gap.
- Player-coach scope varies widely: Ramp's Director posting says the hire begins with one product area before potentially expanding, but does not state the trigger, timing, or decision owner for that expansion.
- Structured interview architecture as signal: Atlassian's published design interview handbook separates Product Thinking from Craft Excellence and adds a Product–Engineering–Design triad at senior levels, making the loop itself a partial map of how the company distributes evaluative authority.
- Regulated AI requirements are hardening: Maven Clinic's Senior Staff Care Delivery role explicitly asks candidates to define how AI communicates uncertainty and when it defers to human care, moving regulated-context AI from a general preference toward a specific design accountability.

