Two findings govern everything below. Read them before the tables.
Agentic Labs publishes no customer or business outcome, and no prepared story will conjure one. What Carrier IQ, Retail Velocity, and Brand Pulse do prove is that you can build working agent interfaces with verification states, evidence attachment, and confidence surfacing. The homepage figures are something else. Five carriers in 90 seconds, 247 venues, 76 live signals: that is demo scope and throughput. Not adoption. Not retention, revenue, or cost. Concede the boundary out loud instead of letting a throughput count occupy the seat where an outcome belongs. TinyFish gives you commercial-outcome vocabulary in conversation. Saying it aloud does not convert it into portfolio proof.
Each of your three impact types is strong somewhere, and never all three inside one case. Product and design judgment is strong everywhere. Organizational evidence peaks twice and on unlike things: decision rights at Equinox+, mandate creation at Alibaba. Outcome evidence peaks at Allē, Alibaba, and Thermo Fisher. Only Alibaba puts organizational evidence and real numbers on the same page. An evaluator hunting for the single case that does all three will keep hunting. Answer at portfolio level and the question closes.
Which archetype you walk into best armed
Ranked, because you are not going to prepare for four in one week. Per-case load ratings sit in the table at the end.
- Healthcare and regulated. Your strongest evidence sits exactly where these evaluators look first.
- Growth-stage platform. The numbers exist. One of the prettiest ones is not yours, and you need to be the person in the room who says so.
- AI-native. Product and design surplus. The missing outcome layer costs you less here than anywhere else.
- Enterprise platform. They weight organizational evidence first, and your organizational thinness is the kind preparation does not touch.
Four kinds of thin. Only one answers to a failure story
This distinction does more work than the grades do. When a cell reads thin, it is thin for one of four reasons, and the remedy is different in each case.
| Gap type | What it means | What recovery does | Where it appears |
|---|---|---|---|
| Narration gap | The event happened. The page does not tell it end to end. | Closes it fully. | Red Cross, Thermo Fisher, the Trust essay |
| Attribution gap | A number sits on the page, but your work did not move it. | Reframes rather than closes. | Equinox+ |
| Measurement gap | Nobody ever produced a number. | Cannot touch it. | Agentic Labs |
| Scale gap | You have not yet held a design function across years, with reporting lines and level structure that outlived a single program. A failure story does not convert into tenure. | Wrong instrument entirely. | Your enterprise organizational cell, and every other organizational thin here |
The recovery layer is narrower and more useful than it looks. It will not fill an empty cell. It keeps your full cells from collapsing under a follow-up question.
How to read the verdicts
Impact types.
- Strategic/organizational. Mandate, decision rights, team, executive influence.
- Product/design. The decisions and the trade-offs themselves.
- Metric/outcome. Movement attributable to your work.
Evidence grades.
- Surplus. More than the archetype needs.
- Adequate. Enough. Not a differentiator.
- Thin. Present, will not survive a probe.
- Absent.
Section grades are relative to that archetype's weighting. Table grades are absolute across the portfolio. One case can be adequate in one and surplus in the other.
Recovery verdicts.
- Survives on published detail. The page already holds what an evaluator needs.
- Survives with prepared narrative. The page holds an anchor; you supply the completion out loud.
- Does not survive. No anchor, and building one now is retrofitting.
Confidence labels apply only to my reads of what evaluators weight. Evidence grades are factual reads of what your pages say today.
One sourcing caveat, because you should be able to calibrate. No senior design posting I reviewed asks for failure evidence in a portfolio. Ambience is the only one that prints the word "postmortems," and it prints it as an on-the-job duty, not an interview requirement. What actually supports the recovery layer is the ordinary machinery of structured interviewing. The federal structured-interview guidance tells interviewers to run predetermined probes: who was involved, what led to the situation, what your specific role was, what resulted, what you would do differently. It also tells them not to challenge the candidate by word or expression. Amazon tells candidates outright to arrive with examples of failing.
So the pressure is decompositional. Nobody raises their voice. The question narrows until the claim has to name your hands, and vague evidence dies there.
1. Healthcare and regulated
Weighting. Product and design first, and specifically on consequence-bearing workflows: who is permitted to act, what stops them, what the system does when it is uncertain. Strategic and organizational second, in the narrow form of compliance ownership and sign-off authority. Metric and outcome third, read in safety and deployment terms rather than growth terms. Confidence: high. The clearest outside corroboration is Ambience's Staff Product Designer posting: embedded clinician research, high-stakes workflows, launch ownership, feedback loops, fixing broken conditions, design present in critiques and postmortems. That posting is built around consequence. Growth barely appears in it.
Evidence read.
- Product/design: surplus. Your best cell in any room. Red Cross publishes duplicate-client interception before case creation, alerts on a changed address or identification, an automatic hold on duplicate payment, mandatory review of anomalous cases, and eligibility approval held by supervisors rather than caseworkers. Thermo Fisher publishes exception routing on an SLA clock, deviation references tied to an audit trail and closure timeline, and a retained human regulatory gate for release. Carrier IQ runs the same logic in an AI context.
- Strategic/organizational: adequate. Product Design Director plus General Manager. Six legacy systems consolidated into one nationally deployed platform in six months. Scope evidence, even absent design-org scale.
- Metric/outcome: adequate. $847,000 disbursed across 1,689 cases in the first two weeks. Deployment reality at modest magnitude. Sufficient here. Thin anywhere else.
What they fear. Harm reaching a person, then nobody able to explain how it got there.
The push. "What was the worst thing that could have gone wrong in that workflow?" Then: "Did it?"
Recovery verdict: survives on published detail for the first question, survives with prepared narrative for the second. The preventive architecture is fully published. What no page does is narrate one completed incident from detection through containment, restoration, and a change to the shipped system. Your anchor is the Red Cross retrospective, where you name three things the original build did not close: no client portal, cascading call-center errors, and a recovery plan that stopped at disbursement. That third admission is the most valuable sentence in your published portfolio for this archetype. It is a designer noticing, unprompted, that she built the path out and not the path back. Pre-load it. In a regulated room, a gap you volunteer will beat a metric you produced.
2. Growth-stage platform
Weighting. Metric and outcome first. Product and design a close second, inspected at production altitude. Strategic and organizational third, and reframed as builder instinct: whether leadership and direct making are still joined, per the earlier builder-instinct read. Confidence: high on metric primacy. Amplitude's Head of Product Design posting asks for a leader who pushes on production details, papercuts, edge cases. Airwallex's Director posting asks for flows, states, permissions, object relationships, and continued engagement through launch and iteration.
Evidence read.
- Metric/outcome: surplus in two cases, contaminated in a third. Allē carries 3.2x redemption, 47% lapsed-member reactivation, $42 acquisition cost against a $92 benchmark, and in-app planners converting at 2.4x front-desk enrolment. Alibaba carries +7% daily active users, +20% daily transactions, +47% new visitors from search, +2.2 NPS, and a 47% drop in buyer-reported security concerns.
- Product/design: surplus.
- Strategic/organizational: adequate.
The contaminated third is Equinox+, and it repays a close look, because the same case is also your best organizational asset. It publishes four designers and a researcher, two of them hired mid-project when scope expanded to five brands. It publishes explicit ownership of information architecture, cross-brand navigation, the design system, token architecture, and named influence over roadmap sequencing and MVP scope. It publishes 600,000-plus members, a 4.8-star launch rating, 90 days to MVP. What it does not publish is any movement attributable to the redesign. The instructor-retention figure, where members who follow an instructor in their first 30 days retain at nearly twice the rate, is a pre-existing relationship you used as a design input. It is not a result you produced. Attribution gap. Do not let it sit in a growth-stage room doing the work of an outcome claim. One structured follow-up separates them. Do it before the interviewer does.
What they fear. A leader who attaches herself to numbers somebody else moved.
The push. "Which part of that was yours, and what did the number do afterward?"
Recovery verdict on Equinox+: survives on published detail as a constraint story, does not survive as an outcome story. Six weeks from launch, engineering capacity covered three of five brand worlds. You shipped SoulCycle, Equinox, and Pure Yoga on instructor-followership density and deferred Precision Run and HeadStrong to the first post-MVP update, where they landed. Detect, contain, restore, documented, told against a named constraint with a named restoration. Use it as that. Nothing more.
3. AI-native
Weighting. Product and design first, by a wide margin. Strategic and organizational second, and weighted as proof you can stand up a function that does not exist rather than run one that does. Metric and outcome discounted, because these companies' own products frequently have no stable outcome metric to point at. Confidence: moderate-high on product/design primacy, moderate on metric discounting. The direct evidence is thin. The one posting I can point to, Amplitude asking what AI-native work the candidate built and what she learned from it, belongs to a growth-stage company. I read it as a cross-archetype indicator of how AI capability is being screened generally, not as proof of how a frontier lab weights. The archetype grammar itself comes from the frontier-AI screening playbook.
Evidence read.
- Product/design: surplus. Carrier IQ's staged pipeline, coverage-delta comparison, human review state, evidence attachment, bind rationale, and re-verification add up to a complete argument about designing for machine output a person has to check before acting on it.
- Strategic/organizational: thin (scale gap). Agentic Labs is solo-built, by your own description. That proves range and removes leadership proof in the same stroke. Recovery is the wrong instrument; no failure story converts into org scale. The mitigant is structural. AI-native evaluators weight organizational evidence second, and many of them are standing the design function up from zero themselves, so the ceiling this gap imposes is lower here than at enterprise.
- Metric/outcome: absent (measurement gap). Unfixable. An evaluator still wants something in that seat, so put the mechanism there: Carrier IQ's verification, review, and re-run states as the argument that you design for consequence, with TinyFish as spoken commercial-outcome context. Never a throughput count wearing an adoption costume.
What they fear. Confident wrongness at scale. An interface that makes an unreliable output look verified, shipped to people who will act on it.
The push. "Describe a time your design made a model's output more trusted than it deserved to be."
Recovery verdict: survives with prepared narrative. The Trust essay publishes three real breakages, and the first answers that question head-on: your quote agent clicked the wrong control, read the wrong row, and returned a wrong result behind a confident green check. The essay names the lesson and the design response. It does not say how the error surfaced, how many results were affected, what you did that day, or what changed afterward. Narration gap, fully closable. Supply those four things out loud and this becomes the strongest single answer anywhere in your portfolio. The second breakage, where a model-provider update silently changed a shipped system and power users were the first to leave, some of them permanently, is your second-best, and more useful than it looks. It is a failure you did not cause and had to absorb anyway.
4. Enterprise platform
Weighting. Strategic and organizational first, with the emphasis on executive influence and cross-surface coherence rather than headcount, consistent with the enterprise playbook. Metric and outcome second, but only once translated into business language. Product and design treated as table stakes. Confidence: high on organizational primacy, moderate on the specific preference for influence over management scale. Atlassian's published design process is instructive: scenario questions on delivering results, cross-functional relationships, and ambiguity; a portfolio review probing choices, personal contribution, success metrics, and learnings; a senior interview run alongside Product and Engineering counterparts.
Evidence read.
- Metric/outcome: surplus. Thermo Fisher's $20 million-plus in annual margin recovery, 83% internal rate of return, commitment from all six pharma partners, and 42% overhead reduction is your most enterprise-legible cluster. And the framing carries: exceptions used to surface two to four days after commit dates, when rescheduling costs three to five times prevention. That is a line a CFO repeats in a meeting you will never sit in.
- Product/design: surplus.
- Strategic/organizational: thin (scale gap), and recovery cannot touch it.
Be precise about the shape of the gap. Alibaba shows you naming a mismatch, building a research case, and securing a mandate before leading work across three surfaces. Mandate creation is the harder half, and you have it on the page. What no case shows is sustained organizational leverage: a function held across years, with reporting lines, level structure, and craft-bar mechanisms that outlived one program. Thermo Fisher was one designer and two engineers. Red Cross was a seven-person cross-functional team. Equinox+ was five people and a scope expansion. Answer with mandate creation and the Equinox+ decision-rights statement, and do not let the interviewer's third follow-up be the thing that surfaces the ceiling for you.
What they fear. A design leader who cannot hold a position when an executive pushes back, and whose overridden decisions stay overridden.
The push. "Tell me about a decision you lost."
Recovery verdict: does not survive on published detail. What the record contains is a dispute you won. Product Management believed removing the EIN registration gate would suppress sign-in; sign-up rose 4%. That is a correct call vindicated, and offering it as a loss reads as evasion in a room trained to ask about personal contribution. A genuinely lost decision has to come from outside these pages, and it has to include what you did after losing it and what changed about how you fight the next one.
Two repairs that are not recovery problems
The Allē member count contradicts itself. The hero claims 30 million-plus members. The problem section says 18 million enrolled. Nothing on the page reconciles them. Enterprise and growth-stage evaluators read closely, and the moment one number on a page turns unreliable, the 47% reactivation figure stops being an asset and becomes a question. Write the reconciling sentence. This one needs a fact, not a story.
Retail Velocity and Brand Pulse expose mechanism without consequence. Scan confidence, sentiment classification, source-backed takeaways, error states: all real, none of it tied to a failed run that got corrected. Carrier IQ is the only one of the three that carries the mechanism argument through a probe, because verification, review, and re-run all live inside it. The other two demonstrate range.
Load ratings by case
| Case | Strategic/org | Product/design | Metric/outcome | Recovery load it can carry |
|---|---|---|---|---|
| Alibaba | Adequate. Mandate creation, no headcount or reporting line | Surplus | Surplus | Won dispute only (EIN gate). No shipped failure or rollback published. Does not survive a failure push. |
| Thermo Fisher | Thin (scale gap). One designer, two engineers | Surplus | Surplus, most enterprise-legible | Detection and containment survive on published detail. Full restoration and learning chain survives with prepared narrative (narration gap). |
| Red Cross | Adequate. GM scope, cross-functional team | Surplus | Adequate at modest magnitude | Controls survive on published detail. Four-part chain survives with prepared narrative, anchored to three admitted gaps. |
| Equinox+ | Surplus. Clearest decision-rights statement | Surplus | Absent as attributable movement (attribution gap) | Constraint and scope restoration survive on published detail. Design-failure push does not survive. |
| Allē | Thin (scale gap) | Surplus | Surplus, undermined by unreconciled member count | Lapsed-member reactivation is lifecycle work, not operational recovery. Does not survive as recovery. |
| Agentic Labs | Absent (scale gap). Solo build | Surplus for AI verification design | Absent (measurement gap), unfixable by narrative | Carrier IQ mechanism survives on published detail. Completed incident survives with prepared narrative only if tied to published evidence. |
| Trust essay | Absent by design | Surplus as a thinking asset | Absent | Three real breakages published with lessons. Survives with prepared narrative once containment and restoration detail is added. |
Every strategic/org thin in that table is a scale gap. Not one of them is a recovery problem, and a failure story prepared for any of them is preparation aimed at the wrong target.
Where the published record stops
Your preventive architecture is published and it is strong. Your completed recoveries are not published at all. Every archetype's characteristic push lands on the second half of that sentence, because structured interviewing exists to narrow from claim to personal action.
Three anchors need four prepared sentences each: how it surfaced, what you did first, what got restored, what changed afterward.
- The Red Cross disbursement gap
- Thermo Fisher exception routing
- Trust essay breakage one
That closes every narration gap in the portfolio. It does not close the measurement gap at Agentic Labs or the scale gap at enterprise, and it never will. Those two you name yourself, early, before somebody else finds them and gets to decide what they mean.
-
Explanations can backfire: A 2023 CSCW study found that feature-based explanations did not improve decision outcomes and increased overreliance when the AI was wrong, while example-based explanations performed better — useful if an AI-native panel asks why your trust work goes past "make it explainable."
-
Friction users dislike works: A 199-participant CHI experiment found cognitive-forcing interventions reduced overreliance more than standard explainable-AI approaches, but participants rated the interventions that helped them most least favorably — the exact trade-off your Decision Gate handoff has to defend.
-
Regulated rooms have a vocabulary: FDA's human factors guidance frames the work as minimizing use-related risk through perception, interpretation, action, and feedback about system behavior, which is closer to your Thermo Fisher exception architecture than any generic autonomy narrative.
-
Provenance is not truth: The C2PA specification records origin, modifications, and AI use in tamper-evident Content Credentials but explicitly states it does not establish whether content is truthful — worth knowing before any creative-AI conversation where attribution gets treated as verification.

