Halfway through a portfolio review, asked kindly, framed as curiosity: most of these cases look like they're from a few years ago, what have you shipped recently?
Nobody in that room thinks you forgot how to design. Directors do not get rusty in five years, and everyone present knows it. The sentence is two questions welded together: can you perform at the level we need right now, and can you prove it?
Different evidence for each. Performance gets settled live, in how you take apart a problem you have never seen. Proof gets settled by artifacts, and artifacts carry dates. Candidates lose this exchange by treating it as one question. They either defend the old work, which concedes the frame, or grab for the newest thing they own, which is usually the thinnest thing they own.
Winnable in most rooms. Unwinnable in four. Working out which room you are standing in matters more than any line of phrasing below.
State the gap flat
Your metric-rich work is not recent.
Red Cross: 2021. Equinox+: 2020. Allē lands somewhere in 2020–2021 and your own site says both, which is a separate problem. The Thermo Fisher mySupply case describes a twelve-month build and attaches no year to it; the first public documentation of that launch is 2021. BCG Digital Ventures runs roughly 2014 to 2021. The engine room of your business-outcome evidence shut down about five years ago.
Alibaba is the newest enterprise case and its date is unsettled. The live case page says 2024. Other parts of your record read closer to 2023. Fix that this week. A candidate who cannot date their own work hands the panel the exact doubt the question was fishing for.
The newest published work, the Agentic Labs builds, has no team, no adoption data, no business outcome. TinyFish is current and cannot be shown.
That is the inventory. Do not soften any of it in a room. Do not soften it to yourself either.
Two assets per case. Three decay curves.
Every case you own holds at least two separable assets.
One is a measured outcome: a lift, a reduction, a deployment count. The other is a designed mechanism, meaning the structural decision you made about how a person works out whether to trust what a system is doing on their behalf, and what happens when they don't.
Metrics are indexed to a market condition. They stay true forever and stay persuasive about three years, because after that the buyer silently appends would that number still hold today and cannot answer. Your conversion lift on a pre-LLM cross-border procurement flow is a real number attached to a problem that has since changed shape. Nobody doubts it. Nobody moves.
Mechanisms are indexed to problem structure. If the structure is still standing, the mechanism still argues, whatever year sits on the label. The structures you worked inside are still standing, and over the last five years they moved closer to the center of the market. A five-year-old mechanism aimed at a problem the market now has at industrial scale beats a two-year-old case aimed at a problem nobody has.
Third curve, and it belongs to published thinking. Indexed to neither market condition nor problem structure but to demand: how many people are currently paying to solve the problem the writing names. That number moves in both directions. For the problem your trust essay names, it went up. Metrics decay, mechanisms hold, a well-aimed essay appreciates. I take that third class apart below, including the reason I am rating it lower than I first intended to.
I have run a version of the mechanism argument before, in the coherence dossier, where the throughline landed on the human-system operating contract: a person deciding whether to trust what a system is doing for them, an interface earning that trust in real time. What is new here is the sequencing rule that falls out of it.
Never let a decayed metric carry a claim a mechanism can carry instead.
In the weekly scans I treat posting age as a gate applied before scoring, never a discount applied after. Run your own evidence the same way. Classify by decay curve, then fix the order, then match the class to the question actually asked. Outcome questions get the number and an immediate re-index to the structure. Judgment questions get the mechanism and no number at all.
The audit, asset by asset
| Asset | Age | What has decayed | What has held | Lead-with confidence |
|---|---|---|---|---|
| Thermo Fisher / mySupply | ~5 yrs from documented launch | Any efficiency figure | Governance-visible workflow: ordering inside a regulated supply chain where the record of who approved what has to survive an audit | High |
| Alibaba | ~2 yrs if the 2024 label holds | Transaction and conversion lift | Trust in information about a counterparty the buyer cannot inspect: provenance, attribution, comparison under uncertainty | High once the date conflict is resolved |
| Red Cross | 5 yrs | Deployment scale | Non-expert trust calibration under time pressure; national rollout in six months | High |
| Equinox+ / Allē | 5–6 yrs | Consumer engagement metrics | Zero-to-MVP speed; category range | Use with caution as core evidence, fine as range |
| "Trust Is the New Interface" | 2 months | Nothing | Names, in operational detail, the problem senior postings are currently hiring for | Moderate, conditional on self-dating and one third-party citation |
| Agentic Labs builds | Current | n/a | Build fluency, agentic interaction judgment, zero-to-one definition | High as capability evidence, never as scale evidence |
| TinyFish | ~1 yr, unpublishable | n/a | Current-role credibility, spoken only | High as spoken context, zero as artifact |
Now the order, because classification without sequence is half a recommendation.
- Open with Alibaba, assuming the 2024 label survives your check. It is the newest dated enterprise case you can put on a screen, and leading with it dissolves the date frame before it forms.
- Follow with Thermo Fisher. Most durable single asset you own, and the one most likely to be discounted on sight for its age. Lead with the mechanism, give the year plainly when asked, in that order. Auditability in enterprise systems is not a 2021 concern that faded. It is the open problem at every company currently trying to let an agent take a consequential action on a customer's behalf.
- Hold Red Cross for mandates involving time pressure, non-expert users, or fast national rollout.
- Park Equinox+ and Allē unless someone asks about category range.
- Never open with the labs. They answer the currency question and lose the seniority one.
If the Alibaba year resolves to 2023, invert the first two.
The appreciating asset, and why I am marking it down
"Trust Is the New Interface" carries a June 2026 date and it is the only thing in your inventory worth more now than the day it published. The five-handoff structure and the watch–verify–delegate ladder land on language sitting in live senior postings. Gusto's Head of Design for its service platform asks the leader to determine "when AI outputs are ready to ship, when a human is needed." Xapien's Principal Product Designer posting names provenance, "calibrated confidence," and keeping people meaningfully involved in high-stakes decisions.
Correction, because my pitch version of this argument overreached. Those postings are contemporaneous with your essay, not downstream of it. The Xapien record published the same month; the Gusto leadership record updated the month after. Nothing in the record establishes that the market picked up your framing, and hinting otherwise invites a check you will lose in front of the one person most likely to run it.
The defensible claim is narrower and still enough: your published thinking names, in operational detail, the problem these employers are hiring senior designers to solve right now. Concurrence, not prescience.
Moderate confidence, two conditions. Say the date yourself before anyone reads it off the page. And get one third-party citation, talk, or panel appearance attached to it this quarter, because a self-published framework is a claim and an externally referenced one is a record.
Hold this next one privately and do not volunteer it: your freshest substantive artifact is a framework, and the gap has exactly that shape.
Be exact about the labs, and about where they stop
Your site currently shows four systems. The set I have been working from is three. Reconcile the count before a reviewer does it for you.
In the AI-native credibility dossier I drew the boundary: these builds demonstrate current hands-on capability and must never be allowed to imply production scale. That boundary still holds, and the research narrows it further.
What they establish. Current judgment about agentic interaction. Build fluency close enough to code that Xapien names functional prototypes, the kind that test a hypothesis before engineering commits, as an explicit want, and Stripe lists as a preferred qualification. Initiative. Zero-to-one problem definition. The ability to make an idea inspectable instead of described. Anthropic's own hiring guidance says it cares what you can do rather than where you learned it, and independent work is admissible evidence of doing.
What they do not establish and cannot be argued into establishing. Adoption. Retention. Revenue impact. Reliability under real failure conditions. Cross-functional leadership. Team development. On org-building mandates specifically, Adobe's published design-leadership guidance is blunt: individual-contributor glimpses matter "significantly less" than what the candidate's team produced.
Which forces a revision to the objection-sorting piece. I filed current AI-building evidence under perception gaps, the kind that live in the buyer's read rather than in the record. For leadership mandates that expect team-produced outcomes, that filing was wrong. The gap is structural. The recent evidence is solo by construction, and no framing converts solo into team.
A fifth lab adds nothing the fourth didn't. Stop building breadth.
TinyFish stays verbal
Roughly a year in seat, since around August 2025, with Mino's public beta landing that November. Three uses, all spoken: enterprise AI, shipping continuously, agent governance handled firsthand rather than theorized. Never portfolio proof. Never a slide. If asked why it isn't in the work, one sentence about current-employer material, then move. Do not perform the discretion. The absence reads correctly on its own.
One repair, entirely in your control
The live site no longer labels TinyFish a case study and no longer exposes the customer figures that were visible a month ago. Good. One defect left: the linked pricing-strategy essay serves its full text, current-employer commercial detail included, to anyone who loads the page. The password field in front of it is a curtain, not a lock. That is why I am not linking it.
There is no counter-narrative available for a defect you control. Fix the access gate. Reconcile the system count. Resolve the Alibaba year. The site is the first artifact any reviewer touches and the only piece of evidence you own that carries today's date.
The clause you can win outright
Everything above services the proof clause. Performance is the other one, and performance has no date on it. It gets settled live: how you reason about a problem you have not seen, how fast you find the real constraint inside somebody else's mess, whether you hold the argument when you are interrupted mid-diagnosis.
So steer toward a working session wherever the loop allows one, and ask for a live problem instead of a take-home. A working session turns a question about when you did your best work into a question about how you think, and the deficit vanishes the moment the format changes. In a portfolio-only loop you are arguing about dates inside a format built to display them.
Four rooms resist the reframe even after you have won the performance clause.
Four rooms where the reframe breaks
One. Mandates requiring recent production experience with AI failure. Gusto's leadership posting names no time window, and neither did any other senior posting I checked at Gusto, Xapien, Stripe, or Atlassian. What they name is production proximity: shipped AI-mediated experiences, uncertainty, graceful failure, human override in a live product. Your mechanism argument is genuinely strong there. Your proof of shipped AI at scale is genuinely absent. Concede the second, compete on the first, and know you are competing at a deficit.
Two. Panels that evaluate through business-metric verification. Some committees anchor on reference-checkable numbers. The seniority rubrics circulating among hiring managers ask for shipped work with measurable quantitative impact at scale and treat self-initiated projects as supplementary. My source there is one hiring manager's published breakdown, which is informal, but the same theme repeats in Stripe's and Xapien's stated requirements and in Adobe's leadership guidance, so I would treat it as real. Against that instrument, five-year-old numbers get discounted and nothing recent substitutes. The reframe asks for reasoning credit. This rubric pays only for verified results.
Three. A slate with someone on it who shipped an AI feature in the last eighteen months with adoption data attached. The hard one. Panels under time pressure take the number over the structure, and they are not irrational to, because the number costs them no inference. You cannot out-argue it. You can only out-scope it, by competing for mandates where the difficult problem is organizational or architectural rather than feature delivery.
Four. Org-building roles where recent evidence has to be team evidence. Covered above. Structural, not perceptual, and the honest move is to say so out loud.
Compounding case. The lab-native dossier already conceded you have not worked inside a frontier model lab. Where a mandate wants frontier-lab context and recent production AI at scale, two structural gaps stack. Misfit, not objection. Decline early and spend the cycles somewhere they convert.
One thing conspicuously missing from that list: not one posting I checked imposed a numerical freshness cutoff. This objection is mostly unstated interviewer instinct, not published criteria. So do not preempt it in outreach or in the opening minutes of a portfolio review. Preempt only where the posting itself encodes production proximity. Everywhere else, wait to be asked.
What actually changes the calculus
One artifact. Not five.
A publishable case that follows a single consequential failure the whole way through: the trace that surfaced it, the diagnosis, the design change, the evaluation threshold that defined what good enough meant, the release decision and whose name was on it, the monitor that stayed in place afterward. Named domain. Real users. A threshold a specific person had to own.
That case answers the first three break conditions at once, because it shows production judgment instead of asserting it. It comes from one of two places: TinyFish work that becomes publishable, or a design partner who trades publication rights for the build. Open the second conversation now, since the first runs on a timeline you do not control.
Two cheap upgrades alongside it. Attach any real adoption cohort to one existing lab, because a small user base beats zero by far more than the size difference suggests. And get the trust framework cited somewhere you do not own.
Pre-interview card
"Most of this work is from a few years ago." "The measured outcomes are, yes. Red Cross is 2021; the Thermo Fisher build launched in 2021. What held is the mechanism. Both were about a person deciding whether to trust a system's output before acting on it, and that is now the central problem in every product with an agent in it." → High.
"What have you shipped recently?" "About a year at TinyFish in enterprise agent infrastructure, shipping continuously. I can go deep on the governance and trust problems; I don't show current-employer work." → High. Then offer the labs as current build judgment, framed as prototypes, never as products.
"The agent systems, how many users?" "They aren't products. They're functional builds I use to test interaction hypotheses in code. They show current judgment about agent supervision surfaces. They don't show adoption, and I wouldn't present them as if they did." → High. The concession is what makes the rest land.
"Are those metrics still relevant?" "The number was true for that market. What transfers is the structure of the decision I designed around, not the lift." → Moderate. Survives only if you name the structure immediately and specifically. Left vague, it reads as evasion.
"Have you worked on production AI at scale?" Do not reframe. Answer straight, name TinyFish as the current context, pivot to the mechanism. → Use with caution. Break-condition question, and the honest answer is a partial no.
Two probes to run in your first ten minutes.
- What does your team currently do when the model is confidently wrong in front of a customer?
- Who decides whether an AI output is good enough to ship, and what do they look at?
Either answer tells you inside a minute whether this panel evaluates on structure or on numbers, which determines how you spend the remaining forty-nine.
- Read the Gusto mandate directly: The Head of Design, Unified Service Platform posting assigns the design leader authority over when an AI output ships, when a human must intervene, and when an output should not be sent at all, which is the clearest published statement of the mandate the mechanism argument is built for.
- Borrow Xapien's vocabulary: The Principal Product Designer posting names evidence, provenance, calibrated confidence, and graceful failure in the same breath, and it is worth mining for the exact words to use when you describe the Thermo Fisher and Alibaba mechanisms.
- The direct-evidence precedent: Anthropic's careers guidance says roughly half its technical staff arrived without prior machine-learning experience and invites independent research, writing, and open-source work as evidence, which is the strongest published support for treating solo builds as admissible at senior level.
- Make trust design falsifiable: The five-study, 731-participant finding from Vasconcelos and coauthors that engagement with an explanation depends on how costly verification is gives you a testable claim to bring into a panel instead of the assertion that transparency produces trust.

