A contract here means: what a published artifact at junochen.com can prove if an interviewer clicks through, and the exact line where its proof ends. Past that line, you need a prepared verbal bridge.
Use this the night before. Find the artifact. Read its contract. If the interview will push past the proof boundary, prepare the bridge.
The loop model has three segments: strategy (you named the problem), artifact + observed behavior (you built something and people encountered it), outcome signal + next decision (something measurable happened and you changed course). Every artifact covers part of this, but none covers all three.
This piece maps your evidence to the loop model specifically. For bounded claims, evidence-chain custody, and claim-to-proof pairing outside the loop framework, that's Portfolio Playbook territory.
Strategy
Proves: You have a first-person framework for the core design problem in AI products — when humans should watch, verify, or delegate to agents. Five named handoffs. A Watch-Verify-Delegate ladder. Specific breakage examples drawn from practice: a quote agent reading the wrong row, a provider update changing shipped behavior, a simplified first interaction collapsing first-task success. These come from building, not from reading. (Confidence: high.)
Proof boundary: The essay proves you can name the problem. It does not prove you built the solutions. The five handoffs are not traced forward into any Lab's shipped interaction design. An interviewer who reads the essay and then opens CarrierIQ will see conceptual alignment — review states, evidence runs, verification prompts — but the essay never says "here is how I implemented this in CarrierIQ." (Confidence: high.)
Contract: Use the essay to establish strategic altitude. Do not claim it as a design specification for the Labs unless you are ready to draw the connection verbally, handoff by handoff.
Artifact + Observed Behavior
The Labs
All three — Brand Pulse, Retail Velocity, CarrierIQ — are live, runnable applications you built solo. They are product surfaces, not narrative case studies. There is no separate case page documenting research, iteration, or outcomes for any of them.
Proves: The richest inspectable behavior of the three. Applicant intake, parallel agent execution across five carriers, per-carrier trace views, evidence runs, re-run verification, coverage deltas, review states (Bindable, Normalize, Referral, Call Review). These are implemented control and review affordances — the human-in-the-loop architecture the Trust essay argues for. (Confidence: high.)
Proof boundary: The traces show a single run. They do not constitute a versioned iteration history. The app does not document a wrong quote moving through detection, diagnosis, a design change, and a verified corrected next run. You have control architecture without a published control record. (Confidence: high.)
Proves: An observable work sequence — search, venue discovery, menu scan, item extraction, brand classification, completion with confidence label — followed by brand distribution, price analysis, and ranked displacement opportunities. The 247-venue figure is demo scope. (Confidence: high.)
Proof boundary: No evidence of a person acting on a ranked opportunity, no validation of the ranking, no documentation of classification failure or subsequent correction. The mechanism is documented but no consequence follows from it. I graded this the same way in "What Survives the Second Question" and the assessment holds. (Confidence: high.)
Proves: Parallel agents searching Reddit and X, classifying mentions by sentiment and theme, assembling a live brand score. Source-labeled mention feeds, keyword frequency, share-of-conversation comparisons, platform coverage. (Confidence: high.)
Proof boundary: Same structure as Retail Velocity — signal ingestion and decision-surface design, no measured human action or business consequence. One minor source tension: the app's intro names Reddit and X, while the showcase data includes Substack signals. If asked, the move is simple: "The showcase includes sample data from additional platforms to show extensibility — the core agents search Reddit and X." Say it once and move on. (Confidence: high for the proof; moderate on whether the source tension gets noticed.)
The three Labs prove you can build production AI systems solo, design staged agent behavior, implement evidence visibility and human-review mechanics, and ship working products. They do not prove team leadership, organizational influence, customer adoption, accuracy validation, or learning from failure.
Traditional Cases as Making Evidence
The five traditional cases — Alibaba, Thermo Fisher, Red Cross, Equinox+, Allē — are more visually interrogable at the component level than the Labs. They show design artifacts, system architecture, and interaction detail that the Labs' live surfaces do not expose. This is the making-altitude inversion I flagged earlier: your older traditional cases show more inspectable design depth than your newer AI work, because the Labs are runnable products while the cases are documented processes.
Contract: When a panel doubts your making depth — the "how involved were you?" question — the traditional cases are your primary evidence carriers, not the Labs. Each traditional case also carries an outcome contract below. (Confidence: high.)
Outcome Signal + Next Decision
Proves: Scale (30M+ enrolled members, 40K+ practices), consumer product design within a regulated ecosystem. (Confidence: high for the published claim.)
Proof boundary: 30M+ is your published claim, not an independently verified AbbVie metric. Claim the mechanism, attribute the metric to the program, in the same breath. The custody rule from Portfolio Playbook applies here. (Confidence: high.)
Proves: A genuine operating tension diagnosed and resolved through design — manufacturing data across internal systems, partners planning against invisible batch timelines, conflicting on-time delivery definitions. The finished architecture makes internal and partner measures visible side by side, moves constraints ahead of commitment, retains human QA approval at regulatory release. Outcomes: commitment from all six pharma partners, $20M+ annual margin recovery, 83% IRR, 42% overhead reduction. (Confidence: high.)
Proof boundary: The page presents the diagnosed problem and the finished control architecture. It does not name an opposing stakeholder, describe a speed-versus-risk dispute, show negotiation over the QA gate, or document the process that produced six-partner commitment. "Six of six committed" is adoption evidence, not governance evidence. (Confidence: high.)
Proves: Mandate creation — you identified the desktop-value mismatch, built the research case, secured the mandate, organized the response as three sprints. One explicit competing-view moment: a PM worried that removing the EIN registration gate would reduce sign-in; sign-up rose 4% after the change. (Confidence: high for the published record.)
Proof boundary: This is the closest you get to published organizational-conflict evidence, and it is thin. No executive sponsor named, no adjudication process described, no evidence of a decision you lost. The contract supports "mandate creation plus one disputed product bet." It does not support a broad governance-leadership claim. (Confidence: high for the bounded description; low for any broader claim.)
Red Cross publishes holds, fraud controls, supervisor gates, and outcome metrics. Equinox+ documents a late scope reduction and later restoration of deferred brand worlds. Both show outcome signals. Neither publishes a completed loop — a shipped regression, diagnosis, and corrected next run. Equinox+ comes closest with the scope recovery, but that is a delivery-recovery story, not a correction of a harmful shipped result. (Confidence: high.)
Current Loop Context
TinyFish
Proves (verbally): You are currently working in production AI — deploying agents, encountering the problems the Trust essay describes, operating at practitioner altitude. (Confidence: high for the positioning value; not applicable for inspectable proof, because there is none.)
Hard boundary: TinyFish has no navigable case at junochen.com. The standing rule from "Four Rooms Doubt Your Title" holds: TinyFish is currency, never collateral. Lead with published proof. Use TinyFish to establish that you are in the loop right now, not to carry claims your published work cannot support.
Transition Gaps
Five seams where no published artifact carries the proof. Your spoken answer does the work.
Gap 1: Strategy → Artifact. The Trust essay names the design problem. The Labs implement solutions to it. No published page connects the two. Prepare one walkthrough: pick one handoff from the essay and show where it appears in CarrierIQ, with enough specificity that it sounds like a design decision you made rather than a post-hoc reading. CarrierIQ is the right Lab for this — it has the richest inspectable surface to point to.
Gap 2: Artifact → Observed Behavior. The Labs show working systems. They do not show anyone using them. No published page documents a user encountering an agent's output, trusting or questioning it, acting or failing to act on a recommendation. This gap is structural — the Labs are solo builds without customer bases. When asked, do not pretend observed-behavior evidence exists. Bridge to TinyFish verbally: that is where you are encountering real users interacting with agent outputs in production. Keep the bridge short.
Gap 3: Artifact → Outcome Signal (Labs). The Labs prove you can build. They publish no business outcomes, no adoption metrics, no accuracy validation. The 247-venue and five-carrier figures are demo scope, not impact. This gap gets tested hardest in enterprise and healthcare rooms, where consequence literacy is the actual test. Bridge by pivoting to the traditional cases — Thermo Fisher's $20M margin recovery, Allē's 30M+ members — and frame the Labs as the current-practice complement to that outcome track record.
Gap 4: Outcome Signal → Next Decision (all cases). The hardest gap. No published case documents a complete failure-to-correction loop: a shipped result that went wrong, diagnosis of why, a design change, and a verified corrected outcome. Alibaba shows one disputed bet that worked. Thermo Fisher shows containment architecture. Red Cross shows fraud controls. The Trust essay's breakage examples are the closest published evidence that you have lived this loop — but the essay states the lessons without supplying incident dates, containment actions, or restoration metrics.
Prepare a verbal example from TinyFish or from an unpublished moment in a prior role. Make it specific: what broke, what you diagnosed, what you changed, what happened after. This is where vagueness costs the most. The question behind it — "what do you do when your design ships and fails?" — is how panels distinguish people who have operated from people who have planned.
Gap 5: Governance and Organizational Conflict. Alibaba's PM objection is the only published moment where someone disagreed with your recommendation and you can point to what happened. Thermo Fisher's six-partner commitment implies negotiation but does not show it. If a panel asks how you built consensus under disagreement, how you navigated competing incentives, or how you secured decision rights when someone else held them, your published portfolio gives you one thin example and a lot of finished architectures. Prepare a second example you can tell verbally — who objected, what they were protecting, what you changed in the decision process, and what the result was.
Your portfolio covers strategy and making well. It is thin on consequence loops and governance under conflict. The contracts above tell you where published evidence supports you and where your spoken answer has to do the work. Know which is which before you walk in.
- Recency as soft preference: Atlassian says it prefers portfolio cases from the prior three to four years, but its process probes product thinking, craft, and collaboration separately — suggesting recency is a preference within a broader evaluation, not a hard cutoff.
- Hands-on means three things: Current postings from Sigma, Gusto, and Babylist reveal at least three distinct expectations hiding inside "hands-on" — visible craft credibility, working artifacts that expose unknown behavior, and removing the designer-engineer handoff — and Sigma's posting explicitly asks designers to decide with engineering what becomes product versus validated experiment.
- Oversight without state visibility: NASA's current human-systems standards require automation to expose system state, projected state, override capability, and notification when a decision aid exceeds its competence — a framework worth studying because the accountability distinction between revocable and non-revocable delegation maps directly onto the Trust essay's Watch-Verify-Delegate ladder.
- Adobe's messy-truth advice: Adobe's SVP of Design tells leadership candidates to present only one or two projects deeply, acknowledge the messy truth and lessons learned, and explain design's voice inside the organization — guidance that directly supports preparing for Gap 5's governance-under-conflict question.

