- FILE TYPE: Machine-readable context layer. Paste it whole into a Claude or ChatGPT project. Positioning sheets, cover letters, and interview prep generate from this file and from nothing else.
- SUPERSEDES: Design Context File v1, 2026-07-24.
- GENERATION DATE: 2026-08-01. Site state pulled the same day from the origin response, not from a rendered view.
- GOVERNING RULE: Every claim carries a confidence label. No company-specific assertion that is not written here. Return UNKNOWN and stop.
- CHANGES_THIS_CYCLE:
- The cases now hang off one continuous record of judgment, tagged by altitude.
- The published trust ladder is retired and replaced by two axes.
- Interview prep is rebuilt around the two frames evaluators are actually running this summer.
Read First
USE_WHEN: first load, every cycle. These two outrank everything below them.
- One blocking evidence gap. Carrier IQ shows an operator reviewing, contesting, and approving. It does not show the contest landing anywhere. Nothing public demonstrates that the next run comes back different. That closing link is the differentiation, and today it lives only in your spoken account. The full argument and the reconstruction task are under Longitudinal Control Record.
- One unfixed exposure. All five full case pages and the private pricing essay still return complete content to an unauthenticated request. No change since July 24. It does not slow your search. It does constrain what this file is allowed to generate. Behavioral consequences are in Prohibited.
Identity
USE_WHEN: opening any generation task. Supplies the frame every other section assumes.
POSITIONING_LINE: Design leader for systems where a human decides
whether to trust a machine.
MANDATE_SOUGHT: Director+ product or design leadership. Staff and
Senior Staff IC qualify only where the seat owns
reusable infrastructure: interaction models, eval
loops, launch gates, cross-team operating patterns.
GEOGRAPHY: Remote-first or US-compatible. Relocation-only roles
are excluded, not deprioritized. HARD GATE.
DOMAIN_SPINE: Pharma supply chain, disaster response, B2B
marketplace, aesthetics practice networks.
ADJACENT_BUILD: Agentic insurance orchestration (Carrier IQ), built
outside the spine on purpose.
CURRENT_CONTEXT: Agentic platforms at TinyFish. Speakable in
conversation, never citable in an artifact.
SORTING_VARIABLE: Decision-rights density. Rank every opportunity by
how much of the product turns on a human trusting,
overriding, delegating to, or reversing machine
action. AI-as-feature ranks below AI output that a
human must judge before a consequential act.Artifact Status
USE_WHEN: deciding whether a surface can be linked, cited, or only spoken.
HOMEPAGE_TITLE_POSTURE: LEADERSHIP_FORWARD
HOMEPAGE_HUMAN_CARDS: Alibaba | Thermo Fisher | Red Cross |
Equinox+ | Allee
HOMEPAGE_AGENTIC_LABS_CARDS: Carrier IQ | UAT Sentinel |
Retail Velocity | Brand Pulse
AGENTIC_EXTERNAL_USE_SET: Carrier IQ | Retail Velocity | Brand Pulse
TINYFISH_CASE_CARD: NOT_PRESENT
TINYFISH_CONTEXT: HERO_ASIDE_ONLY
FULL_CASE_SERVER_GATE: ABSENT_ALL_5 (see Prohibited)
PRIVATE_ESSAY_SERVER_GATE: ABSENT (see Prohibited)
EQUINOX_TEASER_METADATA: GENERIC_TITLE_REMAINS- CITABLE: Homepage, all five teasers, the Trust essay, the live Carrier IQ application.
- NOT CITABLE: Full case URLs. The private essay. TinyFish as a portfolio case. UAT Sentinel.
EQUINOX_TEASER consequence. Every other teaser title carries scope or a number. This one carries neither, so the link does no selling on its own. Put the metric inside your sentence. Never leave it to the anchor text.
Longitudinal Control Record
USE_WHEN: choosing which case opens, in any format, for any audience.
This is the organizing claim of the cycle. The file no longer treats the work as four case studies. It treats it as one property, observed four times at rising altitude, and the property reads:
"A decision made under uncertainty, exposed to contrary evidence, corrected on that evidence, and left behind in a form other people can still inspect and revise."
Altitude means the level of the organization the work actually operated at, which is another way of naming the panel it belongs in front of. Craft is the pixel and the state. System is the pattern across a surface millions of people touch. Leadership is the operating model still running after the designer leaves.
Shipping is table stakes. Everyone on your slate shipped. Director+ panels are testing two other things: whether your judgment survives contact with evidence that says it was wrong, and whether it keeps working once you are gone. Four discrete case studies answer neither question. The sequence answers both.
Sequence
L1 ALTITUDE: CRAFT Carrier IQ (Agentic Labs)
PROPERTY_SHOWN: A wrong output made visible as wrong, in the
second before a human acts on it.
USE_FOR: IC panels, design directors, AI-native product teams.
L2 ALTITUDE: SYSTEM Alibaba.com
PROPERTY_SHOWN: Trust in machine-surfaced information at
marketplace scale. $50B+ GMV context.
USE_FOR: search, RAG, attribution, recommendation, marketplace.
L3 ALTITUDE: LEADERSHIP Thermo Fisher / mySupply
PROPERTY_SHOWN: Review logic moved out of a person's head and
into the system. 0-to-1 pharma supply chain, $20M+ margin,
regulated.
USE_FOR: CPO, VP, hiring committees, any non-craft buyer.
L4 ALTITUDE: LEADERSHIP TinyFish (SPOKEN_ONLY, NO_LINKABLE_ARTIFACT)
PROPERTY_SHOWN: The same relocation, across a multi-product
agentic platform.
USE_FOR: verbal answers only. Never generated into an artifact.
SUPPORTING — altitude declared, position off-sequence:
Red Cross ALTITUDE: SYSTEM · OFF-SEQUENCE
(no disconfirming-evidence record)
PROPERTY_SHOWN: Non-experts deciding under time pressure.
National 0-to-1.
USE_FOR: consumer health, vulnerable populations, crisis UX.
Allee ALTITUDE: SYSTEM · OFF-SEQUENCE
PROPERTY_SHOWN: Dual-surface design across 30M members and
40K practices. Two user classes, one system of record.
USE_FOR: two-sided platforms, provider networks.
Equinox+ ALTITUDE: CRAFT · OFF-SEQUENCE
PROPERTY_SHOWN: Consumer trust calibration without expertise.
USE_FOR: paired with Red Cross, never alone.
Retail Velocity ALTITUDE: CRAFT · OFF-SEQUENCE
PROPERTY_SHOWN: Time-critical alerting turned into action.
STATUS: PROPERTY_UNVERIFIED_IN_FILE. Cleared for external use;
the property read was not confirmed against the live lab page
this cycle. Verify before generating the Rubrik row's proof or
differentiation.
USE_FOR: security, monitoring, clinical alerting.
Brand Pulse UNASSIGNED. Cleared for external use, no sequence
position, no company mapping. The silence here is deliberate.
Generate no claim about it.Sequence rule for generation. Open at the altitude your evaluator occupies. Move one level, either direction, to show range. Never open at L3 in front of a craft panel. Never open at L1 in front of a CPO. Off-sequence cases can be pulled at any point, but none of them carries the longitudinal argument, because none of them has a reversal on the record.
The spine case: Carrier IQ
Parallel auto-insurance quote orchestration. Structured intake, then agent runs fire in parallel across the selected carriers, five of eight on by default, filling forms, extracting quotes, returning a comparison.
Demonstrable in public, today:
STEP_1 BASELINE_SET: 250/500/250 · $500 deductible
STEP_2 AGENT_RUN: navigate, form-fill, extract, with step labels
and carrier URLs visible throughout
STEP_3 DEFECT_SURFACED: State Farm returns 250/500/250 ·
deductible drift flagged
STEP_4 OUTPUT_STATE: $181/mo · normalize flagged
Four states, not two: Bindable | Normalize | Referral |
Call Review
STEP_5 OPERATOR_CONTROLS: verify carrier / re-run verification,
proof attachment, follow-up prompt, case note, approve bind
CARRIER_IQ_SEQUENCE_STATUS:
NO_PRODUCTION_EVIDENCE for STEPS_6-8
(correction to system change to next-run difference)Nothing public shows a correction writing back: no marking of an extraction as wrong, no down-weighting of a source with a bad history, no verification step inserted into a following run, no release threshold traceable to a contested extraction.
Every agent company shipped a review surface this year; the surface itself is commodity. A review surface whose output changes the next run is not commodity, and it is the closing link of the property you are selling.
ACTION, THIS CYCLE, AHEAD OF ALL OUTREACH: Reconstruct steps 6 through 8 from your production record. What the operator marked. What the system did with the mark. What was measurably different on the next run. Then decide what is publishable. Until an artifact exists, generate 1 through 5 as demonstrated fact and 6 through 8 as spoken account, visibly separated. Blurring the two is overclaiming.
The Carrier IQ lab page attributes confidence scores to data completeness. Your working account attributes them to prior-run consistency. One is true. Correct the other surface.
These are not two phrasings of one mechanism. Completeness is a property of the extraction in front of you. Consistency is a property of the source's history. Different mechanisms carry different design consequences. An evaluator who reads the page and then hears the other version logs the inconsistency and never raises it.
Trust Model Revision
USE_WHEN: an evaluator raises the Trust essay, delegation, autonomy, or oversight design.
STATUS: Correction to a published position. The single-axis ladder is retired.
The Trust essay frames delegation as a climb: watched, verified, delegated, under the line Nobody delegates on day one. Prior published work stacked your portfolio proof on that staircase. The staircase does not hold.
DISCONFIRMING_EVIDENCE_BASE_1: Anthropic, "Measuring agent autonomy,"
Feb 2026. Auto-approval rises from ~20% among newer users to >40%
among experienced users; experienced users interrupt more often;
longest sessions grew from under 25 min to over 45 min in three
months. Sample is primarily software work and sampled API traffic.
STATUS: PUBLISHED. Citable. The evaluator can go check it.
DISCONFIRMING_EVIDENCE_BASE_2: the same pattern in your own
production agent review at TinyFish.
STATUS: SPOKEN_ONLY. Corroboration, never the load-bearing citation.What the published data says: experienced operators do not climb. They auto-approve more and interrupt more, at once. Familiarity widens what they let the agent do unsupervised and simultaneously sharpens what pulls them back in, because experience teaches which failures are real and which are noise. Two variables moving in opposite directions cannot ride one ladder.
Two independent controls:
AXIS_1 AUTONOMY_SCOPE — what the agent may do without approval
propose-only, reversible action, recoverable action,
irreversible action
AXIS_2 INTERVENTION_SENSITIVITY — what pulls a human back in
review-all, review-on-anomaly, review-on-threshold,
post-hoc-audit-only
NOTATION: SCOPE / SENSITIVITY
CARRIER_IQ: recoverable-action / review-on-anomaly
THERMO_FISHER: reversible-action / review-on-threshold,
hard Decision Gate ceiling on anything audit-bearing
BASIS: coordinates are readings applied to the artifacts, not
published values. Confidence: medium-high. Usable in conversation.Carrier IQ is the public corroboration that the axes move independently. Agents navigate five external carrier sites in parallel and fill forms with nobody watching: wide scope. The drift flag fires on a coverage mismatch worth a few hundred dollars: tight sensitivity. On a ladder those two facts contradict each other. As a coordinate pair they are coherent, and coherent is the correct answer for that risk profile.
The five published handoffs remap. Use the published names verbatim, because the evaluator can open the essay.
| Handoff | Axis(es) touched | Design consequence |
|---|---|---|
| Intent-Setting | Both, before execution | Not a low rung. The moment both axes get set. |
| In-Progress | Sensitivity only | Visible steps and sources let a human interrupt on evidence. Scope unchanged. |
| Output Review | Sensitivity | Where evidence updates what triggers review next time. |
| Decision Gate | Scope only | Irreversibility caps autonomy no matter how much trust has accumulated. |
| Loop Feedback | Both, over time | The only handoff that moves both. Also the most under-designed surface in the market, which is why it is your opening. |
GENERATION RULE: Two-axis model, published handoff names. Never generate "watched, verified, delegated" as a progression.
Per-Company Positioning
USE_WHEN: generating any company-specific claim. Nothing outside this section may be asserted about a company.
Tiers carry forward from July 24 and are a week old. Act = qualifying role from a direct source, credible channel, strong published-portfolio match, live trigger. Watch = concrete trigger, no current action. Re-verify posting state before outreach; past roughly six weeks, staleness is a gate failure, not a scoring penalty. Four fields per company, structured for the Act-Tier Positioning Sheets. Problem readings are inference from business model and posting language.
Act
VANTA — Head of Design
- PROBLEM: Compliance evidence an auditor has to sign. Automated collection is solved. Adjudicating contested evidence is not.
- PROOF: Thermo Fisher, L3.
- METRIC: $20M+ margin, 0-to-1, regulated.
- DIFF: You have designed the surface where a professional accepts responsibility for machine output. Most of the slate designed the dashboard sitting above it.
SEAT_ORIGIN: BACKFILL (high). Former VP of Design says
publicly she left at the end of June and that
this seat is her backfill. [source below]
ORG_STATE: FOUNDED. Team went 5 to 45 under the incumbent.
NARRATIVE_EXCLUSION: founding-a-design-org. Lead with continuity and
next-stage evidence instead.
STALE_SOURCE: vanta.com/design-careers still lists the
departed incumbent. Stale page, not org signal.Source: Kawamoto departure statement. Upgrades the moderate-confidence read in the prior Vanta brief.
AMPLITUDE — Product Design leadership
- PROBLEM: Analytics only pays when someone acts on a number they did not compute. AI-generated insight raises the trust burden and nobody owns the moment of acceptance.
- PROOF: Alibaba, L2.
- METRIC: $50B+ GMV context.
- DIFF: Attribution and confidence surfaces shipped at consumer scale. Not prototyped in a research deck.
WINDOW: UNVERIFIED. No leadership-change signal is documented in this file. Verify posting state and any recent product-leadership change before outreach; generate no timing claim.
RAMP
- PROBLEM: Agents that move money. Autonomy scope and reversibility are the entire product risk.
- PROOF: Carrier IQ, L1.
- METRIC: Four differentiated review outcomes on a live financial decision surface.
- DIFF: The working agentic financial decision surface already exists.
ARTIFACT: live URL
BREX
- PROBLEM: Spend control, where the policy exception is the interesting case and the exception path is usually unstyled.
- PROOF: Carrier IQ referral and call-review states, L1; Thermo Fisher, L3.
- METRIC: $20M+ margin under audit constraint.
- DIFF: The reject-and-escalate path built with the care usually reserved for the happy path.
MAVEN
- PROBLEM: Clinical guidance handed to non-experts making consequential calls on incomplete information.
- PROOF: Red Cross (SYSTEM, off-seq) with Equinox+ (CRAFT, off-seq).
- METRIC: National 0-to-1 disaster relief platform.
- DIFF: Trust calibrated for users who cannot evaluate the expertise underneath it. Very few AI-native candidates can say that sentence.
Watch
| Company | Design problem | Portfolio moment | Metric | Differentiation | Status |
|---|---|---|---|---|---|
| Capital One | Model output adjudicated inside a regulated institution | Thermo Fisher, L3 | $20M+ margin | Audit-survivable interaction, not model UX | REQUISITION_UNRESOLVED — posting dates conflict, no stable requisition located; verify before any action |
| Headway | Matching clinicians to patients who cannot assess the match | Red Cross, SYSTEM off-seq | National 0-to-1 | Vulnerable-population trust calibration | ACTIVE |
| Rubrik | Recovery decisions under time pressure and partial signal | Retail Velocity, CRAFT off-seq | Time-critical alerting | Alert-to-action, not dashboards (rests on the unverified Retail Velocity property; verify first) | PRODUCT_MOMENT_ONLY — agent-control vocabulary verified, no open design requisition confirmed; interview prep only, not an opening |
| CodeRabbit | AI review output an engineer accepts or rejects all day long | Carrier IQ, L1 | Four review states | You designed the rejection path | ACTIVE |
| Ambience | Clinician sign-off on AI-drafted documentation | Carrier IQ, L1 + Thermo Fisher, L3 | $20M+ margin | Sign-off as a designed surface with the evidence attached | ACTIVE |
| Render | Developer infrastructure. Upstream position; distance from the end user's decision moment reliably predicts a low design influence ceiling | NONE — generate no positioning | NONE | NONE | DEPRIORITIZE |
| Babylist | Consumer registry commerce | Alibaba, L2 | $50B+ GMV | Strongest metric match on the board, weakest decision-rights match | DEPRIORITIZE |
Pending
HackerOne, Gusto, Altana, Stripe Agentic Commerce, Suno, Front, TRM Labs, Abridge, Spring Health carry no assignment. Score before assigning:
- COMPANY_DIMENSIONS (3 points each): AI centrality, stage and equity, design influence ceiling, trajectory.
- ROLE_DIMENSIONS (3 points each): compensation, scope, craft, AI exposure, portfolio value.
- THRESHOLDS: Act = company 10+ and role 12+. Watch = company 8+ or role 10+.
No cash-compensation floor is documented. Generate none.
ALTANA — flags resolved.
FINANCING_STAGE: SERIES_C
LAST_DISCLOSED: 2024-07-29 · $200M · $1B valuation
CURRENT_STAGE_INDICATOR: CERVO_AI_ACQUISITION_2026-07-21
ADJACENT_SIGNAL: FEDRAMP_HIGH_2026-02
EMPLOYEE_EQUITY_WINDOW: UNCONFIRMED
ACTION_POSTURE: fit-qualified; equity unresolvedThe $200M Series C kills the stale Series B read. Recency confidence is moderate-high and rests on absence from the company archive rather than on positive confirmation, so hold it loosely. The Cervo AI acquisition puts agentic customs decisioning at the center of the roadmap, which is a direct decision-rights match to Thermo Fisher, L3.
Anticipated Questions
USE_WHEN: interview prep, screen prep, or drafting a written response to a challenge.
Two frames dominate this cycle. Both are asking whether the judgment is real or rehearsed. Every answer below carries a production artifact.
FRAME 1 — ADJUDICATION. Can you resolve conflicting signals from data, stakeholders, and model behavior?
"Tell me about a time the model was right and the human was wrong." Carrier IQ, L1. The State Farm run came back technically valid at a drifted deductible. An operator optimizing on price binds it. The system declines the binary and routes to normalize. Authority stays with the human while the interface shows the human their own shortcut. FLAG: the downstream change is SPOKEN_ONLY.
"Walk me through a decision you reversed." The trust model. Published a three-stage ladder. Then read Anthropic's autonomy data, experienced users auto-approving past 40% while interrupting more rather than less, saw the same pattern in my own production review, and retired the ladder for two axes because it cannot represent that. Four inspectable parts: the published prior position, Anthropic's data, the revised model, and Carrier IQ as the corroborating coordinate. Name TinyFish only as corroboration you cannot link.
"Who decides when the agent ships?" The four review states operate as release criteria, not QA outcomes. NO_PRODUCTION_EVIDENCE for the threshold-revision story. Do not generate it.
FRAME 2 — TRANSFERABLE JUDGMENT. Does experience in domain X predict competence in domain Y?
"Your domain is pharma and marketplaces. Ours is neither." Carrier IQ is the transfer and it sits at a public URL. Auto insurance was never in the spine. The mechanism came over from Thermo Fisher intact: put the evidence, the deviation, and the sign-off in one surface so a professional can take responsibility for a machine's output. Built in a domain never worked in, on judgment formed in one that looks nothing like it. FLAG: the multi-product version of that same transfer is L4 TinyFish, SPOKEN_ONLY.
"What did you leave behind?" Thermo Fisher, L3. A governance-visible workflow that outlives its designer, because the review logic sits in the system instead of in somebody's head. FLAG: L4 is SPOKEN_ONLY. No linkable artifact.
Confidence and Decay
USE_WHEN: calibrating how hard to state anything generated from this file.
| Claim | Confidence | Decay and handling |
|---|---|---|
| SITE_STATE | high | 7-day decay; re-fetch each cycle |
| TRUST_MODEL_REVISION | high on logic | published base citable; TinyFish corroboration SPOKEN_ONLY |
| TRUST_COORDINATES | medium-high | readings, not published values |
| CARRIER_IQ_STEPS_1-5 | demonstrated | stable |
| CARRIER_IQ_STEPS_6-8 | none | BLOCKING |
| CARRIER_IQ_CONFIDENCE_BASIS | contradicted | undated; high consequence, technical panels only |
| TIER_ASSIGNMENTS | dated 07-24 | 7-day decay |
| AMPLITUDE_WINDOW | none | no documented trigger; verify before outreach |
| VANTA_SEAT_ORIGIN | high | — |
| ALTANA_FINANCING | high on round | moderate-high on recency |
Prohibited
- Never link a full case URL. The server hands complete case HTML to any unauthenticated request; the only guard is a browser-side script reading local storage and redirecting. A link is an unlocked door.
- Never link or quote the private pricing essay. It returns HTTP 200 with the full text sitting behind a client-side overlay. That is not publication.
- Never cite TinyFish as a portfolio case. Speak it. Do not link it.
- Never reference UAT Sentinel externally.
- Never generate watched, verified, delegated as a trust progression. It is retired.
- Never present the TinyFish operator-behavior observation as published or inspectable evidence. Cite Anthropic for the published pattern.
- Never assert that an operator correction changes system behavior on the next run until a showable artifact exists.
- Never state a compensation floor. None is documented.
- Never cite the Vanta design-careers page as current org state. It is stale.
- Never claim confidence scores derive from prior-run consistency until the lab page is corrected to match.
- Never generate a timing or window claim for Amplitude. No trigger is documented here.
- Never generate any company-specific claim absent from this file. Return UNKNOWN and stop.
-
Altana's machine-readable system language: The Head of Product Design posting describes the design system as a semantic, machine-readable grammar that product managers, engineers, and agents compose, which is the sharpest articulation of the constraint-first argument anyone has put in a requisition this year.
-
Vercel's decision-rationale repository: Their account of teaching agents product design stores accepted design decisions alongside their reasoning, because shipped code preserves what was decided but not why, which is the portability-of-judgment problem in infrastructure form.
-
Figma's eval process as design practice: First Round's reporting on Figma Make documents golden prompts, nightly model comparisons, and a mixed regime of deterministic tests, human judgment, and AI judging, and is the closest public template for what a design definition of done looks like on a probabilistic product.
-
Gusto's escalation mandate: The Head of Design, Unified Service Platform role makes AI-to-human routing and graceful failure an explicit service-design brief for a three-person team inside an 80-plus-person org, which is worth scoring before the pending tier closes.

