Current-State Verification
Fetched July 24, 2026. The .html URL returns HTTP 308 to /cs-02-supply-chain; the extensionless URL returns HTTP 200 with full HTML (214,343 bytes, last modified July 19). A client-side script checks localStorage.getItem("pf_access_v1") and redirects to /login.html if absent. The complete case content ships in the response body before that script executes.
Same pattern, same severity as the Essay 02 gating issue flagged July 10. The gate is not a gate. Every heading, stat, mockup, and word of the agentic coda is readable in source. Server-side gating. No other fix matters.
A hiring manager who finds this page before you present it has already formed opinions you cannot reshape. You are holding this case back for interviews. It is the most exposed asset in the portfolio.
The Structural Spine — Derivation, Mapping, Tradeoffs (Criteria 2–4)
I'm treating these three criteria as one thread because that is how a VP Engineering or CPO will process them. Research produced architecture, architecture required decisions, and those decisions had alternatives the panel will want to see. If the chain breaks at any link, craft quality is irrelevant.
The 32 → 5 → 5 Chain (Criteria 2 and 3)
The page names thirty-two interviews. It names five failure modes with operational specificity: a 197-order flat list burying exceptions, 5–10 hours of weekly batch-status calls, Thermo Fisher and partners measuring OTD differently, forecast submission via versioned spreadsheets with manual SAP loading, capacity discovered only after commitments. Each failure mode maps one-to-one to a named module: Orders, Batches, Dashboard, Forecasts, Capacity.
A non-designer panelist can follow this. "5–10 hours of status chasing per batch" is a number anyone holds. "63–99% OTD spread by site" makes an operations-minded evaluator lean forward. The mapping from failure mode to module is explicit, traceable, requires zero design methodology knowledge.
Where it thins out is the analytical reduction. Thirty-two interviews became five failure modes. Why five and not seven? Why not three? The page shows the output of the clustering. It does not show the clustering. Any panel member who has hired research-informed designers before will probe this step, because this is where a designer who executed separates from one who led.
Add one passage between the research framing and the five failure modes. Name the clustering principle. Something close to:
"These thirty-two interviews surfaced dozens of complaints, but five were structural rather than situational — they persisted across sites, roles, and geographies."
One sentence turns a list into a derivation.
Tradeoffs (Criterion 4)
Module descriptions imply tradeoffs everywhere. Orders defaults to exception-first. Forecasts replaces spreadsheets with structured submission and version control. Dashboard shows internal and customer KPIs side by side rather than consolidating into a single metric.
Every tradeoff is embedded in the design decision. None is surfaced as a tradeoff. The panel sees what you built. They never see what you rejected. Why exception-first instead of a filtered list? Why side-by-side KPIs instead of a reconciled single metric? Why an 18-month capacity heatmap instead of a rolling quarterly view?
For at least two modules, name the alternative you rejected and the reason. One sentence each:
"We considered a reconciled OTD metric but rejected it because partners would not trust a number they could not verify against their own."
That kind of sentence is what a CPO remembers after the interview. Everything else fades.
Decision Gate Analysis
The editorial record names this case as the strongest regulated human-agent proof because of the "five agents, one human gate" structure. The Trust essay's §03 handoff #4 uses Thermo Fisher as the Decision Gate example: "Automate what you can verify. Keep a human where the outcome is irreversible."
The agentic coda says the five modules become five continuous agents. One human gate remains: batch QA release, because a regulatory signature is irreplaceable. The UI card shows a batch QA release screen with 99% confidence, all five release criteria verified, complete documents, a closed ticket, actions to approve or view the batch record.
Does it read as a designed Decision Gate or a compliance checkbox?
Compliance. "A regulatory signature is irreplaceable" frames the human gate as a constraint imposed from outside rather than a design decision about where automation should stop. The Trust essay draws a sharper line. The QA release gate is not just where regulation requires a human. It is where the consequence of a wrong release (product recall, patient safety, regulatory action) makes the cost of automation failure catastrophic. The design requires the human. Regulation happens to agree.
Reframe the human gate sentence. Replace "a regulatory signature is irreplaceable" with something closer to:
"The cost of a false release — regulatory action, patient safety, partner trust — makes this the one decision where confidence scores are insufficient regardless of their accuracy. Regulation codifies what the design already requires."
Compliance becomes design reasoning. That is what the Trust essay promises, and what a panel evaluating your judgment will test.
The Count Tension
The prose says five modules become five agents. The mockup's activity log says "3 agents active." The Trust essay says five agents ran continuously. A panel member who reads the essay before the interview and then sees the case will notice the discrepancy.
My read: the mockup likely represents a point-in-time snapshot where two agents have no active alerts. But the safer fix is visual, not textual. Show all five agents in the mockup with varying states (active, monitoring, idle). That stays consistent with both the prose and the Trust essay, and it teaches the panel something about how agent systems actually behave. Changing the prose to explain event-triggered agents introduces a concept the coda has no room to develop. Change the mockup. Keep the prose clean.
Trust Essay Bridge
The Trust essay names Thermo Fisher twice: once in the §03 Decision Gate card, once in the conclusion as "a pharma logistics system with $20 million of margin and a regulatory clock." The case needs to deliver what the essay promises. Two tests.
Test 1: Consistency Over Accuracy (§02)
The essay's §02 argues that model accuracy is invisible at the decision moment. Consistency is what design can make visible. "Accuracy is what the model achieves. Consistency is what the design delivers." The essay names seven properties: confidence, evidence, provenance, traceability, evals, observability, annotation.
The QA release card is the sharpest test of this argument in the entire portfolio. It shows 99% confidence and five verified release criteria. But the case never makes the distinction the essay makes. A panelist who has read §02 should understand, when they see this card, that the 99% number could be wrong and the design accounts for that. What makes the human gate workable is not the accuracy of the confidence score. It is the consistency of what the agent surfaces every time. The same five release criteria, the same document verification, the same ticket status, presented in the same structure so the human reviewer can pattern-match against a known format. If the agent surfaced different evidence in different orders with different framing, the human gate would be unworkable even at 99% accuracy.
The case shows the output of consistency. The card looks the same every time, implicitly. It does not name consistency as the design principle that makes the gate function. That is the gap between the essay's argument and the case's evidence.
Add one sentence in the agentic coda, near the QA card:
"The agent presents the same five criteria in the same structure for every batch, so the reviewer's decision is whether to release this batch, never what information to trust."
Stronger than using the word "consistency" because it shows the principle operating rather than naming it.
Test 2: The Seven Properties
The case currently delivers on confidence (99% shown) and evidence (release criteria verified). It does not use or demonstrate provenance, traceability, evals, observability, or annotation. You do not need all seven. But the agentic coda should use at least two more when describing what the agents surface. If the exception-routing agent flags a late batch, does it show provenance — which data source triggered the flag? Does the forecast-drift agent show traceability — the chain of forecast versions that led to the current alert? Adding these to the agent descriptions closes the bridge without requiring new mockups.
Mandate, Adoption, Vision
0→1 Mandate Under Pharma Constraints (Criterion 1)
The constraint environment lands through hero stats and metadata: twelve months from zero to live, six partnerships, nine manufacturing sites, three geographies, $20M+ annual margin recovered, 83% IRR. The problem section reinforces with operational specifics (50% of orders placed less than 90 days before commit, 45% of data outside PRISM in Excel/PDFs). A panel member absorbs the scale without needing pharma expertise.
What is missing is one sentence that tells a non-pharma panelist why the scale is remarkable.
Add one sentence in the problem framing: pharma partners operate under GxP compliance requirements that make any new system a regulatory event, not just a product adoption. That sentence tells a panel member outside life sciences why twelve months across six partners is a design leadership achievement, not just a timeline.
Partner Adoption as Design Outcome (Criterion 5)
The hero stat says 6/6 pharma partners committed. The homepage card was praised for making adoption the design problem. The case itself does not develop this. How did the design produce adoption? Was the side-by-side KPI dashboard the thing that got partners to trust the platform? Was structured forecast submission the feature that replaced their spreadsheets without adding friction?
Strong stat. Absent story.
Add three sentences in the Dashboard module description connecting the side-by-side KPI design decision to partner willingness to adopt. Partners see their own numbers next to Thermo Fisher's. That replaces the trust deficit that kept them on spreadsheets. Dashboard is also where a panel member's eye will land when evaluating cross-organizational design problems, so the adoption story meets them at the right moment in their read-through.
Agentic Vision Placement (Criterion 6)
The coda exists and does real work. It reframes a 2021 platform as a 2026 architecture. The five-agents-one-gate structure is the right move. End-of-case placement is correct. It reads as "here is what I would build now with the same problem," which is the forward-looking judgment a panel wants.
No structural change needed, contingent on the Decision Gate reframe and count fix prescribed above. Without those two fixes, the coda undercuts the Trust essay instead of proving it. With them, this becomes the strongest agentic proof in the portfolio. It is the only case grounded in a live system with real regulatory stakes and real adoption data.
- Rubrik's reversibility language: Rubrik's AI launch describes autonomous actions as auditable, attributable, and reversible with human review for irreversible actions, which is the closest external product vocabulary to the Decision Gate reframe prescribed here.
- Amplitude Wave's agent loop: Wave's description of an agent that surfaces opportunities, drafts specs, routes work, and verifies outcomes maps almost exactly to the five-agent architecture the coda proposes, making Amplitude the strongest outreach target for this case.
- Gusto's AI principles on control: Gusto's published AI principles argue that decisions should start with whether the business owner actually needs the AI intervention, which mirrors the consistency-over-accuracy argument the QA release card needs to deliver.
- OpenAI's confirmation-before-consequence pattern: OpenAI's cloud browser documentation describes pausing for confirmation before actions with financial, legal, or real-world commitment, the clearest industry precedent for the human gate design this case should be proving.

