Alibaba.com
The human is a professional B2B buyer evaluating cross-border suppliers on a marketplace screen. Real procurement spend on the line. The information gap is everything that would make a supplier credible in a face-to-face meeting but is invisible online: fulfillment reliability, responsiveness, order terms, transaction protection.
Your research found that 25% of traffic generated 80% of transaction value. The highest-value buyers were making the highest-value decisions without the signals that would let them commit. The system was showing a catalog. They needed a credibility dossier.
Three surfaces redesigned: homepage, search, PDP. On the homepage, B2B credibility signals moved above the fold and the EIN gate blocking high-value buyers was removed. In search, trust signals appeared at the card level: Response Rate, On-Time Delivery, MOQ. On the PDP, inline tier pricing, live order calculation, and Trade Assurance entered the primary scan zone.
Each intervention does one thing. It takes information the buyer needs to commit and puts it where the buyer is already comparing. The decision moment here is supplier evaluation before commitment. The design response makes that evaluation possible without leaving the surface where comparison is happening.
Thermo Fisher mySupply
The human is a supply-chain or pharma-operations manager running six partnerships across nine sites on three continents. The stakes are batch integrity, delivery windows, regulatory compliance. What's missing is exceptions. When something goes wrong with an order, a batch, a forecast, or a capacity commitment, the system was not surfacing it until the delivery gate, where fixing it costs 3–5x more than catching it a week earlier.
You did 32 warehouse-floor interviews and translated five failure modes into five modules. Exception-first order management surfaced the few orders that required attention out of a larger set. A batch kanban flagged at-risk batches before delivery gates. Bidirectional KPIs created shared visibility into performance gaps across partners. A forecast portal replaced 30 minutes of manual entry with SAP auto-load. An 18-month capacity heatmap made constraints visible before commitments were made against them.
The decision here is exception detection under operational complexity. The human cannot review every order, every batch, every forecast line. The system has to filter and present only what requires human judgment. Your design made the exception the default view. The first thing they saw, before anything else demanded attention.
Two cases in, two completely different domains. Different industries, different users, different risk profiles. But the structural problem is the same: a human who has to decide, information that's missing or buried, and a system redesigned so the critical signal reaches the human at the moment of decision.
American Red Cross
The human is a disaster-relief caseworker. The adjacent humans are field volunteers, finance officers, and program directors. The stakes are aid decisions for displaced families under surge conditions. The information they needed was fragmented across six legacy systems, which meant families waited while caseworkers toggled between disconnected tools trying to assemble enough context to act.
You built one unified record for three roles on Salesforce SLDS, and the design principle you articulated was "complexity in the system, not the interface." Three roles, three different decisions, one shared case. Each role needed different information from the same case record to make its specific decision. Field volunteers acting at the point of contact, finance officers approving disbursements, program directors allocating resources during surge. None of them could get what they needed without toggling across six disconnected systems. The unified record gave each role the context it needed to act without reconstructing the case from scratch.
The platform was designed to survive 10x volume with zero retraining. That constraint is itself a design decision about the decision moment: when a disaster scales, the system cannot require the human to learn new workflows. The legibility has to hold under pressure that makes learning impossible.
Six months from zero to national deployment. $847K disbursed, 1,689 cases opened in the first two weeks. Those numbers matter because they prove the system worked under the exact conditions it was designed for. Real people, real urgency, real scale.
Allē
The human splits in two: a consumer member in a medical aesthetics loyalty program, and a provider practice team managing patient relationships.
Your strategic premise was that loyalty in medical aesthetics is a trust problem, and that reframe changes what the system needs to show at every surface.
On the consumer side, you turned a static balance into a progress milestone. The treatment catalog became a saved care plan. Appointment receipts gained before/after evidence. Each move shifts the consumer from "how many points do I have" to "where am I in my treatment journey and what comes next." What drives the decision is continuation. Do I return. Do I trust this path. Do I have enough evidence to commit to the next step.
On the provider side, you replaced an alphabetical patient list with urgency sort by default. Checkout gained patient context: tier, expiry, birthday month. Allergan's broadcast emails became a provider-voice campaign builder. What drives the decision is prioritization. Which patient needs attention now. What context matters at the point of transaction. How do I communicate in a way that fits the relationship rather than a corporate template.
Two humans, two decisions, one paired architecture. The consumer needs evidence to trust the journey. The provider needs context to act on the relationship. The system serves both from the same underlying data, shaped differently for each role's decision.
Equinox+
The human is a premium fitness member across a five-brand ecosystem. The structural problem tracks with every case above: a human deciding what to do next, and a system that either makes the right next action legible or leaves them scanning.
Your field research and retention data identified that followership was the real information architecture. You designed instructor-as-navigation: four format-specific design languages, brand-led doorways from a shared component base, a QR club loop connecting app and physical clubs, unified activity history across formats.
What's at stake is engagement continuity. The member finished a class with an instructor they connected with. What's next? The system either surfaces the instructor's next session, the related format, the club where that instructor teaches, or it doesn't. Three months from zero to MVP, 4.8 App Store rating at launch.
Five cases across five domains. The pattern accumulates from the work itself.
Agentic Labs
The pattern extends here because the human is no longer the only actor.
Brand Pulse monitors brand sentiment across Reddit and X in real time. Mentions scored as they land for sentiment, urgency, and source. The system produces an hourly live brand health score and a weekly narrative written by the agent. A brand operator has to decide whether to trust the agent's scoring and narrative enough to act on a live reputation signal. The agent has already interpreted the data. The human's job is to evaluate that interpretation and decide whether to act on it.
Retail Velocity handles CPG market expansion, field intelligence, and channel audits. A market operator deciding which field signal or expansion opportunity deserves action, based on intelligence the agent has already gathered and prioritized.
Carrier IQ automates insurance quoting: form completion, coverage selection. A human deciding whether to trust or correct automated outputs in a regulated context where errors carry real financial and compliance consequences.
The same problem you've been solving since Alibaba, extended into territory where the system acts first. The earlier cases gave humans better information to decide with. The agentic cases give them an agent's conclusions. The human must evaluate whether to trust, override, or delegate further.
Where the Earlier Cases and the Agentic Work Converge
Your Alibaba case closes with a forward-looking extension: a natural-language procurement brief, agent traversal across search and PDP, autonomous order completion when buying signals are clear. The same sourcing decision that required trust signals on a marketplace screen now requires trust in an agent that has done the sourcing for you. The buyer still needs to know: is this supplier credible, is this price right, is this order safe. The agent has already answered those questions. The buyer must decide whether the agent's answers are good enough to act on.
Your Thermo Fisher case closes with a five-agent supply coordinator running continuously, with one non-automated human gate: batch QA release for regulatory signature. Five agents handling the exception detection, forecast reconciliation, and capacity monitoring that the original five modules made visible to humans. But the regulatory release stays human. That is a design decision about where autonomy ends and human authority holds. The agent can flag, prioritize, recommend. It cannot sign off on a batch that enters a regulated supply chain.
The Framework You Already Published
You articulated this pattern yourself. Your essay on trust in agentic systems frames the shift: design moves upstream from a path to a known outcome toward communication of intent, interpretation, and trust across handoffs. Your five-handoff framework — Intent-Setting, In-Progress, Output Review, Decision Gate, Loop Feedback — and your Watch-Verify-Delegate trust ladder are the structural vocabulary for the same problem your cases have been solving in specific domains.
Those frameworks map directly onto the work. The Thermo Fisher batch QA release, where five agents run continuously but the regulatory signature stays human, is a Decision Gate. The Alibaba agentic coda, where the agent completes the order autonomously when buying signals are clear, is Delegate. Brand Pulse's hourly brand health score, where the agent has already scored and the human evaluates the output, is Output Review. The framework came from the same design problem, described at a different level of abstraction.
"Accuracy is what the model achieves. Consistency is what the design delivers."
After walking through the cases, that line lands differently. Accuracy is the agent scoring sentiment correctly in Brand Pulse, or the exception-first module surfacing the right orders in mySupply. Consistency is the design making that accuracy legible, trustworthy, and actionable at the moment the human needs to decide.
The essay is the pattern your portfolio has been generating, stated in its own terms.
Where the Pattern Is Strongest and Where It Needs Qualification
Undeniable: Alibaba and Thermo Fisher. The decision moment is most concrete, the missing information most specific, the design intervention most directly traceable to the gap. A buyer who couldn't evaluate suppliers can now see Response Rate and Trade Assurance in the scan zone. An operations manager who couldn't catch exceptions until the delivery gate now sees them first. Clean line from problem to designed response.
Very strong: Red Cross. The three-role unified record under surge conditions is a pure expression of the pattern. Each role's view is tuned to its decision. Minor qualification: the case study's primary contribution is also about deployment speed and platform architecture, so the decision-moment lens is one of several valid readings rather than the only one.
Strong with qualification: Allē. The paired consumer/provider architecture is real and the trust reframe is exactly the kind of design judgment this pattern describes. But the case study serves multiple arguments — loyalty redesign, multi-role architecture, medical aesthetics domain expertise — and the decision-moment reading shares the stage with them.
Real but lighter: Equinox+. The structural problem is the same. The stakes are lower. Engagement continuity in a fitness app is a legitimate design problem, but it does not carry the weight of procurement risk, regulatory failure, or disaster relief. The pattern fits. The gravity is different.
Most explicit but thinnest on mechanics: Agentic Labs. Brand Pulse, Retail Velocity, and Carrier IQ are where the human-system operating contract becomes the stated design problem rather than an implicit one. But the published evidence is thinner on specific override controls, confidence displays, and escalation triggers than the earlier cases are on their concrete interventions. The pattern is most visible here. The proof is most detailed elsewhere.
That asymmetry matters. Your strongest evidence for the throughline comes from the cases where you weren't yet naming it. Your explicit articulation of the pattern in the agentic work and the trust essay is the framework. The earlier cases are the proof. The framework gives the proof a language. The proof gives the framework weight.
- Explanation design tradeoffs: Chen, Liao, Vaughan, and Bansal found that feature-based explanations actually increased overreliance compared to example-based ones, which matters for how Brand Pulse and Carrier IQ surface agent reasoning to the human evaluator.
- Cognitive forcing and trust calibration: Buçinca, Malaya, and Gajos found that interventions requiring more cognitive effort reduced overreliance but were rated less favorably by users, a tension that sits directly inside the Watch-Verify-Delegate ladder's design problem.
- EU AI Act human-oversight requirements: The regulation specifies that high-risk AI systems must let assigned humans understand limits, monitor operation, detect anomalies, override outputs, and interrupt the system, which maps almost exactly onto the five-handoff framework you published.
- NIST AI Risk Management Framework: NIST's framework is built around incorporating trustworthiness into the design, development, use, and evaluation of AI systems, and the language of "trustworthiness by design" rather than "trustworthiness by explanation" aligns with the consistency-over-accuracy distinction in your essay.

