The full case at junochen.com/cs-01-alibaba is password-gated: an unauthenticated visitor sees a login page. But the complete HTML — prose, metrics, screenshots, prototype — ships in the page source before the client-side check runs. This audit covers everything in that delivered response, as published August 9, 2026.
Three chapters: a diagnostic naming the desktop-procurement gap, three sprints addressing it surface by surface, and a forward-looking agentic sourcing prototype. Four headline metrics sit in the hero: +7% DAU, +20% daily transactions, +47% new visitors from search, +2.2-point NPS.
The foundation works. The diagnostic-to-sprint mapping is clean, the visual evidence is specific, the agentic thread connects to the rest of the portfolio. But there are gaps a Director-level panel will find, and most of them close with a sentence or a single added state, not a rewrite.
If you're revising before a specific search, fix these in order:
- IC + manager proof. The gap most likely to cost you the role. One team-leadership moment fixes it.
- Tradeoff. One unresolved tradeoff in Sprint 2 or 3. Fast to add, disproportionate credibility gain.
- Research as mandate-building. You're underselling a Director-level capability. One sentence naming the assumption the research overturned converts discovery into organizational persuasion.
- Hero metric attachment. Connect each headline number to a sprint.
- Agentic control state. One intervention or uncertainty state in the prototype. Matters most for AI-native searches.
- Sprint rationale and diagnosis provenance. Depth questions for later rounds, not first-screen filters.
Here's the full audit across all six lenses.
1. Problem as Juno's diagnosis
You claim first-person authorship in the hero: you identified that desktop carried 25% of traffic but 80% of transaction value, and that the platform was designed for consumer browsing while its users were conducting business procurement. Three research inputs feed three diagnostic cards that each name a surface-specific failure.
The causal chain is implied, not established. You say you saw what others missed and built the research case. But nothing on the page distinguishes your diagnosis from a pre-existing redesign brief. A brief that said "redesign the desktop experience" could produce the same narrative structure. What's missing is the moment of recognition — and whatever organizational indifference or resistance you had to push through to act on it.
High confidence this gets asked: "Was there already a desktop redesign on the roadmap when you started, or did you create the mandate?" The case can't survive this cleanly because it never shows the before-state of organizational awareness. If you created the mandate, prove it. If the redesign was already planned but you reframed its scope, that's a different and still valuable story — but the current framing oversells the second version and underproves the first.
Add one sentence naming the organizational state before the research: what was the company's existing plan for desktop, and what changed because of what you found. That sentence separates "I led a redesign" from "I changed what the redesign was about."
2. Research as mandate-building
The hero says you "secured the mandate." The diagnostic section presents three research methods and converts findings into sprint-ready problem statements. The research is well-structured and the diagnostic cards are specific.
The persuasion is invisible. No executive audience named. No disputed assumption the research overturned. No decision meeting, budget change, scope expansion. The 32 interviews appear but their organizational function is unstated — were they conducted to convince leadership? To align a cross-functional team? To validate a hypothesis you already held?
This is where the case undersells you most. If you actually changed what leadership believed — converted a "desktop is declining, nothing to do" assumption into "we're creating a trust problem we can fix" — that's organizational persuasion, and it's one of the most valuable things a Director-level candidate can demonstrate. The case claims this in the hero and then narrates it as a standard research phase. The proof is buried under competent methodology.
High confidence this gets asked: "Who did you have to convince, and what did they believe before your research that they didn't believe after?" Nothing on the page answers this.
Name the specific organizational assumption the research overturned and the decision it unlocked. Even one sentence: "Leadership believed the desktop decline was a secular trend; the gaze-tracking data showed it was a trust problem we were creating." That converts the research section from discovery to mandate-building.
3. Sprint rationale
Each diagnostic failure maps to one sprint. The order is stated: homepage first, search second, PDP third. The mapping is clean. The causal chain from problem to sprint is explicit. The causal chain explaining the sprint sequence is missing entirely.
You explain why each sprint exists but never why they run in this order. Did the homepage sprint need to ship before search could be redesigned? Did you sequence by risk — lowest-risk surface first to build organizational confidence? Engineering dependencies? Business priorities?
Without sequencing rationale, the three sprints read as a list you executed rather than a strategy you chose. This distinction is unlikely to filter you out of a first conversation. It becomes a leadership-evaluation question in second and third rounds, where the panel is testing whether you sequence work strategically or work through a queue. The sprint order is where the case could prove that, and currently doesn't.
"Why homepage first?" If your answer is "because that's where the user journey starts," that's journey-mapping logic. If it's "because the PM was most skeptical about the sign-in change and we needed an early win to protect the search and PDP sprints," that's a leadership decision the case is leaving on the table.
Add the strategic reason for the sprint order. If the order was dictated by dependencies, say so. If it was a deliberate risk-sequencing choice, say so.
4. IC + manager proof
The role line says Head of Design & Research, North America. The team line lists design, research, product management, and engineering. First-person verbs describe specific design decisions: replacing the EIN registration wall, adding Trade Assurance badges to search cards, surfacing tier pricing on the PDP.
Authorship is never allocated. Every decision is narrated in first person, which could mean you made the design decisions yourself, or that you led a team that made them, or both at different moments. At Director+ level, the panel is evaluating two capabilities simultaneously — your craft and your leadership — and a case that blurs them lets neither register clearly.
The team appears in the role description and then vanishes from the narrative. The case includes no hiring decisions, design reviews, delegation moments, coaching interactions, or conflicts between your craft judgment and a team member's direction. I've written about this IC-manager hybrid problem before: when two interviewers can't tell which version of you they're evaluating, the debrief stalls. They write down two different candidates. The room resolves discomfort by advancing whoever was easier to describe in one sentence.
High confidence this gets asked: "Which of these design decisions did you make yourself, and which did you guide someone else to make?" Followed by: "How did you run the team?" The case has nothing for either question.
Pick one sprint and show the team in motion. Name a specific moment where you directed a designer's work — gave a critique, made a call the designer disagreed with, delegated a surface and accepted the result. One concrete team-leadership moment changes the evaluative weight of the entire case.
5. Tradeoff
One explicit tradeoff: the product manager worried that removing the EIN sign-in wall would reduce conversion. The case reports that sign-up increased by 4%, resolving the disagreement in your favor. This is the only named conflict, constraint, or sacrifice across all three sprints.
Sprint 1's tradeoff is the strongest moment in the case, and you're underclaiming it. A contested decision with a named counterparty, a specific design bet, and a measured outcome. This is exactly what Director-level panels are evaluating. It's currently one paragraph in one sprint. Promote it — structurally prominent enough that a reader scanning the case encounters it within the first minute.
Search and PDP have no equivalent. Both sprints read as frictionless execution: problem identified, solution designed, metrics improved. Experienced evaluators read frictionless execution as incomplete storytelling, because real projects at this scale always involve something that was sacrificed, constrained, or fought over.
The sign-in tradeoff also has a subtler problem: it resolves too cleanly. The PM was wrong, you were right, the metric proved it. Real tradeoffs often don't resolve that way. You make a call, accept the cost, move on. A tradeoff that always resolves in your favor reads as selective memory.
High confidence this gets asked: "What didn't you ship?" or "What did you have to give up to get the PDP changes through?"
Add one unresolved tradeoff to Sprint 2 or Sprint 3. Something that was cut, delayed, or accepted as a known cost. If the tier-pricing display required a concession to engineering or a regional market, say so. If the search-card trust badges displaced something else above the fold, name what was lost.
6. Agentic sourcing vision
Chapter 3 is explicitly labeled as what you would build today, not as part of the 2024 shipped work. It connects the three redesigned surfaces into a natural-language-to-transaction loop: chat converts intent into a structured brief, search evaluates suppliers against that brief, the PDP confirms terms before an autonomous transaction. The prototype is CSS-built and visually specific.
The homepage tags this case with "Agentic Payment," and the Trust essay provides the conceptual bridge between Alibaba's cross-border trust problem and agent-mediated transactions.
The thread from enterprise trust work to agentic sourcing is coherent. The case earns this connection because the three sprints genuinely address the same trust signals — supplier credibility, pricing transparency, transaction confidence — that an autonomous purchasing agent would need to evaluate.
What's missing is the control layer. The prototype shows an agent that determines what to buy, when to buy, and how to pay. It does not show who defines the clearance rules, what happens when the agent's confidence is low, how a buyer reverses or modifies an autonomous decision, how spending authority is bounded, or what the agent does after a rejection. I've written about this in the two-dials framework: autonomy scope and intervention sensitivity are independent controls, and a prototype that shows the autonomy without the intervention design is showing half the problem.
At an AI-native company, the question will be: "What happens when the agent is wrong?" At a regulated or enterprise company: "Who approves the purchase, and what does the approval interface look like?" The prototype currently answers neither.
Add one state showing the agent uncertain or the buyer intervening. A confidence indicator on a supplier recommendation, a spending-authority threshold that triggers human review, a "why did you choose this supplier" explanation the buyer can interrogate. One control state converts the prototype from a capability demo into evidence of judgment about where autonomy should stop.
The metric problem
The four hero metrics float without attachment to any specific sprint or design decision. The sprint-level metrics are better positioned — they sit adjacent to the interventions that plausibly produced them — but none includes a measurement period, baseline, attribution method, or experimental design.
The number 47% appears twice: as the diagnostic baseline (47% of buyers cited security fears) and as a Sprint 3 outcome (security concerns fell by 47%). A careful reader will notice. A skeptical interviewer will ask whether these are the same number in different frames.
Two scenarios, two fixes. If the diagnostic 47% and the outcome 47% are different numbers — 47% of buyers cited fears before, and that rate fell by 47% from that baseline (to roughly 25%) after — state the resulting rate. "Security concerns dropped from 47% to 25%" is a clear claim a panel can evaluate. If they're actually the same data point — 47% cited fears, and the redesign addressed those concerns — then the outcome tile is restating the problem as if it were a result. Remove it from the outcome row and find the actual post-launch measurement. Either way, the current page leaves a careful reader unable to distinguish between a real outcome and a formatting coincidence.
The custody rule applies to the hero metrics: claim the mechanism, attribute the metric to the program, same breath. The sprint-level metrics are close to this standard. The hero metrics are not.
Attach each hero metric to the sprint or combination of sprints that drove it. If +20% transactions is the aggregate result of all three sprints, say so. If +47% new visitors from search is primarily a Sprint 2 outcome, say so.
- OpenAI Payments role posted: The Product Designer, Payments listing asks for trustworthy design of billing, monetization, and agentic workflows — the exact cost-authority gap the Alibaba prototype leaves open.
- Gusto's trust vocabulary: Gusto's Head of Design, Unified Service Platform posting uses "graceful failure," "human handoffs," and send/review/never-send language that would make the Sprint 1 sign-in tradeoff and Carrier IQ's correction sequence the most literal evidence to lead with.
- CHI 2026 on transparency cost: A peer-reviewed study with 12 participants found that most preferred progressive or on-demand transparency over maximal process visibility, which complicates the agentic prototype's implied design question of how much sourcing rationale to surface by default.
- Giga's non-deterministic framing: Giga's Staff Product Designer posting describes enterprise agents as non-deterministic systems and asks for correction UX — the same "what happens when the agent is wrong" question the Alibaba agentic chapter currently cannot answer.

