Three chapters, three sprints, one agentic coda. Nobody who worked for you appears in any of them.
I opened this expecting the failure I flagged in your other cases: leadership facts narrated at execution altitude. That's not what's wrong here. You take the diagnosis in the first screen, in your own words, "I named it." You sequence research before mandate. You end on a forward-looking procurement chapter rather than a results recap. Most cases at this level miss all three. You clear all three.
The actual problem is narrower and much harder to fix. Read the verbs. You name a problem, you fix a homepage, you rebuild a search card, you restructure a PDP. The object is always a surface. Nowhere does a person report to you, apply a standard you wrote, or decide something you handed them. No principle, no review, no criterion survives Sprint 3. The case proves a system moved while your hands were on it and says nothing about whether your judgment kept working after you let go.
At Director+ that is the whole question.
Delta and sourcing
Opening, diagnostic chapter, three-sprint structure, metrics, agentic coda, captions. All identical to the July capture. Nothing from the last round landed.
The deployed file reports a last-modified date of July 19. That corroborates the finding; it doesn't carry it alone, because a timestamp on a static host can reflect a rebuild rather than an edit. The structural match is what makes it evidence.
On sourcing. You asked for an audit of the page, so I'm auditing what the server returns today rather than what I remember. Worth knowing how I got it. The site still ships the full case body before the browser evaluates the password gate, so a plain non-browser request pulls the entire narrative, the metrics and the captions with no credential. I documented this in issue #5 and I'm not relitigating it. One line and I'll drop it: a portfolio arguing trust calibration should not leak its own lock.
Whose diagnosis this reads as
Present, and stronger than I expected. The opening puts desktop at roughly a quarter of traffic against 80% of transaction value, states that the experience was built for consumer browsing while the actual users were running company procurement, then claims authorship outright. Three methods sit behind it in the diagnostic chapter: 32 cross-functional interviews, a Baymard benchmark audit, gaze tracking on sessions above $500. Three failure cards tie observed behavior to specific surfaces. Sign-in deficit on the homepage. No B2B filtering in search. Hidden price tiers on the PDP.
Thin where it counts. The conclusion is established. The resistance never is. A mobile-first organization at that scale did not leave a $50B-GMV insight lying unclaimed on the floor. Somebody held the position that desktop was a minor problem, and that somebody was wrong. Without them, "I named it" is authorship of an observation. Observations are cheap at this level. Arguments aren't.
Add one sentence to the opening, directly after the traffic-versus-value contrast. I can't draft it, because the fact isn't public:
"The prevailing view inside the org was ___. The signal that changed it was ___."
The second blank does the work. One number that flipped an executive belief outperforms your entire methods list.
How research became authority
Present in sequence. The opening says you "built the research case" and secured the mandate before redesign began. The teaser credits executive buy-in separately. That ordering is rare and it is correct: research as the instrument that manufactured authority, not the ritual performed after somebody hands you the problem.
Thin in mechanism. There is no decision moment anywhere on the page. No forum, no approver, no ask, no objection, no counterparty. "Secured the mandate" reports an outcome and hides the machinery, and machinery is what this buyer is scanning for. The CZI product design interview guide is one of the few published company rubrics you can read rather than infer, and it presses candidates on who their partners were and what they personally contributed inside the group. A mandate with no counterparty fails both halves of that.
Add one sentence at the close of the diagnostic chapter:
"I took the case to [role or forum]. The ask was ___. The objection was ___, and the evidence that answered it was ___."
Don't name the executive. Do name the objection. A record containing only agreement reads as a record assembled afterward.
Why three sprints
Present structurally. The bridge into Chapter 2 assigns each diagnostic failure to a sprint and states the sequence plainly: first impression, then search, then transaction, "in that order." Every sprint carries its own problem, intervention, before-and-after evidence, and two outcome metrics. Cleanest architecture on the page.
Thin in defense. You assert the order. You never defend it. Three sprints instead of one coordinated program is an operating decision with a price attached: sequential validation burns calendar and buys the option to stop. None of that reasoning appears. The skeptical read is that the sprints ran in funnel order because funnel order is the obvious order, which makes the sequence look like it chose itself.
This one I'll draft, because the logic falls out of your own diagnostic and all you have to supply is the gate:
"Three sprints, not one program, because each surface fed the next. If the homepage didn't hold procurement buyers past first impression, search improvements would land on traffic that had already left. We ran each sprint as a gate: ___ had to move before the next one opened. That cost us calendar time and bought us the option to stop."
Fill the gate with the real threshold. If no formal gate existed, say the sequencing was dependency-driven and leave it there. Don't manufacture rigor you didn't have.
Where the case fails
The failure is self-inflicted, and it sits in the gap between your title and your verbs.
Present: the hero brief names you Head of Design and Research, North America, and lists Design, Research, PM and Engineering as participating disciplines. Every sprint runs in hands-on first person, "I fixed that." The IC half of the hybrid is demonstrated repeatedly and demonstrated well. Nobody will question your craft.
Absent: the management half. All of it.
- No team size.
- No reporting line.
- No indication of which of those four disciplines worked for you versus alongside you.
- No design review, no hiring, no delegation, no coaching.
- No moment where a designer you managed applied a criterion you set and you never touched the file.
The title asserts leadership. The narrative demonstrates an exceptional senior IC with unusual executive access.
Hold that against what your actual buyers publish. Vanta's Head of Design scopes an organization near 40 people and names design-system strategy, launch review and definitions of done as the work itself. Amplitude's Head of Product Design defines a 15-person player-coach seat and asks outright for structure the candidate created rather than inherited. A widely circulated recruiter taxonomy of VP design candidates cuts the field three ways: the craftsperson who raises the bar, the translator who converts a business problem into a design mandate, the builder who creates capability that wasn't there before. CS-01 proves translator, cleanly. Then it stops.
Two fact prompts. Skip neither.
"At the start of the program I led ___ designers and ___ researchers. PM and engineering were partner functions; ___ reported to me."
"The North America research function was [created / expanded / inherited] when I joined."
Created or expanded, and that clause moves you from translator to builder inside three seconds of reading. It belongs in the hero brief, not paragraph nine. You ran a Lead UX Researcher search at Alibaba in public, which puts on the record that you were staffing senior research yourself. Pull that thread.
The single tradeoff in the entire case
One instance. Sprint 1 pulls the EIN registration wall out of the browse path, and the case notes that "the PM worried" about sign-in impact before reporting sign-ups up 4%. Best-built paragraph on the page, because the risk gets named before the result does.
And it's the only one, and it resolves in your favor, which makes it setup rather than cost. Three sprints, eighteen annotated callouts, and there is no rejected alternative, no delayed scope, no metric you knowingly surrendered, no cost that outlived launch. Every downside converts to upside inside two sentences. Meanwhile the headline results are carrying the credibility of the entire case. Unbroken good news at that scale reads as curation.
A hiring manager running a live product design search recently published the scenarios he probes for: engineering delay, user goals colliding with business goals, scope pressure, metrics declining post-launch. One post from one manager is thin evidence by itself. It stops being thin when those four scenarios map onto what Vanta and Amplitude ask for in structured form. None of the four appears in CS-01. When a case contains no unresolved cost, the reviewer supplies one from imagination, and the imagined version always looks worse than whatever actually happened.
Add one sentence per sprint. If you only do one, do Sprint 2, because I can draft half of it for you. Ranking search results by supplier response rate demoted something a stakeholder cared about. It had to.
"Ranking by supplier response rate meant demoting [paid placement / incumbent supplier visibility]. [Merchandising / supplier relations] pushed back, and they were right to — it cost those suppliers surface area. We took it anyway, because a procurement buyer who can't get a quote back doesn't convert at any price, and the response-rate data was unambiguous."
One bracketed choice, not four blanks. Same form for the other two sprints.
The agentic coda and what it won't define
Present, and better than I've credited it before. The closing chapter opens "If I were building this today" and walks a natural-language procurement brief through agent-side filtering, supplier and PDP inspection, and an order-total check, landing on three human choices: place the order, request a sample, compare more suppliers. The connective tissue to your current trajectory is on the page. Issue #4 established Alibaba as your intent-setting proof; this coda makes the link explicit.
Thin at exactly the point an AI buyer cares about. The chapter has the agent completing autonomously once buying signals clear and never defines "clear." No confidence threshold. No handling for contradictory evidence. No interruption trigger, no provenance, no recovery path when the agent buys the wrong thing. Same gap I catalogued across eight agentic patterns: trust vocabulary, no observable trust behavior.
I'll write this one in full, because it's design judgment about a proposed future rather than a claim about your past:
"The agent completes autonomously only when supplier signals agree. When Trade Assurance status and response-rate history contradict each other, or when the order total exceeds the brief beyond stated tolerance, it stops and shows the conflict rather than resolving it. Procurement buyers don't need an agent that's right most of the time. They need one whose confidence is legible enough to override."
That last sentence is your bridge into agentic-sourcing conversations. Use it verbatim.
One relocation
In July's teaser audit I argued that the multi-surface component atlas (navigation, search, cards, PDP, onboarding, tokens) was drowning the leadership narrative on the teaser. It hasn't moved. The teaser still carries it. The full case still doesn't.
Swap them. On the teaser it's noise. Inside the full case it becomes the only visible evidence that something you built had a shape that outlived a sprint. Place it after Sprint 3 and caption it with what it governed and who used it when you weren't in the room.
Verdict
CS-01 does not clear the system-leadership bar. It misses on one axis. Diagnosis, mandate sequencing, sprint architecture, forward vision: all present, mostly well executed. What's absent is an organization. No team, no reporting line, no review mechanism, no standard that persisted past your attention. A buyer scanning for builder mode finds a superb translator and has to guess at the rest. At this level nobody guesses upward.
Do the team-size and research-function lines first. Five-minute inserts, and they should have shipped with the first published version of the page. They are not the leverage.
The leverage is one paragraph between Sprint 3 and the agentic chapter. Head it "What stayed." Name the criterion, principle or review that outlived the sprints. Response-rate visibility becoming a standard in later search work. Trust-signal hierarchy entering the design system. Whatever the true version is. Then name who applied it after you stopped touching it. Three sentences, four at most. It is the only edit on this list that proves judgment operating outside your own hands.
If nothing stayed, write that instead, and write what you'd install now. That still beats silence.
I'm not drafting it. It's the one paragraph whose facts I can't establish from any public source, and it's the first thing a VP Design follows up on in the room. It has to be true, and it has to be yours.
-
Altana's machine-readable system language: The live Head of Product Design posting at Altana describes the design system as a semantic, machine-readable language of patterns and interaction grammars that PMs, engineers, and agents compose, which is the most explicit statement yet of what "what stayed" means when the inheritor is a machine rather than a person.
-
Oversight is not a staircase: Anthropic's measurement of agent autonomy found experienced users both auto-approving more often and interrupting more often, which complicates the Watch → Verify → Delegate ladder and argues for treating autonomy scope and intervention sensitivity as two separate controls in the CS-01 coda.
-
Reversibility as shipped product: Linear's agent-assisted editing release visually distinguishes agent-written changes, preserves the agent as author, and creates restorable checkpoints, giving you a live reference for the recovery path your Alibaba agent sequence currently skips.
-
What "done" means for a probabilistic product: First Round's account of Figma's AI eval process documents scoring rubrics, golden prompts, nightly model comparisons, and mixed human and AI judging, which is the release-gate machinery a Vanta or Amplitude evaluator will expect you to describe when they ask what your definition of done looks like.

