Every posting in this cluster says "AI experience." Ramp's Director, Product Design posting wants design for agents, memory, and PRs. Headway wants leadership for AI-assisted clinical documentation. Gusto wants AI-native design practice applied to payroll trust and compliance. Sierra wants someone who can build evaluation frameworks, scorecards, QA gates, and agent personas.
Read those postings side by side and you see a credentialist surface: shipped AI products, led design for AI-powered features, understands how to design with and around models. Have you done this before? Show me.
Accurate, and almost useless as a guide to what actually happens in the room.
The Actual Evaluation
The buyer at these companies shares a professional fear they will not name in a posting or an interview debrief. They have witnessed, studied, or vividly imagined what happens when an automated system is confidently wrong in a moment that matters. Nursing-home patients' coverage terminated by algorithm while families exhaust appeals. Consulting reports with fabricated legal citations that looked authoritative until someone checked. Chatbot vulnerabilities that could let attackers hijack customer conversations in financial services and healthcare.
These incidents live in the buyer's peripheral vision. They never cite them in interviews. But the incidents shape what the buyer reacts to when reviewing a portfolio or listening to a candidate describe their work.
What they evaluate is consequence awareness. A candidate who shipped a recommendation engine at a consumer app has AI experience. A candidate who designed exception-first order management in pharma supply chains, where a missed flag at the batch level costs 3-5× more to fix at the delivery gate, has consequence experience. The buyer cannot always articulate the distinction. They feel it. It shows up as "something about her felt more senior" in the debrief, or "he really understood our problem" without anyone being able to say which sentence did it.
Ramp's VP Design Diego Zaks has said the quiet part publicly:
"Flags situations that benefit from human review while automating the rest."
That sentence is the actual evaluation criterion wearing casual clothes. The buyer wants evidence you know where the human gate belongs, why it belongs there, and what the interface looks like at that gate.
Sierra's trust architecture wraps LLMs in supervisory layers. Their Agent Studio runs experiments with statistical significance before promoting changes to all traffic. That is a design-research methodology applied to AI deployment: controlled exposure, measured outcomes, confidence thresholds before rollout. If you have research-to-architecture instincts, this is your language. Use it.
Headway's AI-assisted progress notes generate compliant, insurance-ready documentation, but the provider reviews before submitting. That "final review" step is load-bearing. The design question the buyer is hiring someone to answer: what does that review moment look like so the provider actually catches errors instead of rubber-stamping?
All three companies, across three different domains, are hiring for the same anxiety: who designs the moment where the human decides whether to trust the machine's output?
Three Relationships to Failure
The buyer's evaluation behavior shifts based on proximity to failure. Know which one you're facing.
Has witnessed it
Their company has already had an AI output cause a support escalation, a compliance flag, or a user trust incident. This buyer is the most specific in evaluation. They ask about failure modes before features. They want to hear you describe a system that was wrong and what you built to make the wrongness visible before it reached the user. They are allergic to demo-quality AI showcases. They have seen the demo work and the production version fail.
First-contact angle: "I've spent the last three years designing systems where a missed exception costs 3-5× more to fix downstream than if caught early. Your [agent/automation] work looks like it's hitting the same design problem: how do you make the system's uncertainty visible at the moment someone needs to act on it?"
Is building toward it
Their company is deploying AI into workflows where failure would be consequential but hasn't happened yet. Headway putting AI into clinical notes. Gusto applying AI to payroll compliance. This buyer is anxious but less precise. They evaluate through proxies: does this candidate talk about edge cases unprompted? Do they mention error states before I ask? Do they describe users who are skeptical of automation, or users who are delighted by it? The candidate who only talks about delight loses this buyer.
First-contact angle: "Your AI-assisted [notes/payroll/compliance] work means you're designing the review moment where a provider or employer decides whether to trust what the system generated. I've designed that exact moment in pharma supply chains and federal disaster relief, where rubber-stamping wasn't an option. I'd love to share what I learned about making review meaningful instead of ceremonial."
Fears it abstractly
AI is a strategic priority but the consequential deployment is still six to twelve months out. This buyer evaluates for vocabulary and framework more than for specific incident evidence. They respond to structured thinking about when humans should watch, verify, delegate, and when they should be able to interrupt or reverse. In postings, you'll spot them by what's absent: heavy "AI" language in the job title or description, but no mention of specific AI products already in production. They emphasize "AI-native design practice" or "design for AI-powered experiences" without naming the system. They are pre-deployment, buying the architect before the building.
First-contact angle: "I published a framework for trust in agentic systems — Watch, Verify, Delegate — and I've built three live agentic products that test it. If your team is designing the trust architecture before deployment, I've already been through the version where you have to decide what the human sees, what the agent handles, and where the override lives."
The Tells
How to read which buyer you're facing, in real time.
They ask about failure before success. "Tell me about a time something went wrong" in the first fifteen minutes, rather than as a late-interview curveball, means this buyer has been burned or is actively worried. Lead with Red Cross or Thermo Fisher. Both are stories where failure had direct human or operational consequences and you built systems that made failure visible early.
They probe the human-AI boundary. "How did you decide what to automate and what to keep manual?" is the tell that their energy is on the gate, on the boundary itself. Your Thermo Fisher case answers this directly: 5 agents running continuously, one human gate at batch QA release because the regulatory signature is irreplaceable. Your trust essay's Watch→Verify→Delegate ladder is the framework version of the same answer.
Watch for where their energy goes. If they lean in when you describe what happens when the system is wrong and go flat when you show the happy path, stay on the error state. You are talking to this buyer.
Sierra-style language without naming Sierra. "Evaluation frameworks," "guardrails," "confidence scoring," "controlled rollout." This vocabulary signals they have internalized the problem even if their company hasn't shipped the solution yet.
Generic AI credentials get dismissed. Snap's Director of Product Design Imani Ritchards has said publicly that "I've tried it" is insufficient. She asks why you built it, what problem it solves, whether real users tested it. The same instinct operates across this buyer archetype at the Director+ level: they are filtering for judgment, the kind that exposure alone never proves.
What Wins Them
Ranked by how often each factor is decisive.
Describe their reality back to them before they describe it to you. If you're talking to Headway: "The hard design problem in AI-assisted clinical notes is the review moment — making it real enough that the provider catches errors instead of approving on autopilot." You have just demonstrated that you understand their specific trust problem without them having to explain it. This is the single strongest first-contact move with this buyer.
Lead with consequence. Red Cross: $847K disbursed, 1,689 cases in two weeks, federal oversight built into the platform so it never became a separate gate. Thermo Fisher: exception-first design, surface the 3 orders out of 197 that need attention. These are consequence stories. They prove you have operated where errors have real cost.
Pair each consequence case with its AI-native successor in the same breath. Don't inventory your AI work separately. Close the Thermo Fisher story with: "I've already designed the five-agent version of this system, with one human gate at batch QA release because the regulatory signature is irreplaceable." Close Alibaba with the agentic sourcing loop where the agent traverses search and PDP autonomously but the buying signal threshold determines when a human confirms. Same problem, pre-AI proof, then the AI-native architecture you've already designed. The buyer's unstated fear is hiring someone who will need six months to understand AI failure modes. You collapse that fear in one sentence.
Name the human gate and defend it. "My default is to dissolve constraints into the product so they don't become process." Then give the example where the human gate is irreplaceable and you kept it. Thermo Fisher's regulatory signature. Red Cross's federal compliance. This tells the buyer you know where the line is, and you hold it with equal conviction in both directions.
What Loses Them
Leading with AI tool fluency. "Built with Cursor and Claude" as a portfolio headline tells this buyer nothing about your judgment. They care what you built, why, and what happens when it's wrong. The toolchain is irrelevant.
Happy-path portfolios. If every case study shows the system working beautifully and none shows what happens when it fails, this buyer checks out. They have seen beautiful demos. They need to see that you've designed for the moment the system breaks.
Abstract trust language without operational proof. "I believe in human-centered AI" is a sentence this buyer has heard from every candidate. It carries no weight without a specific system where you made a specific decision about where the human gate belongs and can show what happened as a result.
Defending AI experience you don't have. If you haven't shipped a production AI feature, do not pretend. Show that you've been solving the problems AI is now being deployed to solve, and that you've already designed the next version. That reframe is more credible than a stretched credential. This buyer, more than most, can smell the stretch.
Your Opening by Buyer Variant
Has-witnessed-it buyer (Ramp, companies post-incident): "I designed exception-first systems in pharma and federal disaster relief where a missed flag had direct operational and human cost. Your team is solving the same problem with AI agents: how do you make the system's uncertainty visible before someone acts on bad output? I've already built the architecture for both versions."
Building-toward-it buyer (Headway, Gusto, regulated growth-stage): "Your AI-assisted [clinical notes / payroll compliance] work means you're designing the highest-stakes review moment in your product: where a human decides whether to trust what the system generated. I've designed that moment across pharma, disaster relief, and medical aesthetics, and I've published a trust framework — Watch, Verify, Delegate — for exactly this class of problem."
Fears-it-abstractly buyer (earlier-stage AI companies, pre-deployment enterprise): "I've built three live agentic products and published a framework for trust in agentic systems. But the reason those exist is twelve years of designing for domains where automation under consequence was the core problem before anyone called it AI. If your team is building the trust architecture before deployment, I've already been through the hard version."
- Ramp's buyer route is clear: Diego Zaks is identified as VP Design on Ramp's own author page, making him the most direct design-leadership contact across your current target list.
- Headway's VP Product confirmed: Jake Poses joined as VP Product overseeing product, product design, and product marketing per a March 2025 leadership announcement, but the current Design Director posting doesn't name a hiring manager, so verify the reporting line before tailoring outreach.
- Figma's CEO on blurring roles: Dylan Field said AI is causing designers, engineers, and PMs to cross into one another's roles more often, which reinforces why this buyer archetype values judgment across boundaries over narrow AI tool credentials.
- Pennsylvania sued Character.AI over chatbot authority: The state alleged chatbots illegally held themselves out as licensed doctors, a concrete example of the trust-boundary failure that lives in this buyer's peripheral vision when they evaluate your portfolio.

