Four different capabilities hide behind the phrase "AI design experience." Companies rarely specify which one. A posting that says "experience designing AI-powered products" could be testing authorization and trust design, design-system infrastructure for AI surfaces, code-level prototyping alongside models, or end-to-end ownership of consequential automated workflows.
You have strong evidence for the first, moderate evidence for the third, and genuine gaps in the second and fourth. Matching the wrong evidence to the wrong construct wastes your strongest card on a question nobody asked.
The activation-rules framework is published. What follows applies it to AI-native credibility specifically: how to diagnose which construct is live, which evidence layer to lead with, and where honesty about limits serves you better than a stretch.
Diagnosing the construct
Three checks before you map evidence to anything.
Pass 1: Vocabulary clusters. The posting's word choices sort into families.
- Authorization and trust: permissions, approval, confirmation, reversibility, evidence, provenance, human-agent boundaries, delegation, "on whose behalf"
- Design-system infrastructure: standards, primitives, reusable patterns, evaluation workflows, quality criteria, cross-surface consistency, tokens
- Code-level prototyping: named tools (Claude Code, Cursor), "prototype in code," working prototypes, implementation trials, build exercises
- Workflow ownership: named operational flows, multiple actors and states, process replacement, adoption measurement, accountability across delivery lifecycle, cycle time, error rates
Pass 2: Check the product against the posting. Read the company's help docs, API docs, and product pages. If the product manages permissions, delegation, or agent identity but the posting emphasizes "design systems," the authorization construct is still live — the product tells you what the designer will actually touch. When the product and the posting diverge, trust the product.
Pass 3: Team structure and public statements. Five minutes on LinkedIn and the company blog. Where does design report? If design reports into engineering and the company's public AI narrative emphasizes autonomous action, authorization design is likely the primary construct regardless of whether the posting names it. Check what the CPO, CEO, or founders have said publicly about AI. Founders who publish about trust, safety, and human oversight are screening for Construct 1. Founders who publish about shipping velocity and technical capability are screening for Construct 3 or 4. A mismatch between the founder's emphasis and the posting's vocabulary tells you the posting was written by someone other than the decision-maker. The decision-maker's framing is the one that matters in the room.
Most postings activate more than one construct. Identify the primary — the capability that, if missing, ends the conversation — and lead with evidence for that.
Authorization and trust design
What they're testing: Can you design the systems that determine what an agent is allowed to do, on whose behalf, with what evidence, subject to what human gates, and how authority expands or contracts over time?
Posting signals: Permissions, access control, agent identity, approval flows, confirmation states, evidence requirements, provenance, recovery from wrong actions, irreversible-action gates, human-agent decision boundaries.
You'll hear: "How would you design the approval flow for an agent acting on a user's behalf?" or "How do you think about what an agent should be allowed to do without checking with the user?" or "Walk me through how you'd handle an agent that can take an irreversible action."
Lead with: The Trust essay framework — the five handoffs (Intent-Setting, In-Progress, Output Review, Decision Gate, Loop Feedback) and the Watch/Verify/Delegate progression. Then Carrier IQ as a working surface: structured intake, four distinct review states (Bindable, Normalize, Referral, Call review), per-carrier re-verification, and an Approve Bind gate that makes those handoffs inspectable. The essay gives the vocabulary, and Carrier IQ shows it running.
Confidence: High. The essay defines the problem space with original terminology. Carrier IQ demonstrates authorization design where different exceptions route to different human actions, and an operator can challenge one carrier result without restarting the full workflow. Thermo Fisher's regulatory QA release adds shipped production weight for irreversible-action gates.
When asked: "The Trust essay defines five handoffs where human authority must be designed, not assumed. Carrier IQ implements three of them — structured intake, staged evidence review with exception routing, and a consequential approval gate. I can walk you through how the review states handle different exception types differently."
Design-system infrastructure for AI products
What they're testing: Can you convert judgment about AI interaction patterns into governed, reusable infrastructure that other teams ship with?
Posting signals: Standards, shared components, pattern libraries, design tokens for AI surfaces, evaluation workflows, quality criteria, feedback loops, cross-team adoption of design primitives.
You'll hear: "How would you build a shared component for AI confidence display that works across three product surfaces?" or "Walk me through how you'd establish design standards for AI-generated content across our product." or "How do you think about consistency when multiple teams are building AI features?"
Lead with caution. Your closest evidence is the Labs considered as a set — the repeated patterns across all three (traces, evidence display, state management, review controls) suggest systematic thinking about AI interaction primitives. The Trust essay's shared vocabulary functions as a conceptual pattern language. A design-systems evaluator wants repeatable structures, and the cross-Lab consistency demonstrates that instinct even without a published library. But no public artifact shows a governed AI pattern library, a shared component system, or adoption tracking across multiple product teams.
Confidence: Use with caution. You can credibly discuss the reasoning behind reusable AI interaction patterns. You cannot point to a published artifact demonstrating infrastructure governance — the standards, versioning, and cross-team usage that a design-systems evaluator looks for.
When asked: "Across the three Labs, I've been working with a consistent set of interaction patterns for agent transparency — traces, evidence display, confidence signals, review states. The Trust essay codifies the reasoning. I haven't published a standalone pattern library, but I can show you the shared logic across the three applications and talk about how I'd govern it at scale." Be direct about the gap. An evaluator screening for systems infrastructure will respect that more than a stretch claim.
What would shift this: The Confidence/Authority Matrix and Contract-Variance Disclosure artifacts, if published, would give this construct dedicated visual surfaces with reusable rules and state definitions. They remain unbuilt.
Code-level prototyping alongside models
What they're testing: Can you build working things with AI tools — make them run?
Posting signals: Specific tools named, "prototype in code," codebase collaboration, implementation trials, build exercises. The code-construct dossier already covers the three distinct versions of "can you code" — prototyping dependency, cross-functional technical fluency, and production-engineering ownership. Identify which version is live before choosing evidence.
You'll hear: "Can you show me something you've built?" or "Walk me through how you'd prototype this — what tools, what sequence?" or "Tell me about a time you went from idea to working product without handing off to engineering."
Lead with: The three running Labs. Retail Velocity is the clearest live pipeline — data acquisition to ranking to operational recommendation. Carrier IQ is the richest stateful application. Both are inspectable: an evaluator can open them, run them, see the interaction architecture. Brand Pulse demonstrates parallel agent orchestration with source-labeled evidence. A live application the evaluator can use answers the prototyping question more directly than any process narrative.
Confidence: Moderate. Running products are stronger evidence than mockups or case studies. The gap is process visibility — no public repository, commit history, prompt history, or rejected-iteration record. The evaluator sees what you built but not how. For a company screening for prototyping fluency, the running product may suffice. For one screening for development-process judgment, narrate the build process verbally and be specific about what you tried that failed.
When asked: "All three Labs are live — you can run them now. Carrier IQ is the most complex: structured intake, parallel carrier agents, four review states, re-verification, and an approval gate. I built these using [name the tools honestly]. I can walk you through the build decisions, including what I tried that didn't work." Evaluators screening for code-level fluency want to hear iteration judgment.
Retail Velocity was built on TinyFish Search, Fetch, and Agents. Describe it as a Juno-authored application and interaction layer built on named TinyFish infrastructure. Do not claim it as wholly independent.
Ownership of consequential automated workflows
What they're testing: Have you owned an automated process end-to-end — from diagnosing the problem through designing the workflow, shipping it, measuring it, and iterating on it in production?
Posting signals: "Own" or "redesign" a named process end to end. Configuring, deploying, monitoring, and improving agents. States, routing, permissions, handoffs, escalation, auditability. Cycle time, quality, error rates, adoption, operational outcomes. Replacing or automating an existing manual process. Named operators, clinicians, support teams, or business functions whose work changes.
You'll hear: "Tell me about a workflow you owned from problem diagnosis through production iteration." or "How would you redesign our [named process] to incorporate AI agents?" or "What would you measure after launch, and how would you use those signals to change the system?"
Current postings that clearly activate this:
- Notion's Workflow + Process Designer, AI Enablement — owns business processes across intake, triage, execution, approvals, escalations, and reporting, then ships the agents that run them.
- Assembled's Product Designer, AI Agents — owns the operator lifecycle for AI support agents from configuration through production improvement.
- 3Y Health's AI Product Designer — drives end-to-end delivery of agentic healthcare workflows including production QA and instrumentation.
Lead with: Carrier IQ for the AI-specific workflow surface — intake, execution, exceptions, evidence, re-verification, and approval joined into a single operational flow. Then layer Thermo Fisher and Red Cross for shipped high-consequence outcome evidence. Thermo Fisher's regulatory QA release is particularly strong: a workflow where the downstream error is irreversible and attributable, and a human retained the gate. Lead with Carrier IQ rather than the traditional cases alone because workflow-ownership evaluators want to see that you can design the agent-specific operational states — exception routing, re-verification, evidence trails — not just the human process around them. Thermo Fisher proves you've shipped consequential outcomes. Carrier IQ proves you can design the AI-native workflow architecture this construct specifically tests for.
Confidence: Use with caution. Carrier IQ demonstrates the interaction architecture of a consequential workflow. It does not publish customer adoption, field outcomes, or end-to-end production metrics. The distance between "designed this workflow" and "owned this workflow through production iteration and measured its outcomes" is exactly what this construct tests for. Thermo Fisher and Red Cross partially bridge it — shipped products with real operational stakes — but they predate the AI-native framing.
When asked: "Carrier IQ demonstrates the workflow architecture — structured intake through exception routing through approval. Thermo Fisher is where I've owned the highest-consequence version of this: regulatory QA release where the agent surfaced batch and exception data but a human retained the release decision because the downstream error was irreversible. I can speak to what I'd measure and how I'd iterate in production." That signals awareness that ownership extends past launch, which is the gap this construct probes for.
Classification and what shifts it
The Issue 8 objection classification tagged AI-native credibility as hybrid: part perception gap (evidence exists but is not being seen or correctly mapped), part evidence gap (specific proof is genuinely missing). That holds.
The activation rules above manage the perception-gap portion — leading with the right evidence for the right construct so the evaluator encounters it before forming a judgment.
The evidence-gap portion requires building. No public artifact currently shows a human correction being recorded, incorporated into a versioned system change, and followed by improved behavior in later runs. That temporal record — bad output, diagnosis, change, verified improvement — would partially address AI depth, hands-on craft, and technical process simultaneously, as covered in Issue 8. The Carrier IQ Correction Lineage artifact is specified for this purpose. It remains unbuilt after five specification cycles. Publishing it moves the broadest portion of the evidence gap. Publishing the Confidence/Authority Matrix and Revocation Sequence makes authorization reasoning inspectable beyond the essay. Neither exists as a linkable portfolio surface today.
Until they do, the classification stays hybrid. The activation rules handle the perception side; the unbuilt artifacts address the evidence side.
Exit condition
A role whose entry requirement is production-engineering ownership of ML systems or multi-year tenure inside an AI research lab is beyond what any reframe bridges. If the posting requires shipping model infrastructure, training pipelines, or evaluation frameworks as an engineer — not designing the interfaces around them — the gap is structural. Recognize it early. Move to the next target.
Quick-scan reference
Before outreach or an interview, identify the primary construct and pull the matching row:
| Primary construct | Lead with | Confidence |
|---|---|---|
| Authorization and trust | Trust essay (five handoffs, Watch/Verify/Delegate) + Carrier IQ review states and approval gate | High |
| Design-system infrastructure | Labs as a set (shared patterns) + Trust essay vocabulary. Name the gap | Use with caution |
| Code-level prototyping | Running Labs (Retail Velocity for pipeline, Carrier IQ for stateful complexity). Narrate the build process | Moderate |
| Consequential workflow ownership | Carrier IQ workflow architecture + Thermo Fisher/Red Cross for shipped consequence weight. Signal production-iteration awareness | Use with caution |
Before every conversation: which of these four things is this company actually testing when it says AI design experience? Answer that, then choose your evidence.
- Anthropic's AI-use policy: Anthropic encourages AI for application prep but prohibits it in assessments unless explicitly allowed, which means the same company can screen for AI fluency and independent judgment in adjacent interview rounds.
- OpenAI Identity role specifics: The current OpenAI Identity designer posting names agent-to-agent authorization and mental models for delegated access, making it the clearest live example of a Construct 1 role at a frontier company.
- Correction lineage remains unbuilt: The Portfolio Playbook's five-cycle accountability report confirms zero published forward-looking artifacts, which means the evidence-gap half of the hybrid classification has not moved since Issue 8.
- Workflow ownership vocabulary is sharpening: Notion's AI Enablement posting explicitly separates states, routing, permissions, auditability, and human-in-the-loop rules as design responsibilities, giving Construct 4 its most concrete employer-side definition yet.

