The cost of conflation
When a hiring committee says "AI experience," they are testing for one of two capabilities. Sometimes both, but one is almost always primary.
AI-tool fluency: you can build and ship with AI in production. Agent orchestration, trace inspection, governance, auditability. The evaluator's private question: Has she actually worked with these systems?
AI-product fluency: you can design the user-facing surfaces of AI products at the frontier. How inference cost shapes UX. How intent-based interaction replaces click-path interaction. How agent workflows create design problems around trust, delegation, and rollback. The evaluator's private question: Can she see what these products need to become?
Lead with tool fluency at an AI-native company that's trying to figure out what its product surfaces should look like, and you read as an implementer. Lead with product fluency at an enterprise company that needs someone to stand up AI-assisted workflows inside an existing org, and you read as theoretical. Both readings are hard to reverse once they land.
Your evidence stack has assets for both. They live in different places, they're in different states of readiness, and they prove different things.
AI-Tool Fluency — Your Production Evidence
Three assets, each doing different work.
TinyFish — past-tense context only
Your most recent production role. AI agent infrastructure, shipping real systems. It gives you recency and technical credibility.
Four approved uses: most recent role context, proof of technical AI currency, grounding for forward-looking ideas ("I saw this pattern at TinyFish, which is why I designed..."), and bridge narrative back to design leadership. Past tense always. Not a case study, not portfolio proof.
TinyFish proves you were inside AI production work. It does not, on its own, prove what you designed, decided, or owned.
Agentic Labs — three live, publicly accessible applications
Brand Pulse, Retail Velocity, and Carrier IQ. Built solo, live, and inspectable. These are your primary external-use proof that you build AI systems with your hands.
Brand Pulse runs parallel agents across Reddit and X to collect, classify, and synthesize brand mentions into a scored competitive view. An evaluator can inspect multi-agent coordination, sentiment classification, evidence-backed synthesis with source quotations, agent traces, and rerun controls.
Retail Velocity discovers independent restaurants, scans menus, extracts beverage products, classifies brand presence, and surfaces ranked displacement opportunities for CPG field teams. An evaluator can inspect an agentic research pipeline that converts unstructured web data into ranked operational recommendations with opportunity scores, confidence labels, and evidence snippets. One note: the live page states it's powered by TinyFish Search, TinyFish Fetch, and TinyFish Agents. If an evaluator asks about the infrastructure relationship, be straightforward.
Carrier IQ is the strongest of the three for most conversations because it crosses categories. It orchestrates parallel insurance quote agents, then exposes the full control surface an operator needs: Session/Navigate/Fill/Extract/Verify stages, per-carrier traces, review states (Bindable, Normalize, Referral, Call Review), carrier verification, case-file evidence, operator notes, and an Approve Bind action. For tool fluency, it shows staged agent behavior and trace visibility. For product fluency, it shows a human-review and approval layer designed for a regulated workflow. That dual proof is why Carrier IQ appears in deployment guidance for nearly every tier below.
What the Labs collectively prove: you can design and build working agent interfaces with trace inspection, evidence presentation, and human oversight gates.
What they leave open: customer adoption, business outcomes, accuracy validation, a documented failure-to-correction-to-next-run sequence. An evaluator who needs outcome evidence gets it from your traditional cases. The Labs prove you build. The traditional cases prove you ship at scale with measurable results. They're complements.
Traditional cases — outcome proof that maps to AI conversations
Thermo Fisher, Alibaba, Red Cross, and Allē are not AI evidence. But each answers a question that AI-hiring evaluators ask and that the Labs and Trust essay cannot answer alone.
Alibaba proves you've designed across multiple product surfaces at genuine scale, with cross-surface coherence. In an AI conversation, this answers: Can she hold a system together when it spans many surfaces and teams? Lead with Alibaba when the role involves unifying AI-assisted workflows across a large product ecosystem.
Thermo Fisher proves shipped ownership with measurable outcomes in a regulated, high-stakes domain. In an AI conversation: Has she shipped in an environment where getting it wrong has real consequences, and can she show the results? Lead with Thermo Fisher when the evaluator needs outcome data or regulated-domain proof.
Red Cross proves design leadership in a mission-critical, resource-constrained organization where human welfare is directly at stake. In an AI conversation: Does she understand that some systems serve vulnerable populations and the design constraints that follow? Pair with Thermo Fisher for healthcare and regulated tiers.
Allē proves you can own a product end-to-end and drive measurable business results. In an AI conversation: Can she connect design decisions to business outcomes? Lead with Allē at growth-stage companies where the evaluator is screening for ownership and impact, not AI sophistication.
These cases fill the proof gap the Labs leave open. When an evaluator probes past the Labs — "These are impressive demos, but have you shipped something like this at scale?" — the traditional cases are the answer.
AI-Product Fluency — Your Frontier Evidence
Two asset types. One published, four in progress.
The Trust essay — published, linkable
"Trust Is the New Interface" is your strongest current AI-product fluency asset. Public, first-person, and it makes a specific argument:
- Agent interfaces move design upstream — from optimizing a deterministic path to designing how intent is communicated and interpreted.
- Aggregate model accuracy doesn't tell a user whether this specific output is safe to act on. The interface must make consistency, confidence, evidence, provenance, and system change visible.
- Every agentic workflow contains five handoffs: Intent-Setting, In-Progress, Output Review, Decision Gate, and Loop Feedback.
- Users build trust through watching, verifying, and then delegating. Products need evidence and control appropriate to where the user currently sits on that ladder.
The essay's core line —
"Accuracy is what the model achieves. Consistency is what the design delivers."
— reframes a technical problem as a design problem in a way that sticks. It positions you as someone who has thought through the frontier design problems companies are actually hiring for: intent capture, reliability signaling, human-agent handoffs, decision gates, trust calibration.
What the essay cannot do alone: it's a framework, not a shipped product. Evaluators who have been burned by strategy-only hires will want to see the framework connected to production. That's where the Labs and TinyFish context do their work.
The four forward-looking artifact domains — all in progress
Issue #8's portfolio playbook described four domains where forward-looking visual artifacts would strengthen your AI-product fluency proof: Inference-Aware UX, Intent-Based Interaction, Agent Infrastructure as UX, and Human-Agent System Design.
As of this week, none has a published URL or a cleared artifact you can share. The domain briefs exist. The conceptual foundations are partially public through the Trust essay and the Labs. But no evaluator can currently open a link and see a dedicated visual artifact for any of these four.
This changes the deployment sequence I recommended in Issue #8. I wrote then that forward-looking artifacts should lead AI-native conversations. That assumed at least one domain would be presentable by now. Since none is, the Trust essay and Carrier IQ carry more weight than I originally assigned — they're doing double duty as both your published AI-product thinking and your closest-to-frontier production proof.
What each domain will prove once an artifact is ready:
Inference-Aware UX — you can make model quality, speed, cost, and uncertainty visible as product choices rather than hidden infrastructure. Highest value for enterprise platforms managing cost-quality tradeoffs at scale.
Intent-Based Interaction — you can replace action-by-action interaction with intent capture without turning ambiguous language into accidental authorization. The Trust essay develops this conceptually; a visual artifact would show intake, confirmation, and revision states.
Agent Infrastructure as UX — you can turn agent infrastructure (traces, task state, evidence, permissions, rollback) into an actionable control surface. Carrier IQ already implements part of this; a dedicated artifact would show the complete lifecycle including correction and later-run behavior.
Human-Agent System Design — you can choreograph how work, evidence, and temporary decision authority move between people and agents across a consequential workflow. Highest value for healthcare and regulated companies.
You can reference these domains verbally and describe what you're building, but do not promise a link you can't yet deliver.
Deployment by company tier
Use this as a lookup before outreach.
AI-native (OpenAI, Midjourney)
Lead with: Trust essay for product-fluency framing, then Carrier IQ for working-system evidence. TinyFish as past-tense production context only.
After a forward-looking artifact clears: The artifact becomes the new front door. Essay becomes foundation; Lab becomes production grounding.
These evaluators test product fluency first. Tool fluency is assumed. If you lead with "I built three agent apps," you read as an engineer who builds well rather than a design leader who sees what needs building. Lead with how you think about the design problems their product faces, then prove you also build.
Growth-stage platforms
Lead with: Thermo Fisher, Red Cross, or Allē for shipped ownership and outcomes. Add the Lab that matches the company's product problem: Brand Pulse if they deal in content intelligence or competitive analysis, Retail Velocity if they run data-to-recommendation pipelines, Carrier IQ if their product involves review, approval, or compliance workflows. When in doubt, Carrier IQ — its control layer translates across more contexts than the other two.
After a forward-looking artifact clears: Use it as an AI-product-fluency layer on top of shipped execution proof. It supplements; it doesn't replace the traditional cases.
These evaluators need to see that you've shipped at scale with results. AI fluency is a bonus that strengthens the candidacy. Don't let AI evidence crowd out your strongest outcome stories.
Enterprise platforms (Stripe, Capital One)
Lead with: Alibaba for scale and cross-surface coherence, then Thermo Fisher. Use Carrier IQ when the mandate includes agent auditability, review, or approval workflows.
After a forward-looking artifact clears: Add Inference-Aware UX or Agent Infrastructure as UX when the artifact matches the platform's specific control problem.
Enterprise evaluators often say "AI experience" but mean "can you stand up AI-assisted workflows inside a large, complex org without breaking existing systems or trust." Carrier IQ's approval gate and evidence layer speak to this more directly than Brand Pulse or Retail Velocity.
Healthcare and regulated (Headway)
Lead with: Thermo Fisher and Red Cross for regulated-domain proof. Trust essay and Carrier IQ to connect preserved human gates to current agent systems.
After a forward-looking artifact clears: Human-Agent System Design is the highest-value addition here. Its current public basis is the Trust essay's conceptual frame — strong but not standalone proof.
Regulated-domain evaluators screen for whether you understand that some decisions cannot be delegated, and that the interface must enforce that boundary. The Trust essay's Decision Gate concept and Carrier IQ's Approve Bind action are your best current evidence for this.
What to do before your next outreach
Trust essay, three Labs (Carrier IQ strongest), TinyFish as context, traditional cases for outcome proof. Forward-looking artifacts are in progress — reference verbally, do not promise links.
Before any outreach, make one decision: is this evaluator primarily testing tool fluency or product fluency?
Two cues from the posting. First, where does AI appear in the requirements? A posting that mentions AI alongside "building," "shipping," "production systems," or "agent infrastructure" is testing tool fluency. A posting that mentions AI alongside "user experience," "product vision," "trust," "interaction models," or "frontier design problems" is testing product fluency. Second, who does the role report to? Reporting to engineering or a CTO skews tool fluency. Reporting to a CPO, CDO, or CEO skews product fluency. When both signals are present or neither is clear, the role likely tests both.
When you're testing both, decide whether to layer or separate. Layer tool and product fluency together — Carrier IQ does this naturally — when the conversation is about a single product where the design leader will own both the system architecture and the user-facing surface. Keep them separate when the evaluator has already signaled which one they care about, or when the role splits those responsibilities across different people. Layering when the evaluator wants to see one thing clearly makes your signal noisy. Separating when they want to see range makes you look narrow.
If you're still unsure after reading the posting: Trust essay plus Carrier IQ. One proves you think about these problems with precision. The other proves you build solutions to them. That combination covers both fluencies and buys you time to read the room.
- Correction-loop artifact priority: The failure-to-correction-to-next-run sequence remains the highest-leverage missing proof across all four forward-looking domains, and completing one for Carrier IQ would close the Labs' most visible evidence gap.
- Retail Velocity's TinyFish dependency: The live page explicitly credits TinyFish infrastructure for its search, fetch, and agent layers, which means an evaluator who probes provenance will surface the boundary question before you raise it.
- Anthropic's design-engineering hybrid: Anthropic's published account of Claude Code product design shows designers making substantial implementation changes, which collapses the tool-fluency and product-fluency distinction in ways that would require a different lead than the Trust essay.
- FDA decision-support guidance: The January 2026 Clinical Decision Support Software guidance distinguishes software that supports professional judgment from software that replaces it, giving the Human-Agent System Design domain a regulatory anchor once an artifact is ready for healthcare conversations.

