The gap is hybrid
The overarching objection — "does this candidate have real AI experience?" — is a hybrid of perception gap and evidence gap. The ratio shifts by role.
Perception gap where proof exists but gets routed wrong. You have direct evidence of human-agent control design. You have three running applications. You have a published framework for delegation and trust. When this evidence lands in front of someone screening for oversight architecture, it works. When it lands undifferentiated in front of someone screening for production code authorship, it looks adjacent to what they asked for.
Evidence gap where no proof object exists regardless of routing. You cannot currently demonstrate a commit history, evaluation pipeline ownership, or shipped consumer generative-media product work. No amount of repackaging fixes a missing artifact.
The cost of treating this hybrid as one problem: you deploy the full stack against every AI role. The evaluator screening for code sees a thought-leadership essay. The evaluator screening for oversight design sees a prototyping demo. Each one concludes you pattern-matched on the word "AI" rather than their actual design problem.
Issue #10's construct decomposition identified four assessment families. The current posting landscape requires six.
Six constructs the market actually screens for
Each is a distinct design problem that a hiring committee evaluates for. They share the surface requirement of "AI experience" and almost nothing else.
1. Human-agent control and oversight. The evaluator wants evidence that you've designed the boundary where a human decides whether to trust, override, or approve an agent's action. OpenAI's Identity role and Slack's Principal Architect search both center this. The screen: do you understand delegation as a design problem with states, gates, and failure modes?
2. Production-code authorship. The evaluator wants proof that you personally write and ship code. Not that you directed engineers. Not that you prototyped in Figma. That you authored production software. OpenAI's Codex role is the clearest example. The screen: a repository, a commit history, a deployment you own.
3. Evaluation infrastructure. The evaluator wants to know whether you've built or owned the systems that measure model quality — prompt pipelines, graders, comparison sets, regression suites, release gates. Anthropic's Evals & Prompts role screens for this directly. The screen: operational ownership of the measurement layer.
4. Rapid AI-tool-assisted prototyping. The evaluator wants to see that you build working software quickly using current AI tools and that the result is a functional application, not a deck. Midjourney's Design Lead says it plainly: "Prototypes > presentations." Headway's Provider role and ServiceNow's SalesCRM posting both ask for front-end prototyping and scripting. The screen: can you make things that run?
5. Consumer generative craft. The evaluator wants taste, emotional resonance, and the ability to invent interaction paradigms around generative capability. Suno's Head of Product Design emphasizes creative leadership and consumer craft; AI-product experience is listed as a plus, not the admission gate. The screen: can you make a generative product feel like something a person wants to use?
6. AI-native enterprise workflow. The evaluator wants proof that you can translate complex business processes — quoting, configuration, deal workflows, field operations — into conversational and agentic interaction models. ServiceNow's SalesCRM role combines enterprise-domain depth, conversational workflow design, rapid prototyping, and high craft inside a data-rich product. The screen: domain fluency married to AI interaction design.
What each proof object actually demonstrates
Your three Agentic Labs applications are not interchangeable. Each shows a different design problem. Deploying the wrong one against the wrong construct costs you credibility.
Carrier IQ
What it shows: The user sits at a consequential approval gate. Session stages (Navigate, Fill, Extract, Verify), review states (Bindable, Normalize, Referral, Call-review), coverage deltas, proof attachment, operator notes, re-verification controls, and an Approve Bind decision point. Direct evidence of designing the boundary between agent preparation and human authorization in a regulated context.
Strongest match: Control and oversight (Construct 1). Your best proof object for any role where the primary screen is oversight, authorization, or human-agent control design. Confidence is high for the visible interface — the approval gate, review states, and exception routing are inspectable and directly relevant. The limitation: Carrier IQ still lacks a visible record of a wrong output being corrected and that correction changing a later run. It shows the gate architecture without showing the full correction loop.
Do not deploy for code authorship or eval infrastructure screens. Carrier IQ demonstrates design architecture, not a codebase or measurement system.
Brand Pulse
What it shows: Parallel source agents moving through visible execution states, source-labeled evidence attached to synthesized conclusions, an editable intelligence mission. Multi-agent visibility and evidence-beside-synthesis presentation.
What it does not show: An approval gate, a correction control, a human override that changes agent behavior, or confidence fields on individual conclusions. The user can cancel and rerun a mission, but those are execution controls, not authority over agent decisions.
Strongest match: Prototyping (Construct 4). Confidence is high — it's a working application that an evaluator can run, and it demonstrates agentic system design with visible agent orchestration. Useful as a secondary reference for oversight roles when paired with Carrier IQ, to show range across sensemaking and transactional contexts.
Do not present as equivalent evidence of human-override design. An evaluator who inspects it will find the gap, and that damages the credibility of your stronger evidence by association.
Retail Velocity
What it shows: Agent-collected market data translated into ranked operational priorities with categorical confidence labels, evidence snippets, and drill-down to source pages. A live activity stream exposes the full extraction pipeline.
Retail Velocity names TinyFish infrastructure (Search, Fetch, Agents) as its technical dependency, so you cannot present it as proof that you built the underlying agent stack. And it does not show human override or correction — the user inspects and chooses which account to pursue but doesn't change agent behavior.
Strongest match: Prototyping (Construct 4) and enterprise workflow (Construct 6). Confidence is high for prototyping — the full acquisition-to-ranking pipeline is visible and runnable. Moderate for enterprise workflow: it shows operational prioritization and field-intelligence design, but it doesn't establish enterprise-domain tenure in CPQ, CRM, or sales workflows specifically. Use the TinyFish dependency as context for your production AI experience, past tense.
Do not deploy for control and oversight roles where the evaluator needs to see approval authority. Retail Velocity shows inspection, not authorization.
Trust essay
Presents two frameworks — five handoffs (Intent-Setting through Loop Feedback) and a Watch-Verify-Delegate trust ladder — and argues that consistency is the design layer, not accuracy. Includes three first-person failure episodes from production AI work.
What it does not contain: interaction specifications, rejected alternatives, a product state map, an implemented correction flow, or the longitudinal record of bad output to diagnosis to design change to measured improvement. As Issue #8 identified, that temporal record remains the single largest gap in the AI evidence set. Confidence: high.
Deploy for oversight roles as a framing layer paired with Carrier IQ. The essay provides the vocabulary; Carrier IQ provides the craft. Also useful for enterprise-workflow roles where the evaluator values systems thinking about human-agent boundaries.
Do not deploy as standalone craft evidence. An evaluator screening for implemented design decisions will find frameworks and principles, not artifacts. Leading with the essay alone against a craft screen reads as conceptual rather than operational.
BCG Digital Ventures cases and Alibaba
Equinox+ (consumer 0-to-1 at speed), Thermo Fisher (regulatory QA release with a human gate before pharmaceutical batch shipment), and Red Cross (consequential workflow under operational pressure) — from your Product Design Director role at BCG Digital Ventures. Alibaba (enterprise platform trust and buyer-facing B2B redesign) — from your Head of Design and Research, North America role.
Deploy for enterprise workflow (Construct 6), where domain depth and delivery at consequence matter. Thermo Fisher and Red Cross directly demonstrate human authority in high-stakes automated workflows. Alibaba establishes enterprise platform trust at scale. Equinox+ serves the consumer-craft construct as evidence of 0-to-1 speed and product sensibility.
Do not deploy for roles where the primary screen is AI-native tool fluency or generative-product interaction paradigms. These cases predate the current AI toolchain.
Three gaps that routing cannot fix
Production-code authorship. The running Labs establish that you make working software. They do not expose a repository, commit history, code review process, or deployment ownership record. As Issue #9 covered, the process-visibility gap remains open. What would close it: publicly visible evidence of code-authorship process — something that connects the running application to the person who wrote it and shows how it was built. If a role's primary screen is personally authored production code, you do not currently have public proof that clears it. Confidence: high that this is a real gap.
Evaluation infrastructure ownership. The Trust essay discusses evals conceptually. Carrier IQ exposes traces and re-verification. Neither proves ownership of Python pipelines, graders, comparison sets, or model-launch release systems. What would close it: evidence of operational ownership of a measurement system — not discussion of why evals matter but a visible record of building and running the quality layer. If Anthropic's Evals & Prompts role is the target, this is an admission-level gap. Confidence: high.
Consumer generative craft at scale. Suno and Midjourney screen for taste, emotional resonance, and interaction-paradigm invention around generative media. Your current evidence set does not include generative-media product work or creative-tool design at consumer scale. The Trust essay and Labs address agent architecture, which is a different design problem than making a music-creation product feel right. What would close it: evidence of designing a generative-media product experience where emotional response and creative quality are the primary success criteria. Confidence: high that this is a distinct construct. Moderate on whether Suno treats it as a hard floor, since the posting lists AI-product experience as a plus.
These gaps don't mean avoid these roles. They mean knowing, before outreach, whether your strongest evidence addresses the actual screen — or whether you're relying on adjacent credibility to clear a construct you can't directly prove.
Pre-outreach matching lookup
Pull this up before any AI-company outreach. Read the posting. Figure out which question below it's really asking. Select evidence accordingly.
Does the posting center approval gates, override controls, or delegation boundaries? Lead with Carrier IQ's approval gate and review states. Pair with the Trust essay for vocabulary.
"The design problem I keep returning to is where human authority sits in an agentic workflow — not as a theoretical question, but as states, gates, and recovery paths. Carrier IQ is the clearest example: the operator reviews coverage deltas and evidence before an Approve Bind decision, and the system routes exceptions to human review rather than defaulting to automation."
Do not lead with Brand Pulse or Retail Velocity. They show inspection, not authorization.
Does the posting ask you to build working software, prototype with AI tools, or show running applications? Lead with all three Labs as working applications. Emphasize Retail Velocity's full extraction-to-ranking pipeline and Carrier IQ's stateful complexity.
"These are running applications I designed and built — Brand Pulse orchestrates parallel source agents for brand intelligence, Retail Velocity turns live menu extraction into ranked field priorities, Carrier IQ prepares insurance submissions through a multi-stage agent pipeline with human approval before binding."
Do not claim infrastructure authorship where TinyFish dependency is visible. Do not lead with the Trust essay — it's a written argument, not a built thing.
Does the posting emphasize enterprise domain depth, workflow automation, or translating business processes into agentic interactions? Lead with Retail Velocity for operational prioritization, then bridge to Thermo Fisher and Red Cross (BCG Digital Ventures) for consequential enterprise delivery.
"Retail Velocity translates agent-collected market data into ranked account priorities with confidence labels and evidence drill-down — the field rep sees why this restaurant matters before they walk in. That's the same design problem I worked on at BCG Digital Ventures with Thermo Fisher's QA release workflow, where the human gate before batch shipment was non-negotiable."
Do not lead with generic trust language without naming the specific workflow. The ServiceNow evaluator is screening for domain fluency.
Does the posting emphasize creative vision, emotional resonance, or inventing new interaction paradigms for generative products? Lead with Equinox+ (BCG Digital Ventures) for consumer 0-to-1 speed and product sensibility. Reference the Labs for current AI-tool fluency. You have a real gap here — own it cleanly if pressed.
"My consumer product work is in fitness and wellness — Equinox+ was a 0-to-1 launch at BCG Digital Ventures. What I'd bring to a generative product is the same instinct for how a product should feel, plus current AI building practice across three running applications. I haven't shipped a generative-media product."
Do not lead with the Trust essay or Carrier IQ. Agent oversight architecture is not what this evaluator is screening for.
Does the posting require a personal code repository, commit history, or ownership of evaluation pipelines and model-release systems? You have an evidence gap. The Labs demonstrate maker capability but do not demonstrate code-authorship process or eval-system ownership. Know this before the conversation, not during it.
"I build working applications — three are running publicly right now. What I haven't published is the development process behind them. I can walk you through how I work and what I built, but I don't have a public code-authorship record to point you to today."
If you pursue these roles, the conversation needs to establish what you can demonstrate — working applications, production AI context from your TinyFish work (past tense), systems thinking about evaluation — while being direct about what you haven't published. An evaluator who discovers the gap themselves will weight it more heavily than one you've already addressed.
Hiring committees that paused over Labor Day are reconvening this week. Roles that went quiet are cycling back into active review. The window between now and mid-October is when most of these AI roles move from posted to interviewing. Your evidence is strong for three of these six constructs. For the other three, know what you're walking into before you reach out.
- Carrier IQ's case page: The former case-page routes for Brand Pulse and Retail Velocity now return 404 responses, which means the only inspectable evidence for those two Labs is the running applications themselves — confirm whether Carrier IQ's case page has the same routing problem before relying on it in outreach.
- The correction-loop artifact: The single highest-value evidence build remains a complete failure-to-trace-to-change-to-rerun record, which Issue #8 named as the largest gap and which none of the three Labs currently demonstrates.
- Retail Velocity's TinyFish label: The application explicitly credits TinyFish Search, Fetch, and Agents on its interface, while Brand Pulse does not name TinyFish — that asymmetry changes which infrastructure-authorship claims are safe for each app.
- Anthropic Core Apps vacancy state: The direct job URL returned an error on September 12 even though a recent search capture retained the role text, so treat it as a recently expressed mandate rather than a confirmed open role until the listing is independently verified.

