The short version
You have one portfolio artifact that should be built before anything else: a published record showing an AI agent producing a bad output, your diagnosis of why, the design or system change you made, and verified behavior on a subsequent run.
This has been the top item in the Build Queue for five cycles. The specification keeps getting sharper while the published portfolio stays the same. That has a cost: every interview where you explain the correction sequence verbally instead of pointing to a record is an interview where the evaluator has to trust your narration rather than inspect your work.
Build it now. Here's why.
Why this is an evidence gap
The competency almost certainly exists. Your Agentic Labs work required diagnosing agent failures and changing system behavior across runs. But the published portfolio doesn't make that sequence inspectable. An evaluator who needs to see it can't see it.
Your published Labs — Brand Pulse, Retail Velocity, Carrier IQ — show staged execution, trace evidence, review states, coverage deltas, operator notes, and approval flows. Carrier IQ in particular demonstrates re-verification and evidence attachment. What none of them exposes is a wrong output leading to a rejection, a diagnosis, a design change, and a verified subsequent run. The competency is implied by the interaction surfaces that are visible, but implication isn't inspectable.
The earlier AI-native credibility dossier made a version of this argument. The Portfolio Playbook retrospective flagged that it hasn't shipped. I'm making the case again because the math keeps getting simpler: one artifact partially addresses three objections that currently require three separate verbal explanations. It doesn't close any of the three completely. But it reduces the verbal load on all three simultaneously, and that compression is what makes it the priority.
What "temporal design" means here
The term exists in design research — Pschetz and Bastian used it in 2018 to describe treating time itself as design material, and Wiberg and Stolterman extended it in 2021 into HCI. I'm using it in a narrower sense specific to agent systems:
Designing for the states, evidence, intervention rights, and behavior changes that unfold across sequential agent runs — not a single screen, not a single session.
The design problem that spans what happened last time, what went wrong, what changed, and whether the next run improved.
The closest named pattern is Zhou et al.'s Confirmation-Diagnosis-Correction-Redo sequence from CHI 2026, which studied how users supervise multi-step agents. But CDCR addresses intervention within a long-horizon task, not provenance across system versions or separate production runs. Anthropic's agent evaluation guidance gets closer — distinguishing traces, trials, outcome verification, and regression evals — but frames these as engineering practices rather than design proof.
The artifact you need sits in the gap between those adjacent frameworks.
Nobody is publishing this yet
I reviewed five representative AI-design cases published between 2024 and 2026 — two company-published (Gemini Enterprise agentic tools, Headspace Ebb) and three practitioner portfolios (Thomas Kelly's Agentic Studio, Alok Sharma's AI Agent Settings Framework, Mathieu Fortune's AI-accelerated design process). All five show interface states, flows, research iterations, monitoring surfaces, or governance structures. Several show explicit forms of human control and review.
None publishes the complete unit: failed output, inspectable evidence of what went wrong, a design or system change tied to that diagnosis, and verified behavior on a comparable later run.
Fortune's case comes closest — it connects diagnosed generation problems to changed inputs and visibly improved output. But it lacks predefined evaluation conditions, matched rerun criteria, or evidence that adjacent behavior didn't regress.
Five cases is a purposive sample. I can't claim nobody anywhere has published this. But the pattern is consistent enough to say: if you build and publish a complete correction-loop record, you will have something rare in the current portfolio landscape.
Why the shelf is empty. The absence is structural. Production agent failures are usually proprietary — companies don't publish records of their agents getting things wrong. The proof category itself hasn't been named in design practice, so designers aren't documenting the sequence even when they're doing the work. And portfolio conventions are still organized around screens, flows, and research artifacts — formats built for capturing interface states, not cross-run behavioral sequences. These conditions won't resolve in a quarter. The differentiation window is longer than a cycle, which makes the build investment worth it.
Three objections, one artifact, honest limits
For each: the bold text is the question an evaluator would actually ask. Then what this artifact makes visible, and what it leaves unresolved.
"Have you actually designed for model behavior, or only AI interfaces?"
The AI-native depth objection. Your current published work shows consequential delegation and control-surface thinking, as the earlier dossier established, but doesn't expose formal evaluation infrastructure or longitudinal correction evidence.
What the artifact resolves: It shows you reasoning about non-deterministic behavior — diagnosing why a specific output failed, deciding what to change, checking whether the change worked. That is designing for model behavior in a way an evaluator can inspect without taking your word for it.
What remains unresolved: Frontier-lab tenure. Ownership of foundation-model training or fine-tuning. Production adoption metrics at scale. Years of formal eval-infrastructure ownership. If the evaluator's actual filter is "has this person shaped model behavior at the parameter level inside a lab," one artifact from your own projects won't clear that bar. The artifact addresses the design-reasoning dimension. The organizational-context dimension needs other evidence.
"Are you still hands-on enough to make difficult interaction decisions yourself?"
The craft objection. The IC+Manager scorecard analysis identified that evaluators at this level probe whether a candidate still makes design decisions or only directs others who do.
What the artifact resolves: It exposes specific interaction decisions you personally made — which states to surface, what evidence to show the operator, how to structure the correction interface, what the re-verification experience looks like. Design decisions with visible consequences.
What remains unresolved: Team leadership. Repeated craft direction across a product portfolio. Production visual polish at scale. How your judgment propagates through other designers. The artifact proves individual craft judgment. The management and scaling dimensions need other evidence.
"What did you actually build, test, and change?"
The technical process objection. OpenAI's recruiting guidance emphasizes ownership, builder execution, and how you reached an answer. Anthropic asks for direct evidence — independent research, writing, or open-source work.
What the artifact resolves: A correction-loop record shows trace reading, failure classification, a versioned intervention, evaluation conditions, a rerun, and an outcome. It answers "what did you change and how did you know it worked" with inspectable evidence.
What remains unresolved: Repository history. Production-code authorship. Prompt and eval-set provenance. Deployment ownership. Monitoring infrastructure. Regression safety beyond the demonstrated case. Anthropic's own guidance on regression evaluation — testing whether previously working behavior still works after a change — is the standard your artifact should acknowledge even if it can't fully demonstrate at portfolio scale. One successful later run is evidence about that run. It is not proof of generalized improvement. The artifact should be honest about this scope. Showing what you checked and what you didn't is itself a credibility signal.
Confidence boundary
I looked for published evaluation rubrics, interview guides, or recruiter commentary from OpenAI, Anthropic, and Headway that explicitly name failure-recovery reasoning, correction sequences, or multi-run behavioral design as design-candidate criteria. I did not find any. The public materials test adjacent constructs — problem-solving reasoning, ownership, impact, technical ability — but don't expose a correction-loop scoring rule.
The artifact's hiring relevance is an inference from current role scope and AI operating practice, not a confirmed evaluation filter. I'm confident in the inference — these companies build agents, agents fail, someone has to design the correction experience. But the confidence label matters. You can say this artifact addresses capabilities these employers value. You should not say it addresses a known evaluation criterion.
Verbal bridge — use until the artifact publishes
For any conversation where these objections surface:
"I haven't worked inside a frontier lab. My AI work is grounded in how humans supervise, verify, correct, and delegate to agentic systems. The published Labs show the interaction and control layer. The full failure-to-correction-to-next-run record is what I'm building now — I can walk you through the sequence, but the published version isn't live yet."
If the evaluator probes specifically on the technical-builder dimension:
"The Labs are live interfaces I designed and built. What they don't show publicly is the iteration trail — the failed outputs, the diagnosis, the versioned changes, the rerun evidence. I have that history; it's not yet in portfolio form. I'd rather show you the real sequence than a polished summary that skips the failures."
Two things to watch in delivery. First, say the gap once and move to the capability. The spoken-answer discipline from Issue #7 applies: diagnose whether the evaluator doubts competence or integrity, use the smallest defensible claim, spend the answer on the target capability, stop. Second, if a second probe comes on the same objection, switch evidence carriers. Don't repeat the same verbal explanation with different words. Move to a specific failure, a specific decision, a specific outcome. The verbal bridge opens the conversation, but the specificity of your follow-up is what sustains it.
What to build
The artifact specification, refined across five Build Queue cycles:
- A specific agent output that failed, with the evidence visible
- The diagnosis — what went wrong and how you identified it
- The design or system intervention, with enough detail that an evaluator can see the decision, not just the outcome
- Predefined evaluation conditions for the next run — not post-hoc selection of a run that happened to look good
- The next-run result under those conditions
- An honest note on scope — what you checked, what you didn't, what would need further verification
Items 4 and 6 are what separate this from a before/after screenshot. They're also what make it credible to an evaluator who understands how agent systems actually behave. Anthropic's guidance on regression evaluation is the standard your artifact should acknowledge even if it can't fully demonstrate at portfolio scale.
The specification is ready. Ship it.
Pre-interview reference
Pull this section before any conversation where AI-native depth, hands-on craft, or technical process objections might surface.
Three questions you may hear and what to lead with:
- "Have you designed for model behavior, or only AI interfaces?" → Lead with the Labs' control-surface and delegation design. Concede no frontier-lab tenure. Bridge to the correction sequence you can walk through verbally.
- "Are you still hands-on?" → Name specific interaction decisions you made in the Labs — states, evidence surfaces, re-verification flows. These are your decisions, not a team's.
- "What did you build, test, and change?" → The Labs are live interfaces you designed and built. The iteration trail (failures, diagnosis, versioned changes, rerun evidence) exists but isn't yet published. Offer to walk through a specific instance.
Verbal bridge (memorize the structure, not the words):
Acknowledge the gap → claim the decision layer → offer the specific sequence → stop.
"I haven't worked inside a frontier lab. My AI work is grounded in how humans supervise, verify, correct, and delegate to agentic systems. The published Labs show the interaction and control layer. The full failure-to-correction-to-next-run record is what I'm building now."
Artifact spec (the build target):
- Failed agent output, evidence visible
- Diagnosis of what went wrong
- Design or system intervention, decision visible
- Predefined evaluation conditions for next run
- Next-run result under those conditions
- Scope note — what was checked, what wasn't
- CHI 2026 confirmation patterns: Zhou et al.'s study found that 81% of participants preferred intermediate confirmations over confirm-at-end when supervising multi-step agents, which has direct implications for where you place correction affordances in the artifact's interaction sequence.
- Anthropic's regression eval standard: Their engineering guidance explicitly warns that capability tests should be paired with regression tests to catch breakage elsewhere after a fix, a principle your artifact's scope note (item 6) should reference honestly.
- NIST monitoring gaps persist: A March 2026 report found that deployed-AI monitoring remains fragmented, with unresolved questions around human-AI feedback loops and drift — which means the correction-loop proof category you're building toward addresses an industry-wide gap, not just a portfolio gap.
- OpenAI's player-coach model: Their Growth design leadership role explicitly combines direct design work with management of a small team, which means the artifact's proof of individual craft judgment lands differently there than at a company screening for pure management.

