Trust
Coca-Cola × TinyFish · Agentic Trust Design · 2026

An agent looked at 5,505
restaurants. How do you know
when to trust its call?

5,505
Real PA restaurants scanned
1,386
Flagged for human judgment
163
Design decisions explored on real data
3
Real stop-conditions now enforced
01
Chapter 01
The Real Product

Not a screenshot.
The actual app, walking itself through it.

Real evidence links, a real filter toggle, a real conflict — the shipped product demoing itself.

coke-intelligence.app/agent-portal ● Live product, auto-playing
Evidence → coverage → conflict, on a loop. Hover to pause. Nothing here writes to the real audit log — it's a passive walkthrough of the same real components and real data used throughout this build.
02
Chapter 02
The Object

One restaurant record.
Not an agent run, not a workstream.

Every screen in this build reads and writes the same object. Here's why that was the decision — and where its real edges are.

1/186
Real captured trace — Nucleus Raw Foods, a real browser-use.com session
185/186
Reconstructed trace — built from real final-state fields, never shown as literal logs
≤ 2
Runs, derived on demand — buildRunHistory(), not stored
real
Errors — causeClass, cause, recoverable
Real per-workstream QA depth behind this evidence: 100% qualification · 5% POS cross-validated · 7% menu.
writes into
Decision-ready Partial / Degraded Waiting / Conflicted Unknown Real nuance: two state pairs intentionally share one color — disambiguated by label, not hue.
RestaurantRecordthe only
top-level entity
FieldsClaim state — real distribution, 186-restaurant sampleReal policy
Qualification
status, evidence[], orderingSurface
100% decision-ready (186/186) — 0% unknown in this sample
Gate
POS
vendor, status, evidence[]
74.2% decision-ready25.3% partial0.5% conflicted
Escalate
Menu
beverages[], brand, price
98.9% decision-ready1.1% unknown
— none yet
Checkout
extractionStatus, upsell, bundle
17.2% waiting82.8% unknown
Recover
Identity
id, name, address, menuUrl, outputUrl
No claim derived — pure identity, nothing to classify.
not part of claims.ts
Stages
qualify → pos → menu, checkout, bundle, image
A meta-summary of the other rows, not itself classified.
not part of claims.ts
Failure
causeClass, recoverable
Feeds Checkout's claim state directly — not drawn as a separate edge.
read by Checkout, above
Human Review
humanReview[], promotedToEval
Downstream of Escalate, not a claim source — see HITL below.
not part of claims.ts
Declared, not wired to UI
POSCheckoutBundle

stage-dependencies.ts's one declared edge feeds exactly one line of copy in the Judgment Queue — it doesn't gate or hide anything today.

Proposed — not yet exercised
Has bundle?Bundle workstream

Same mechanism as the proven qualify→menu gate. Honest because it has to be: bundle/image have run for 0 of 186 real restaurants so far.

"Agent run" isn't a stored object here — buildRunHistory() derives it from a restaurant record's own fields, on demand, every time.

when escalated
humanReview[]
Historical, on the record — fieldPath, agentClaim, humanCorrection, promotedToEval.
ground-truth-store
This session's live annotations, browser-local — its own comment calls it "the app's single real decision/verification log."
"Written to this restaurant's audit log and added to the eval set... Future runs of this exact pattern won't reach this queue again." — AdjudicationView.tsx
Two distinct real logs, not yet unified into one — a human correction here doesn't (yet) rewrite the record above.
"

I would not add breadth now. I would redesign only three things deeply: Judgment Queue → Restaurant Intelligence Record → Temporal Recovery.

Real CPO design review, this pilot — CPO_REVIEW_TRACKING.md

The same map,
worked through with real evidence.

Every claim state above already exists in claims.ts, wired five times into the shared record component. Here's each row, with its real rule, its real policy, and the real restaurant behind it.

Real, functional — gate
Qualification: stoppedMenu: not run
Proven. A marketplace-only or static-PDF stop hard-short-circuits every downstream stage — real, in trace-template.ts.
Plaza Azteca Mexican Restaurant, Harrisburg — real trace graph showing Qualify complete and every downstream stage genuinely not run
Real, functional — recover
Store closedWaitingResume
Designed, not a bug. RecoveryTimeline.tsx's own comment: "waiting is a valid, designed agent state, not a failure with a retry button." Real hours, real next-open time.
Real recovery timeline for a closed restaurant, resuming from checkout at the next open hour
Real, functional — escalate
Disputed claimJudgment clusterHuman
Already shown, not asserted. A real judgmentReason routes into a real judgment cluster for adjudication — this is Vecchia Sicilia Pizza Shop's two disagreeing runs, seen live in the Chapter 01 demo and the converged screenshot below.
Honest gap — not gated yet
beverages: []Unknown
No real policy yet. Only qualification gates downstream work today. Plaza Azteca Mexican Restaurant · Harrisburg's own stop is what drives its empty menu, above — same restaurant, same real event.
Declared, not wired to UI
POSCheckoutBundle
Real, but inert. stage-dependencies.ts's one declared edge feeds exactly one line of copy in the Judgment Queue — it doesn't gate or hide anything today.
Proposed — not yet exercised
Has bundle?Bundle workstream
The next real extension. Same mechanism as the proven qualify→menu gate. Honest because it has to be: bundle/image have run for 0 of 186 real restaurants so far.
Rejected
AgentPortalRunDetail's own "Restaurant Truth"
+
DataPortalRunDetail's own "Restaurant Record"
Two independently-built representations of the same restaurant — free to drift apart, and they did.
Converged
The shared RestaurantIntelligenceRecord component, real evidence, real claim states
One shared component. Both portals render the exact same RestaurantIntelligenceRecord.tsx — the superseded summary component was deleted outright, not deprecated.
03
Chapter 03
Shipped

Every restaurant is a row.
Every claim links to its own evidence.

5,505 real Pennsylvania restaurants. Status words became clickable evidence; Qualification became a real, enforced rule. This is the actual shipped product.

● Live, interactive — not a screenshot
Real headline, full populationShipped
Dashboard home — real 1.85:1 competitive ratio, real brand-exclusivity quadrants
Real stop, real evidence, real trajectoryShipped
Plaza Azteca Mexican Restaurant, Harrisburg — real trace graph showing Qualify complete and every downstream stage genuinely not run, with the cited static-menu-PDF evidence
Real policy documentation, honesty preservedShipped
Agent Setups page — real stop-condition policy with real failure-mode rates and unresolved [CONFIRM] placeholders left honest
186 restaurants × 6 workstreams = 1,116 real cellsShipped
Restaurant x workstream matrix — 186 restaurants, 6 workstreams, blocked cells highlighted
Real backlog · one dot each
1,386
Qualification outcome · full 5,505 population
96.3% qualified 3.5% marketplace-stopped 0.2% static-menu-stopped
Brand exclusivity · Coke vs. Pepsi, real full population
1.85 Coke-only 1 Pepsi-only
Two real sessions disagreed — both shown, neither silently mergedShipped
Vecchia Sicilia Pizza real two-run conflict, both sessions cited side by side, resolve action
04
Chapter 04
What Looked Fine, Wasn't

The first version of "trustworthy"
was just confident-looking.

Three real bugs, shipped and reviewed before anyone caught them. Drag each slider.

Before: POS Platform shown as Toast · 94% confidence · verified, no link, no evidence
After: real evidence links, decision-ready status, and an explicit adjudication note when sources disagree
Before After
01 · A precise number measuring nothing. "94% confidence" was a flat binary — 94 or 22, nothing between, no way to check it. Drag to see the fix: a real URL and an explicit call-out when two real sources disagree.
Before: caption claims ChowNow ranks first with 2 real cases, while the only card on screen is Toast
After: caption honestly reflects the one real cluster shown
Before After
02 · A fabricated restaurant, standing in for a real one. The caption claimed "ChowNow" ranked first — the only real card on screen was Toast. Drag: the copy now derives live from the same data the card renders.
Before: raw truncated Python dict/JSON dump as the primary text
After: legible one-line summary, raw output kept visible as secondary evidence
Before After
03 · Real activity, rendered unreadable. Genuine agent output, dumped as raw truncated JSON. Drag: same real data, read aloud in one line — raw output kept, not hidden, just demoted.
05
Chapter 05
The Method

Reject a pattern on real data
before you get to ship one.

163 explorations, 6 threads, tested against the pilot's own real data until each one broke or held up.

RecordBehaviorHealth HITLErrorCoverage
Rejected
Rejected: Agent Health — per-agent tiles, one row per run
Agent Health — per-agent tiles told you which workstream was busy, never whether the restaurant record itself was trustworthy enough to act on.
Converged
Converged: real restaurant x field matrix, 7 restaurants x 5 fields
Intelligence Health — six real states per field, not one blended score. 35 real cells across 7 real restaurants, the actual disagreement highlighted in place.
Also Converged

Intervention-Role Taxonomy

None / Verify / Wait / Adjudicate / Teach / Approve — sorted by who should act, not by severity. Real split across this pilot's cohort: 4 in-the-loop, 0 on-the-loop, 3 out-of-the-loop.

Also Converged

Temporal State Model

READY → RUNNING → WAITING → RECOVERING → COMPLETE. A restaurant closed for the night becomes a valid, first-class "waiting" state — not a failure with a retry button.

06
Chapter 06
Principles, Distilled

What building trust into
agentic UX actually requires.

Real 'Direct ordering' evidence link on a POS Platform card
01

Evidence beats labels. A status word with no link is a claim with no way to check it.

Real Qualification card showing a policy exclusion with cited evidence
02

Every stop needs a real, cited reason. Not a silent success that quietly skipped a hard case.

Two real agent runs cited side by side, both shown as strong evidence
03

Real conflicts should look like conflicts. Merging disagreement into one confident answer destroys the signal.

Real captured session badge on an agent trace
04

Say what's real and reconstructed, on the artifact itself. Not in a README — right on the trace.

Restaurant matrix filtered to blocked cells, showing individual rows
05

A rollup hides exactly the restaurant that needs a human. Show the row, not just the average.

Go Deeper

The full interactive
build.

Explore the Canvas