An agent looked at 5,505
restaurants. How do you know
when to trust its call?
Not a screenshot.
The actual app, walking itself through it.
Real evidence links, a real filter toggle, a real conflict — the shipped product demoing itself.
One restaurant record.
Not an agent run, not a workstream.
Every screen in this build reads and writes the same object. Here's why that was the decision — and where its real edges are.
top-level entity
stage-dependencies.ts's one declared edge feeds exactly one line of copy in the Judgment Queue — it doesn't gate or hide anything today.
Same mechanism as the proven qualify→menu gate. Honest because it has to be: bundle/image have run for 0 of 186 real restaurants so far.
"Agent run" isn't a stored object here — buildRunHistory() derives it from a restaurant record's own fields, on demand, every time.
fieldPath, agentClaim, humanCorrection, promotedToEval.I would not add breadth now. I would redesign only three things deeply: Judgment Queue → Restaurant Intelligence Record → Temporal Recovery.
Real CPO design review, this pilot — CPO_REVIEW_TRACKING.mdThe same map,
worked through with real evidence.
Every claim state above already exists in claims.ts, wired five times into the shared record component. Here's each row, with its real rule, its real policy, and the real restaurant behind it.
trace-template.ts.
RecoveryTimeline.tsx's own comment: "waiting is a valid, designed agent state, not a failure with a retry button." Real hours, real next-open time.
judgmentReason routes into a real judgment cluster for adjudication — this is Vecchia Sicilia Pizza Shop's two disagreeing runs, seen live in the Chapter 01 demo and the converged screenshot below.stage-dependencies.ts's one declared edge feeds exactly one line of copy in the Judgment Queue — it doesn't gate or hide anything today.
RestaurantIntelligenceRecord.tsx — the superseded summary component was deleted outright, not deprecated.Every restaurant is a row.
Every claim links to its own evidence.
5,505 real Pennsylvania restaurants. Status words became clickable evidence; Qualification became a real, enforced rule. This is the actual shipped product.
The first version of "trustworthy"
was just confident-looking.
Three real bugs, shipped and reviewed before anyone caught them. Drag each slider.



Reject a pattern on real data
before you get to ship one.
163 explorations, 6 threads, tested against the pilot's own real data until each one broke or held up.
Intervention-Role Taxonomy
None / Verify / Wait / Adjudicate / Teach / Approve — sorted by who should act, not by severity. Real split across this pilot's cohort: 4 in-the-loop, 0 on-the-loop, 3 out-of-the-loop.
Temporal State Model
READY → RUNNING → WAITING → RECOVERING → COMPLETE. A restaurant closed for the night becomes a valid, first-class "waiting" state — not a failure with a retry button.
What building trust into
agentic UX actually requires.

Evidence beats labels. A status word with no link is a claim with no way to check it.

Every stop needs a real, cited reason. Not a silent success that quietly skipped a hard case.

Real conflicts should look like conflicts. Merging disagreement into one confident answer destroys the signal.

Say what's real and reconstructed, on the artifact itself. Not in a README — right on the trace.

A rollup hides exactly the restaurant that needs a human. Show the row, not just the average.