Skip to content

The substitution error: a census of 135 retros (S264 → S527)

The substitution error: a census of 135 retros (S264 → S527)

Section titled “The substitution error: a census of 135 retros (S264 → S527)”

Error class under study. Substituting an artefact’s existence, prevalence, or provenance for evidence that the artefact is CORRECT.

Coverage. All 135 retro files read in full — S264.md through S527.md, 9,462 content lines / 838 KB, concatenated and read end to end. No sampling, no truncation. The MemPalace secondary sweep was not run; the retro sweep consumed the budget and returned enough signal that the secondary would have been marginal. Every row below is traceable to a quoted line in a named retro.

Headline count: 161 occurrences across 118 of the 135 retros (87%). Seventeen retros carry none — nearly all of them short, single-lane, or _none_-heavy.


The five observed variants all recur, heavily. The corpus forces three additions and one major generalisation.

#VariantCountStatus
V1Consumer-counting24confirmed
V2Population-as-evidence25confirmed
V3Task-directive-as-authority34confirmed
V4Stamp-as-warrant51confirmed — and far wider than docs
V5Absence-as-proof31confirmed
V6Landed-as-live12NEW
V7Signal-as-state11NEW
V8Authored-therefore-applied7NEW — and the key to why naming fails

V4 generalises far past documents. DR-106 framed V4 as a docs problem. The corpus says otherwise: the same shape appears as a Checker PASS, a green test suite, a CI gate, a --check run, a ratification label, an audit finding, an agent attestation, and a generated tombstone. In every case a verification artefact certified only what it actually executed, and was read as certifying the claim. This is the largest bucket in the corpus by a wide margin.

V6 — Landed-as-live. Existence in source / spec / ledger substituted for existence in the running system. Distinct from V1 because nothing is being counted; the claim is about reachability, and the artefact is real and correct — it just is not the thing that runs.

V7 — Signal-as-state. A status, health, completion or exit signal substituted for the state it purports to report. Distinct from V2 because the signal is designed to report the state; it reports a proxy.

V8 — Authored-therefore-applied. The agent’s own recent output — a comment it wrote, a rule it quoted into a brief, a fix it just reviewed, a correction it just issued — treated as external evidence. This is the smallest bucket and the most important one; see Part 6, RC6.


Columns: Session · Claim · Substitute · Caught? · What caught it · Cost. Caught codes: BEFORE (caught before acting) · AFTER (work shipped on it, then caught) · PROD (reached a live surface) · AMBIG.

SClaimSubstituteCaughtWhat caught itCost
S385”F3 (token-naming) churn at ~87 files”file count of var(--color-*) usersBEFOREdeterministic Python extraction — “measure blast radius before quoting it”none; a wrong gate quote to the owner
S393dedup fold has “2 importers”grep from '@/lib/X'BEFOREtypecheck; real count 16 (5 dynamic imports + 9 vi.mock)none
S398”{50.12} = 146 routes”count of TODO(OPS-T1) markersBEFORErecon workflow re-baselined pre-fan-out; real = 86, “~73 were stale comments above ALREADY-bound schemas”none — “value of grounding scope before orchestrating”
S406S3 template-coverage recon completedir-scoped grepAFTER”Dir-scoped recon greps miss consumers (types/unified-gap.ts slipped)“rename shipped incomplete
S409buildBidSummary blast radius HIGHgitnexus transitive impactBEFOREmaxDepth:1 re-query → LOW, 2 callers; “HIGH was process-step aggregation”avoided a wrong escalation
S411Unit-D wire-field map completesingle Explore agent’s consumer listAFTERChecker full-suite (2 fails) + orchestrator; “the map traced only one” of two test consumers2 broken consumers, 1 silent ?? []
S438touch-point inventory ~9, then ~22prose inventoryBEFOREChecker sweep-grep found T23; “Sweep-greps, not prose inventories, are the ground truth”2 wrong inventories
S443gitnexus caller counts completetool outputAMBIGsubo executors grep-verified every set; “Treat gitnexus caller counts as a floor”repeated re-work
S443find_related_items DROP — “no surviving caller”caller sweep, owner-ratified on itAFTERS443 types regen surfaced 3 tsc errors; survivor was the legacy IMS item pageratified DROP on a false premise
S444govfacet_b migration’s “disposition list” (5 objects)migration author’s declared footprintBEFOREfull-corpus grep found 5 more live readers incl. flow.pywould have broken at GO#2 apply
S444GO#2 gate = “~8 TS consumers + 3 callers”continuation promptBEFOREexecutor traced the actual table; real gate included SQL fn bodies, a Python INSERT, an api.* viewnear-miss on a breaking rename
S455{127.24} content_items impact LOW / 0 callersGitNexus verdictBEFOREexecutor escalated; real radius = 4 signatures + ~15-20 pytest files, 15,211 linesnone
S457”keep use-application-types.ts — shared with the API route”filename-substring grep matching a prose commentBEFOREChecker’s import-path re-verification, round 3would have shipped stranded dead code
S466knip: form_extractors is orphaneddead-code gate + a false TECH.md claim that a second writer existedPRODlater archaeology; it was “the sole form_template_fields writer, 2,235 LOC”2,235 LOC of load-bearing code deleted
S472memtrace get_impact can serve impact-before-edit18 reported callers for getAuthorisedClientBEFOREGitNexus cross-check: CRITICAL / 151 production routes→ DR-071
S475gitnexus impact usable in agent worktrees0-direct / not-found verdictsBEFOREgrep ground truth, independently by 3 executors + a curatorwasted verdicts
S484”knip output enumerates the dead UI”knip counts __tests__ as consumersBEFOREproduction-importer-zero sweep found 5 more orphans knip structurally cannot flag5 built-not-wired features hidden
S498”six DR ids were swept unrecorded”heading grep of the registerAFTERgit log -S audit: it was 34; 610 live citation sites across 40 retired idsDR-087 re-issued; DR-090 burnt
S499”citation counts are facts”per-line grepAFTERexact-match per-id recount; 250 → 193 lines → 21 ids, “each shift a counting-method change, not drift”3 wrong figures in the ledger
S507”~80+ columns are built-but-unwired”ast-dataflow censusBEFOREaudit deflated it: 135 are the deliberate PG-default convention; the tool is blind to declarative TableSchema writes and PG-function readsavoided a mass-retire
S515record-run.ts is live — 18 callerscaller countBEFORE(owner)owner challenge; only one was migrated, and the S507 audit had already diagnosed the module as pre-cocoindex vintage1 of 4 overturned verdicts → DR-104
S522id-412 AC says 17 consumers, DR-117 says 24two different counting methods, neither declaredAFTERretro-miner; “nothing says so”two unreconciled counts live in the ledger
S524every count in the task fileprior sessions’ measurementsBEFOREthe AC’s own “re-measure first”; 17→16, 24→26, 99→87 — “16 at 63602df9, S522’s own measurement commit, so it was never right”one wrong figure repeated to the owner
S525restore the deleted schemas — “2–3 referencing files”grep -rl | wc -lBEFORE(owner)owner: “is it the case that the live files still referencing these are accurate and correct?” Every one was a test; two referenced the names as string literals in codemod fixtures”That is DR-104’s trap verbatim — which I had written into the taxonomy-spike brief as a binding rule in this same session”
SClaimSubstituteCaughtWhat caught itCost
S363→S366”GO on mxbai embedder”recall@k probe over stagingAFTERS366: “a staging-confound artifact” — 527/546 rows synthetic, one 154k-char gold docan architecture decision reversed
S416win-rate parity test proves the rewritepre == postBEFOREfidelity review: “VACUOUS against live data — all 12 workspaces have NULL outcome; the join returns 0 rows, so pre==post trivially”the sole gate on a CRITICAL rewrite proved nothing
S420staging is data-sparse (0 procurement ws)S418 diary figureBEFORElive query: 422, churning to 429+ mid-sessionwrong session plan
S421CV rename is “FK-safe: 0 rows reference pqq”row count on one column, on sparse stagingAFTERprod apply failed 23503; form_template_requirements had 66 pqq rowsprod migration failure (rolled back clean)
S421win-rate parity suite green”3 passed” in 2 msBEFOREsuspicious speed → 3 no-op skips; real run 1.2 s. “Treat a sub-100ms ‘pass’ on a DB-touching suite as a skip signal”vacuous gate
S427”zero data lock-in / fully reversible”content_items=33, q_a_pairs=0 on the Platform DBBEFORE(owner)owner; client-prod held 631 content_items / 926 chunks / 55 thumbnailsa reversibility ruling on the wrong DB
S438raw-pool sd write worksrow-COUNT assertions greenBEFOREChecker; bare MagicMock auto-configures __aenter__, so await conn.execute() never raises and the write silently no-opsvacuous test
S448governance review/action works0-row UPDATE returning successPRODS447 subo’s B3 checker; “every gov-facet reader/writer 0-rows, silently false-succeeding” — nothing mints the facet rowslive silent no-op
S456three tests green[] == [] under wrong labels after a positional shiftAFTERa signature change surfaced them; “Tests were green without exercising real behaviour”3 vacuous tests across 2 executors
S460/api/search date filters workfully green suite, ISO-Z-only fixturesBEFOREadversarial checker probed the pinned runtime library; <input type=date> emits bare YYYY-MM-DD → 400 on every searchnear-miss on a user-facing 400
S481the api-views migration appliedsupabase migration list local == remoteAFTERcatalog check (to_regclass); a cross-lane stamp collision meant the file was silently SKIPPED2 smoke runs burned → DR-081
S481{145.47} citations fix correct136 green tests on injected fixturesBEFOREchecker; SS-D1 was “structurally unreachable in production” (wrong FK axis)re-scope
S483nightly classification eval is measuring somethingit ran and produced scoresPRODS483 audit; the 91-item gold fixture references pre-id-131 UUIDs that “resolve to NOTHING on staging”weeks of eval comparing against zero rows
S483mcp-eval cleanup workstry/catch “best effort”PRODS483; deletes hit DROPPED tables, error swallowed”leaked every eval-created row into staging for weeks”
S494”notifications have never been created”one environment’s rowsBEFOREverification split (3 confirm / 1 qualify / 2 agent-corrects-Coordinator); id-143 had recorded 124 notifications on prod”distrust universal negatives”
S506”Run #64 is the e2e baseline”a run existedBEFOREe2e-triage agent: it was a scoped workflow_dispatch, 5 tests/shard; re-baselined on #63PR #139 wrongly implicated
S507”Run #37’s zero give-ups proves 1.0.18”a zero countBEFOREfence-state check: “vacuous — the walk never ran. A proof run’s denominator must be verified before its numerator means anything”avoided a false upgrade verdict
S514rebuilt test reads verify the merge-route pingreen assertions dropping the error channelAFTERpost-hoc review; “failed queries pass vacuously — including the merge-route pin stamp, the exact failure Inv-9 exists to close”Major discipline defect across 4+ files
S515column is live — “a row exists on staging”row populationBEFORE(owner)owner; “both Platform DBs are internal dev, pre-launch”re-opened a retire verdict → DR-104
S516the length-floor guard is testeda passing testBEFOREmutation testing; the test set PIPELINE_CLIENT_ORG="x" and asserted on a string containing no x — removing the floor changed nothingvacuous guard on a credential-leak path
S516_dsn_password_variants is coveredsuite greenBEFOREmutation: stubbing it to return set() “failed no test at all” — despite closing 5 of 14 adversarial caseszero coverage on the load-bearing function
S520the rooms mine worked9 rooms populated, 0 general, divergence 0BEFOREGROUP BY room: room=decisions was 82% not-decisions. “‘The field is populated’ is the same trap as DR-104’s ‘there are rows in the DB‘“a discriminator “worse than useless”
S522IP-leak sweep cleanexpects 8 files, finds 0, reports passPRODhuman reading the file; the filter points at a deleted directory, on a public repoID-68 PC-40 identity sweep vacuous
S523consumer census complete (twice), then a thirdgit grep returning 0BEFOREexpected-count assertion; 3 mechanisms — :! pathspec fatal with stderr suppressed, untracked files, a persisted cd”S522 named this class. Naming it did not prevent it.”
S524URL landing path is verifieda test inside the nightly’s file selectionPRODS524; describe.skipIf on a manually-set env var — “proven once, by hand, at S319, and never since”~5 months of no coverage
SClaimSubstituteCaughtWhat caught itCost
S357executor file-ownership boundariesthe PLAN’s narrative proseBEFORE (×3)executor escalation each time; “derive executor boundaries from the actual tool→file registration map”zero wasted edits — the escalation contract worked
S359bl-322 premise: classify.py loads the prompt at runtimebacklog entryBEFORE{114.1} RESEARCH + orchestrator verify; classify.py is goneprevented duplicate PRODUCT scope
S373GH secrets/CI artifacts out of ID-68 scopethe spec’s stated scopeBEFOREthe flip itself surfaced a new leak vectorflip held
S424bl-165 is ready/half-donethe backlog ledger recordBEFOREempirical backlog triage w/ a grounding doc; both halves already shipped as ID-61.4”the grounding-doc-first pattern stopped agents taking stale records at face value”
S429two ratified-spec premisesTECH M3 / the briefBEFOREexecutors verifying before edit; both factually wrong vs the code”S426 specs authored fast in one night → factual drift”
S434”drop the forms route (forms = manual-upload, okf-v3 §8.2)“the re-spec premiseBEFOREthe removal executor proved forms are corpus-walked on main today→ DR-014, BL-392
S437{131.11}/{131.21} details are currentsubtask details proseBEFORET4 brief made §9 governing; “an executor following details verbatim would have built owner-rejected surfaces”near-miss
S438ratified decisions reach sibling recordsthe ledgerAFTER”Back-propagation defect class recurred FOUR times” — incl. the Checker catching the audit itself missing the fourth4 stale records
S443”re-point onto source_documents.rel_pathan owner ruling’s stated targetBEFOREre-scope Planner + Checker vs DDL: no such column”Owner-facing option labels must be column-verified before boards go out”
S448ID-131 endgame “fully shipped”the S447 handoff + mirror statusBEFOREgit log/origin-main; only ~half integrated, nothing on main, subo dead”trust git-log over a died-session’s self-report or mirror status fields”
S448migrations “authored, NOT applied”mirror journalsBEFOREDB check; staging was already ahead through 20260704130000near-miss on a coordinated GO
S449route deletablea migration-prose citationAFTERChecker; the POST leg had a live /library caller, and “the hook’s fetch-mocked unit test kept the suite green through the 404”route deleted, caller broken
S453”Prod-apply … completed”a task-list titleBEFOREorchestrator; the session had applied nothing (verify-only)false causal fact written into bl-423
S455brief’s 3-file list defines completionthe briefBEFOREsub-o self-identified; the testStrategy said grep-clean e2e/** and a 4th file blocked the goal”testStrategy scope, not the brief’s file list, is the acceptance surface”
S457bl-414’s 5 owner questions are openbacklog statusBEFORE(owner)owner’s instinct that S456 already answered them — correct; 4 of 5 mapped to DRsitem deleted
S475”status enum enforced server-side” (TECH INV-3)the specBEFOREexecutor grep of the write path: zero references at v0.10.1→ minted {148.13}
S475”executors follow commit-per-slice briefs”the brief lineAFTER~985K tokens, 84 dirty files, zero commits. “The brief-carried convention is not a mechanism”orchestrator rescue-commit
S477product-retros.json has no mirrors”the 156.1 brief premiseBEFOREretro mirrors had shipped under ID-148.10/12”Brief premises need a freshness check at dispatch”
S484the journalled close-gate inventory is completethe journalBEFORE (×2)executors deviated correctly; 4 dead exports not 3, and a direct importer missing from the file listbrief-composition gap
S488dispatch-brief line citationsbundle_writer.py:998, “2197 pytest baseline”BEFOREexecutors re-derived from source; real guard at :1250, baseline 2199”briefs should cite symbols, not line numbers”
S498{163.20} is doneledger [x] + a recorded decisionAFTER, near-total losspre-reset diff-audit; the code “existed nowhere in git”, surviving only as uncommitted changes in an Intent workspace, “one git checkout . from loss”recovered as 8aa52164
S499{163.19} is doneledger [x], no landing SHAAFTERthe same pre-reset audit; “Two orphans from one workspace, both hidden by a done-checkbox”recovered
S500{128.21} is pending/unadjudicatedthe task file’s own lineAFTER”shipped-but-half-rolled-out for 11 days… the task file’s own line hid the incompleteness — two sessions carried a stale carry”11 days
S501workflow RULES block baseline = 263a pinned figureAFTERwent stale inside the same workflow (308 by wave 4)wrong verdicts
S512the inv-inventory brief’s premisebrief-encoded naming assumptionBEFORE(luck)an API storm prevented the dispatch; the inline sweep found Inv-N is a per-spec namespace”Grounding briefs on untested naming assumptions produces confidently wrong artefacts”
S513”~46 Inv-citing files … verified, not an assumption”an inherited ledger count promoted to verifiedBEFOREthe grep in the brief’s own step 1; actual 122 files / 1,262 cites, 2.6× wrong; “id-80 owns no register at all""Never promote an inherited figure to ‘verified‘“
S519”It is ratified, so it is executable”owner ratification 8-of-9BEFOREprojecting the mine found 2 config bugs + 1 design gap in one pass, “none visible by reading”
S519DR-112 Q4 (move 01-vision.md) is soundthe ratificationBEFOREinitiatives matches four levels earlier; “Ratified rulings inherit their premises’ defects”would have changed a published URL for no effect
S520”The owner approved the move, so execute the move”approvalBEFOREid-386 had already ruled KEEP-IN-PLACE with evidence (8 siblings citing by bare filename)“The approval was real; the premise under it was not”
S522id-401 “NOT on main”; id-407 “no PR opened”status_noteAFTERgit cat-file -e main:<path> — both merged two sessions earlier”Reading a status_note is not reading state”
S522id-404 “stays core-product per the S522 owner rulingan unticked multiSelect boxAFTERretro-miner; the owner never mentioned id-404”attributing to a ruling what was silence on a checkbox”
S523id-416 “Owner ruled at S523 that this is a separate task”the owner’s “See previous response” (a deferral)AFTERself-corrected at session close. “S522 recorded this exact shape and naming it was not enough to stop it recurring”second occurrence in two sessions
S523nine ACs metone replace("- [ ] ", "- [x] ")AFTERAC-9 was untouched, leaving two claims the session had falsified standing in the repo. “An AC checkbox is a claim, and a bulk edit cannot make nine claims”live falsehood on disk
S525AC-11’s premise (keep live-verify.sh as a thin wrapper)the ACAFTERreading the file end to end: it “could not run at all”, collapsing the choicetwo options presented to the owner against a false premise
SClaimSubstituteCaughtWhat caught itCost
S348{101.9} parity RED, 8 write targetsa snapshot docBEFOREID-45 RESEARCH vs live flow.py: 9 targets, {101.9} closed”verify the forward map against live code, not the snapshot”
S354reference_ingest not on proda migration-file comment (“GATED operator step, NOT applied”)BEFOREpg_proc on both prod projects — it was LIVE”verify prod STATE, not migration-file annotations”
S358narrative forward-docs reflect current statedoc currencyBEFORElive checks overturned doc claims repeatedly in one session; the owner caught several→ forward-doc freshness protocol
S363cocoindex 1.0.3the briefBEFOREinstalled and pinned are 1.0.7
S370”prod already half-flipped, api schema empty”three independent docs agreeingBEFOREexecute_sql + Management API: api fully applied on all 3 remotesmis-scoped {115.13} as risky live DDL
S370§7 REVOKE gotcha is load-bearinga diary entry re-asserted in the S369 promptBEFOREalready superseded by S368’s measured finding; “a stale gotcha carried forward across sessions”
S388”Decisions A–D are in the saved output”the saved session output existingBEFOREgrepping raw transcripts; the dump was “a lossy recursive summary that REFERENCED A-D without enumerating them”forensic detour
S390subtask introduces no regressionper-subtask Checker green on touched modulesAFTERthe wrap full-suite gate; a transitively-dependent suite broke2 regressions
S4006 branches merge cleanpairwise git merge-treeBEFORE”Pairwise-clean did NOT mean sequentially-clean” — a cross-stream guard ratchet was breachedconsolidation re-plan
S401platform ref-data populated (15/57/4/24)the S400 continuation promptBEFOREit was 0/0/0/0; carried from S399 without re-verifying”never trust a carried-forward ‘done’ for ephemeral DB state”
S405”#68 already merged / main HEAD cfc4fff3a sub-agent’s factual claimBEFOREorchestrator verified; that was a feature-branch HEAD, main was 600d44de3 doc refs corrected
S406bid_questions/bid_responses retain names”an S403 recon claimAFTER5-agent adversarial verification incl. live prod DB. The claim had propagated into the S403 continuation prompt AND two ledger journals2 subtasks blocked on a phantom gate
S410the squash baseline carries the DBa migration baselineAFTER”SPINE OF THE SESSION: schema ONLY — not core DATA rows, not function ACLs… This single fidelity-gap class bit THREE times this session”0-row application_types; anon EXECUTE on 72 RPCs
S416procurement-workspaces spec [CURRENT-CANONICAL] RATIFIED-S242the ratification labelBEFORE17-agent grounding workflow; the ratification predated both the ID-61 rename and the S391 glossary, and the satellite table is empty”Treated as a straw man to interrogate, not authority”
S421fresh reset is provendb reset exit 0 (measured before the regen) + --checkAFTERa real fresh reset: 42P01. --check is NOT a fresh-reset proof — it cannot catch migration ORDERING bugs”latent on-branch
S424the ratified G0 compose/recipe is deployableratificationBEFOREimage probe found four latent defects incl. “NO v1.38.0 CONTAINER tag exists""tool-probe ANY stock image”
S438the newer artefact is the current oneauthorship dateAFTER”Recency of ratification, not of authorship, is the staleness signal”inverted staleness judgement
S444SQL fn rewrite is correctCREATE OR REPLACE succeededBEFOREempirical scratch-Postgres: an ambiguous column ref (errors only on first CALL) + a COUNT(*) fan-out”validate against real data, not ‘CREATE succeeded‘“
S445DR-030 covers api-surface stalenessthe DR existingAFTER”api.-surface staleness fired 3x in one GO#2 despite DR-030 existing… the guardrail is docs-only — no automated check exists”*3 live breaks, 1 blocking runtime
S451Task-131 close: “zero live production reads of any M6-dropped relation”a residue-sweep PASSPROD P0get_guide_content 500’d prod; Postgres does not dependency-track LANGUAGE sql bodies/guide/[slug] down since the M6 GO
S453fixtures are byte-verifiedtwo independent executor attestationsBEFOREchecker re-ran the real writer: 53/54. “Attestation is not verification… the pre-commit hook invalidated them after the fact”
S453”ID-134 D10 defers the Platform OKF repo”a TECH.md coordination noteBEFOREDR-027/R6 + an owner ruling supersede it. “Coordination notes in spec prose are point-in-time”2 journals corrected
S456Sentry causes the /login pageerrorthe S455 diagnosisBEFOREcontrolled A/B rebuild (DSN unset, 23/23 repro) disproved it; real cause @vercel/analyticsavoided a wasted fix cycle
S457”gitnexus detect_changes is unrunnable in worktrees”standing guidance in briefs and skillsBEFOREit ran successfully twice this sessiondegraded discipline for N sessions
S459”targeted memo invalidation is infeasible”the {66.14} rulingBEFOREthe pinned 1.0.7 documents a memo-state API; the ruling predates cocoindex v1(then S460 re-falsified the replacement hope by reading source)
S459the gov.uk fixture URL is stableS415’s “repoint to a verified-live page”AFTERit rotted again by S458. “live-URL seeds are structurally fragile”recurring nightly red
S472the x509 fix is in effectit was committedAFTER, ~30 sessionsgrepping the installed binary’s own schema: the key was nested wrong and “the settings validator drops misplaced keys silently”30 sessions of gh/gh-axi TLS friction
S475vendored lib/ledger matches the pinned taga drift workflow existingAFTER”FALSE for ~5 releases… the drift workflow watches schema assets only and is non-blocking, so its standing warning was ignored”a 5-release delta absorbed at once
S476”CI failures are pre-existing, safe to merge”the labelBEFOREdiagnosing the reds found a real W1 search regression; reframed as “real-but-already-broken-on-staging, no live users”would have shipped broken search to prod
S477”the 12 test failures are pre-existing on main”an executor claim verified by reverting its own fileBEFOREthey began the moment v0.12.0 was pushed. “A revert-check scoped to your own commit cannot clear environmental causes”misread root cause
S478”Checker PASS means safe to cherry-pick”the PASSBEFORE147.5 passed with deliberately deferred tsc debt owned by 147.6held the commit instead
S481per-lane checker PASS ⇒ integration greenper-lane PASSesAFTER (×2)post-pick full-surface run; shared test files carried semantic conflicts invisible to any single lane23 broken tests
S481”S480 self-gate knip clean”a self-gateAFTERit meant no-new-flags; bun run knip exits 1 on a ~96-item baselineclose-gate re-scoped
S484”The CI ANTHROPIC_API_KEY is live”it being setAFTERowner-confirmed dashboard-disabled; 3 W3 markers “structurally CANNOT pass”3 markers deferred
S488the VPS bundle clone is currentit existingBEFORE2 commits behind origin; a 07-15 conformance migration never fetchedreconcile before the gated commit
S494annotating a defect discharges ita “Sync drift” section cataloguing 4 defects and fixing noneAFTERaudit; “leaving the doc wrong for 23 days while telling readers it was wrong”23 days
S495an agent that live-proves in-DoD claims is reliable on side findingsrigour observed inside the DoDAFTER”Every out-of-DoD side finding was wrong or badly wrong… Rigour was brief-induced and stopped exactly at the brief’s edge”3 wrong facts relayed to the owner
S495a hedge survives the relaythe checker’s UNPROVEN verdictAFTER”the Coordinator restated it to the owner as settled fact with the hedge stripped”3 framings reached the owner wrong
S498DR-035 names its enforcement mechanismthe DR’s title and bodyAFTERits own implementing migration proves the mechanism a no-op. “Only the Status note carried the truth — and start-session greps headings, so the documented read path surfaced exactly the false half”a false ruling in the register
S498the register is the recordthe registerAFTER”The register accumulated COPIES of content that lives elsewhere, and the copies went stale while the originals stayed correct”DR-035 + DR-049 drift
S499”verification would confirm the tombstones”machine-generated tombstone metadataAFTERverification reversed 8 of 39; three provably false substance_moved_to values were hard-coded in the generator. “Generated hedge text is a work-list, never a record”8 wrong retirements
S500”the build gate covers scripts/“two green gatesAFTER”Two green gates shared one blind spot; only the third, differently-configured compiler caught it” (import.meta.dir is Bun-only)one full measurement round
S504”the T1 audit nits still need fixing”audit findingsBEFOREalready fixed by S503’s salvage. “An audit finding has a shelf life”
S506the nightly’s guard covers the webhooka ::warning printed in every run for weeksPRODunmasked sidecar artefact logs. “A ::warning that a human must scroll to is not a guard”weeks of dead webhook
S507”extractor-version-cross-ref fails on webhook-dead-era debris”the S506 hypothesis, written into the continuation promptBEFOREa 90-line test read; orphans are skipped cleanly. “Read the failing test before trusting a failure hypothesis”an owner ruling granted on a wrong basis
S511”the Stage-5 tests are structurally unsatisfiable”the S507 framingBEFOREa read-only inventory agent; PR #146 had landed distinct-bytes fixtures before the session started — yet it was “carried verbatim through the S509 board, the id-396 mint, and this session’s brief, and nearly became the spec’s premise""the session’s highest-value catch”
S516”Grounded in documents” ⇒ groundeddoc A citing doc BAFTERdoc B says the opposite. “The lane had been briefed on DR-104 and followed it; the rule was incomplete, not ignored”a WIRE verdict withdrawn
S516”an ERD in the architecture docs is design intent”09-diagrams.mdAFTERits own §1 boundary is “no new schema or flow content” and it sources column lists from database.types.ts. “DR-104’s code-evidence trap one step removed”two prior sessions treated it as prescription
S518”S517 restored recall without a rebuild”a green repair-statusAFTERthe palace re-broke within hours. “A green repair-status is not evidence the vector path works”a wrong all-clear
S519”both smoke tests passed, so the repair worked”smoke testsAMBIG”True here, but that is the same reasoning S517 used and S517 was wrong” — the real discriminator was a re-check from a later, separate process with a vector-path assertion
S524verify_driver.py is the ONE parameterised primitivethe file’s own docstring (“this module is NOT a fork”)BEFORE(owner)the two files share no code, it has zero runtime callers, and an S521 owner finding had already routed it to id-46 as a stale duplicate”Recall is not a session-start ritual; it fires before a verdict”
S525live-verify.sh is the manual entry point”preserved specifically as the manual against-a-live-box entry point”AFTER, monthsit hard-dies at :187. “That justification survived precisely because the thing it justified was never run… A preserved-for-manual-use artefact needs a date of last use, not a rationale”months unrunnable, root cause of W3
S527the mine verified cleanthe session’s own verification scriptBEFOREit diffed against the delete manifest (palace-side, 2,947 files) not the projection (disk-side, 2,934). “A verification step that can produce a false alarm is a verification step that will eventually be waved through”false 11-row alarm
SClaimSubstituteCaughtWhat caught itCost
S363ci-summary not requiredthe classic branch-protection API query returning nothingBEFOREowner flag + a ruleset query; required via production-protection ruleset 15785019a mis-read gate
S366subo-113 hung ~2h8 stale events under one of two event schemasAFTERparsing both .timestamp and .ts; the worker had committed 2 subtasks”cost a long forensic detour”
S3724 executors died”No task found” + no worktrees + no commitsAFTERthey had FINISHED; notifications batched ~25 min lateduplicate commit, one discarded
S373Staging variables inventory completeGitHub API default per_page=10BEFORE10 of 15 returned; the TEST_USER_*_EMAIL vars were hiddencaught pre-flip
S393prior harness-migration pass was incompletean absent final reportAFTERreading the prior commit + diff: it had deferred-by-design. “An absent report != unfinished work”wasted worker dispatch + a revert
S407”5 orphaned eval specs run 0 assertions in CI”an audit premiseBEFOREa fresh-context skeptic agent; 2 of 5 had non-gated unit tests running in CI shards”a naive delete-5 would have broken two eval targets”
S414the onprem deploy failed on configboth deploy jobs skippedBEFORE”the build job FAILED on a post-build telemetry step… the UUID code path was never reached”wrong root cause + a wrong conclusion, both in the ledger
S427”no manual q_a_pair create path exists / build new”a writer-grep on q_a_pairsBEFORE(owner)owner; a mature Q&A editor had shipped at S198, writing via the generic app/api/items route. “The grep concluded ABSENCE on the wrong layer”nearly greenfielded an existing feature
S429W2 wedged/deadlagging events.jsonl + no OQ recordBEFORE(owner)asking the owner before killing it — it was alive and working”ask-before-destructive-action paid off directly”
S449”the re-upload-diff trio is fully orphaned”executor claimBEFOREthe 17-final executor’s pre-deletion grep found a live caller in qa-detection.tsfile correctly kept
S449rg content_items gate is a retirement proofgrep on the headline tableAFTERblind to .from('content_history') embedding content_items!inner (5 live files)“grep every table in the drop set”
S450”production code is content_items-clean”the rg gateAFTER~15 production + ~50 test files tsc-broken at regen. “Same gate-blind-spot class as S449… third recurrencemass tsc breakage
S450the prod-residue executor’s tsc-0 claima broken grep pattern (unescaped alternation)BEFOREthe checker’s independent recount
S451source_documents has no ingest-axis columnsearching old column namesBEFOREID-138 M1 had added live origin_type + retention_class at S445. “A cross-task-state blindness that a ledger slice-read would have caught”wrong {133.5} flag
S455an API-error-killed agent produced nothingthe errorBEFORE”ECONNRESET often hits the RETURN channel after the work is complete”; a complete artefact was already on diskrecovery instead of redo
S457the 4-shard nightly red is an app regressionshard failure countsBEFORE”The bail-at-4 fail-fast made visible failures a floor, masking the true blast radius”wrong attribution
S478shard counts are the failure countPW_MAX_FAILURES=4AFTER103 tests never ran; “any shard count is a floor”untrustworthy green
S479”staging OKF deploy wiring does not exist yet”a stale docs page + a repo env-grepBEFOREthe {132.35} journal’s in-container probes prove it EXISTS. “env config lives in Coolify, so repo greps cannot disprove deployment state”wrong research verdict
S494a DR-NNN was never issuedgrepping the current registerAFTERgit log -S; 34 entries had been swept without trace. “Grepping the current decision register is a lying oracleDR-087 re-issued for an unrelated ruling; DR-090 burnt
S495latest DR is 089grepping ### DR-NNN headingsBEFORE”burnt ids and reserved ids are invisible to it by construction”; DR-090 burnt, DR-091 reservedre-allocated to 092/093/094
S496client prod is finenothing reporting otherwisePROD, 6 weekslist_branches; it had sat MIGRATIONS_FAILED since 13/06. “An 85-migration drift had a machine-readable signal for six weeks that nothing watched”6 weeks of drift
S498”there is no shared nav registry” (DR-041)a point-in-time finding promoted to a permanent rulingAFTERcomponents/shell/nav-config.ts — which cites DR-041 by name in its header — is that registry. “the same task then built the thing that falsified it”a false ruling in force
S498”the DR-087 code is on a branch, just unmerged”a cited SHAAFTERzero commits absent from main, zero occurrences of the symbols; “the cited SHA resolves nowhere”
S501”zero in-corpus .from(CONST) sites”the ts-morph corpus boundaryAFTERretro-miner + a direct query; the boundary is import-reachability, not tsconfig excludewrong coordinator conclusion
S501”workflow agent errored ⇒ work incomplete”the errorBEFOREimpl + fixtures + tests complete on disk, suite greenavoided a redo
S510{377.1}’s input file existsthe task itemBEFOREdeleted by an unrelated docs audit without striking the itemowner struck it
S512local-fs-platform/corpus does not exist (id-396 RESEARCH)the RESEARCH claimBEFOREit existed on disk as a stale derived copy carrying a DR-014-violating forms/ tree”the owner’s deletion was safe — but by derivation, not by any gate”
S515”grep found nothing”a grep over the column censusBEFORE(owner)owner; it searched the wrong surface — the corpus fixture gate already had a hardcoded DRIVER_MANIFEST_DEST_PATHS1 of 4 overturned verdicts → DR-104
S520”Q6 is blocked on #1665”the continuation promptBEFOREreading searcher.py: room=/wing= are genuine ChromaDB pre-filters; the 0-result shape was HNSW/metadata inconsistency. “I nearly built a client-side workaround for a healthy filter”recall-grounding still wrong today
S521”these .lavish boards were never rehomed”a filename-exact checkBEFOREcontent hash found 4 of 12 already rehomed under session-prefixed names
S522the docs/testing relocation is verifiedit was verified as a correct file moveAFTER”sweeping for what LIVED in it is the easy half. The half that gets missed is what POINTED INTO it” — 5 dangling citations + a specified-but-unbuilt destinationid-406 was 3 sessions from dispatch against a tree that no longer existed

Source/spec/ledger existence substituted for runtime existence.

SClaimSubstituteCaughtWhat caught itCost
S382the /extract cutover shippedit landed on code/mainPRODID-111.11’s E2E; /extract 503 on staging and prod. “Deployed client URL-ingest extraction was silently non-functional”a silent prod outage
S410PullMD is retiredretirement in the spec/code (ID-112) and the platform compose designBEFORElive VPS compose still deploys the full 4-container stack. “Retirement is in the spec/code only — NOT deployed”VPS right-sizing mis-sequenced
S445apply the batch and the id138 tests go greenthe migrations were authoredAFTERpost-apply test:integration; the fns were PostgREST-unreachable (no api wrappers). “Migration-authored != client-reachable”3 reds unfixable by any apply → DR-032
S456”the deployed image is current”the image being pinnedBEFOREstaging pinned a June sha while main had moved 200+ commits. “Deploy-branch/image-pin lag is the default state, not the exception”4 stacked live breaks
S472the settings fix is in effectit was committedAFTER ~30 sessionsthe validator silently drops misplaced keyssee V4
S481deploying from the track ref carries the compose fixthe fix being on the track branchBEFORECoolify clones compose from main; “the track-side compose fix had zero deploy effect”surgical PR #120
S481the bid-worker poller runsDocker reporting UpPRODCNB /cnb/process/web ENTRYPOINT silently ignores compose command:. “The processing_queue poller had NEVER run anywhere”never ran on either env
S498{163.20} shippedledger [x] + a recorded decisionAFTERcode existed nowhere in gitnear-total loss
S509CI exercises the shipped dependency closureCI installing from the lockAFTER”CI installs unpinned transitives while images ship the pinned lock — CI tests a closure nothing ships”rfdetr 1.9.0 broke the nightly
S523requirements.lock describes the imagethe lockfileAFTERgoogle-22 buildpack runs pip install with no --no-deps. “The lock is advisory, not authoritative”the whole dependency posture → id-416
S523the nightly exercises the deployed imageboth run pack buildAFTERonprem-deploy.yml adds a post-pack apt layer (git, openssh-client, LibreOffice) that the nightly does not. “The lane has been exercising an image missing four packages”routed to id-412
S527the push landed under branch protectionthe push succeededPRODGitHub reported ci-summary is expected + Missing successful active Staging deploymenttwo files on main with no CI
SClaimSubstituteCaughtWhat caught itCost
S456the cocoindex worker is healthyCoolify’s TCP-only healthcheck showing running:healthyPROD, days{127.30} LEG-2 curled /health. “/health (the app’s own endpoint) is the truth; the container check is not”days of a boot-crashed worker
S472mempalace recall is operationalthe SessionStart digest appearingAFTER”the hook’s raw FTS read survives corruption that kills every MCP recall path, so the digest is not evidence the palace is healthy”every mid-session recall failed since ~S458
S473an agent’s work is completeidle_notificationBEFOREa background Planner idled while its revision was incomplete, then completed later; a fresh Planner created a near-clobber race”verify the agent’s OUTPUT, never the idle/completion signal”
S475idle ⇒ donethe signalBEFORE~67 idle notifications, several silent-idles, every one recovered by a nudge-for-outputzero work lost — the discipline held
S511the diary landeddaemon job stateAFTERsubmits threw connection resets while entries landed anyway; “the only reliable confirmation is the lock-free sqlite FTS read”duplicate diary rows
S514the fixer completedstatus: completed on the task-notificationAFTERits own text read “waiting on the TS suite”; separately, a background gh pr checks --watch never fired its resultowner merged manually, watch output never read
S517the Retro Miner completed successfullyan idle signalBEFOREtwo idles fired before a model-config error surfaced; a third fired after delivering. “Treat idle as ‘check the output’, never as ‘work is done’“
S518the daemon is aliveit held the port, answered ps, had a pidfile and endpoint.jsonAFTERget_client_if_running returned None; sample showed every thread at 0% CPU behind a blocked worker”deadlocks on its FIRST job after boot, and looks alive while doing it”
S521the palace is corruptrepair-status reporting −9,768 divergenceBEFOREit reads OK with the note “HNSW has extra flushed elements; treating as safe” — and DR-110 makes the obvious reaction actively harmfula moment’s doubt, three times
S525the suite passeda background task’s reported exit code 0BEFORE (×2)reading the output file; the 0 was tail’s and the suite had 12 failures. “This is a harness trap, not an operator error""the catching discipline is written down nowhere”
S527the delete commit failedASSERTION FAILED printed on every successful commit passBEFOREanticipated from the S521 retro; verified externallynone — a retro was read

V8 — Authored-therefore-applied (NEW, 7)

Section titled “V8 — Authored-therefore-applied (NEW, 7)”

The agent’s own recent output treated as external evidence.

SClaimSubstituteCaughtWhat caught itCost
S513the TRIAGE.md text says Xthe coordinator’s memory of writing it (“every walk re-stamps” vs the file’s “every row is always current-walk”)BEFOREtranscript comparison; no hook had touched the file. “Quote the file, not the memory of writing it”mis-attributed to a formatter hook
S513Option B is rightthe coordinator had authored Option C’s own mechanism inside Option A’s description, then recommended B on sunk costBEFORE(owner)the owner’s “verify with the /cocoindex skill” rider; one read of api_reference.md:176-179 settled itwould have superseded an owner ratification on a nonexistent engine constraint
S516the fix is correct”because I just wrote it and reviewed it”AFTER (×3)three of four review rounds found a regression in the fix for the round before, all in code written and reviewed as correct minutes earliera credential leak that passed CI and a first review
S516the correction is rightit was a correctionAFTERthe Coordinator withdrew a WIRE verdict, swung to “retire-eligible” on an equally incomplete read, then to “not earnable either way”. “Two reversals on one question in one session”
S521the DR-104 rule is appliedit was quoted verbatim into a sub-agent dispatch brief in the same sessionAFTERthe owner; the same session mis-cited DR-034 (which says the opposite) then DR-041 (a navigation ruling). “Quoting a rule into a brief is not the same as applying it”2 wrong CLAUDE.md commits
S523the option menu is soundthe coordinator wrote itAFTER”I contradicted myself inside an owner-facing decision menu” — an option claimed the nightly could verify a .pdf while the session’s own context said it had never walked the corpusthe owner was asked to decide from a claim the session knew to be false
S527the pipeline room keyword is correctthe config’s own header block documents M3 (compounds match only single-token parts) — thirty lines above the violating keywordBEFOREthe projection; cocoindex-flow cannot reach cocoindex_pipeline, silently losing 95 files. “Documenting a mechanism is not the same as applying it. The guard was the measurement, not the understandingnone — the projection caught it

Part 3 — Distribution, and whether the mix is changing

Section titled “Part 3 — Distribution, and whether the mix is changing”

Rate. 161 occurrences / 135 retros = 1.19 per retro across the corpus. But the corpus is not uniform: retro length roughly triples from S264-era (~40 lines) to S515-era (~120 lines). Per-retro raw counts are therefore confounded by how introspective the retro is, not only by how often the error occurred.

Normalising by era:

EraRetrosOccurrencesPer retro
S264–S424 (May 25 – Jun 26)54460.85
S427–S488 (Jun 28 – Jul 18)40511.28
S490–S514 (Jul 22 – Jul 30, pre-DR-104)17301.76
S515–S527 (Jul 30 – Aug 3, DR-104 onward)13342.62

The rate goes up, monotonically, including after the DRs. I do not claim this is a real increase in error frequency — a large part of it is that the retros got much better at seeing the class, and S522/S524/S527 explicitly hunt for it. But the honest reading is: there is no era in this corpus where the rate falls.

Mix over time. V1/V2/V5 (the “code evidence” family, what DR-104 named) are distributed evenly across all four eras. V3 (task-directive) climbs sharply from the S437 back-propagation cluster onward and is the dominant variant in the S494–S527 ledger-heavy era. V6 (landed-as-live) is concentrated in infra/deploy sessions and does not decline. V4 is the only variant present in every era at roughly constant share (~30%).

The change that IS visible is not in frequency, it is in who catches it. See Part 5.


No. And the corpus contains four independent natural experiments, three of which predate DR-104.

Experiment 1 — DR-030 (S444) → breached 3× in S445, the next session

Section titled “Experiment 1 — DR-030 (S444) → breached 3× in S445, the next session”

S445: “api.-surface staleness fired 3x in one GO#2 despite DR-030 existing… Mocked bun run test is structurally blind to the whole class; only post-apply test:integration catches it. DR-032 written, but the guardrail is docs-only — no automated check exists yet.”*

S445’s own unresolved question asks whether a CI check should gate it mechanically. It was never built. This is the cleanest before/after in the corpus and it is three weeks older than DR-104.

Experiment 2 — DR-071 (S472, “GitNexus is the impact authority”) → three consecutive sessions of zero compliance

Section titled “Experiment 2 — DR-071 (S472, “GitNexus is the impact authority”) → three consecutive sessions of zero compliance”

S523: “DR-071 says GitNexus is the impact authority, and I retired a test file and added a new lib/corpus/ module on a plain git grep.” S524: “Zero code-intelligence calls in a session whose central question was ‘who consumes this path’. 216 Bash invocations, no memtrace, no GitNexus… S523’s retro recorded the same omission; recording it did not change the behaviour.” S525: “Zero code-intelligence calls, for the third consecutive session… Recording it twice changed nothing.”

S525 also contains the A/B that settles it: the one subagent whose brief mandated the tools returned impact(_load_canonical_content_types)47 transitive importers, “graph-only; invisible to path grep” — the exact row the main thread’s own worst error that session (judging blast radius from grep -rl | wc -l) needed.

Experiment 3 — DR-104 (2026-07-30) → breached within hours, and by sessions that cite it

Section titled “Experiment 3 — DR-104 (2026-07-30) → breached within hours, and by sessions that cite it”
  • S516, same day: “‘Grounded in documents’ is not the same as grounded… The lane had been briefed on DR-104 and followed it; the rule was incomplete, not ignored. — and separately, “Citing [the ERD] is DR-104’s code-evidence trap one step removed. Both prior sessions treated it as prescription.
  • S520 (Aug 1): “‘The field is populated’ is the same trap as DR-104’s ‘there are rows in the DB’.”
  • S525 (Aug 3): “That is DR-104’s ‘N callers import it, so it is live’ trap verbatim — which I had written into the taxonomy-spike brief as a binding rule in this same session.”
  • S525, again: “The pattern is not ‘ask more questions’; it is open the artefact before forming the verdict, which is DR-104 restated and which this session violated while quoting it.”

Worse than “no effect”: DR-104’s remedy created a new failure mode. DR-104 says docs outrank code. Within one day, S516 recorded a lane that followed that rule and reached a wrong verdict because it read doc A’s citation of doc B without reading doc B. DR-106 was written the next day to patch that (not every docs-site doc is a north-star doc). Then S521 got it wrong in the other direction — citing DR-034 (which says the opposite of the claim) and then DR-041 (a navigation ruling making no claim about concern boundaries), and concluded:

S521: The transferable lesson is about authority selection, not citation hygiene. I spent two rounds arguing from decision records because DR-104 says docs outrank code, and never asked which artefact actually ratifies this class of fact.”

The resolution came from the schema — a migration’s CHECK constraint and naming comment. DR-104’s binary (docs beat code) has no slot for that.

Experiment 4 — the same shape recorded, then repeated, one session later

Section titled “Experiment 4 — the same shape recorded, then repeated, one session later”
  • S522 recorded “an unticked multiSelect box logged as an explicit owner ruling”. S523 did it again with a deferral (“See previous response”) and wrote: “S522 recorded this exact shape and naming it was not enough to stop it recurring.”
  • S522 named “silently failing command read as data”. S523: “S522 named this class. Naming it did not prevent it. The mechanism that catches it is asserting the expected count, never reading a zero as agreement.”

There is a clean positive control, and it is not a doctrine — it is a line in a template:

SessionBriefs without the report-to-main lineBriefs with it
S5133/3 went silent-idle6/6 self-reported
S5142/2 went silent-idle3/3 self-reported
S5154/4 self-reported
S5164/4 self-reported

S516: “The report-to-main standing line now has three clean trials. Settled.”

And one more, from the most recent retro in the corpus:

S527: “Two known false alarms were correctly anticipated rather than re-earned… Both were called out before they fired and verified externally. This is what a retro being read looks like — and it is worth noting explicitly, because the two prior retros led on the absence of that.”

Both of those transferring findings share a property the DRs do not: they name a specific tool’s specific misleading output (delete_by_source returns no count on commit; a background task’s exit code is tail’s), or they are a mechanical line in a template. Findings about a reasoning posture did not transfer once, anywhere in 135 retros.


Part 5 — What actually prevents it, mined from the record

Section titled “Part 5 — What actually prevents it, mined from the record”

I classified the 118 “what caught it” cells. Four mechanisms account for essentially all catches.

A. Measurement — running the thing rather than reasoning about it (≈45% of BEFORE catches)

Section titled “A. Measurement — running the thing rather than reasoning about it (≈45% of BEFORE catches)”

This is the corpus’s clearest signal, and the retros repeatedly state it as a first-class rule in their own words:

  • S385: “Deterministic-extraction-over-agents: when an analysis task’s core is numeric… a Python script gives exact ground truth.” It also survived the API outage that killed an 8-agent workflow (211k tokens, 0 returns).
  • S408: “Verify DB-alignment claims with a column-level md5 fingerprint, not just matching object COUNTS — counts can match while structure diverges.”
  • S456: a controlled A/B rebuild disproved the inherited Sentry hypothesis.
  • S500: One control observation beats N suspect observations — a zero-dependency PR failing identically exonerated 14 dependabot PRs at once.
  • S516: Run the values; don’t read the regex. The \b anchoring bug was invisible to reading and obvious the moment real values went through the real function.” And: Mutation testing is what distinguishes a real test from a decorative one. Every guard this session was verified by disabling it and watching for red.”
  • S519: Project the operation before running it… turned three latent defects into checkpoint failures while no palace state had changed. Reading the config would have caught none of them.” Plus: Write the assertions as assertions — PASS/FAIL per checkpoint, not a table to eyeball.
  • S520: Let the projection fail first. Demonstrating the hazard beat asserting it.” And “Project competing policies, not just the chosen one.”
  • S521: Re-implement the dependency, do not model it. A passing suite proves only that the code agrees with your model of the dependency.” (18/18 assertions had passed against a fiction.)
  • S523: Reproduce the defect before fixing it and Build the RENDERED artefact, not an equivalent of it — it costs one sed.”
  • S527: Probe a keyword against the real tree before adding it. It turned every keyword choice from an argument into a measurement, and it is what killed three owner-proposed keywords and confirmed the fourth in the same pass.”

Corpus verdict: yes, measurement is the dominant preventive. It is also consistently described as cheap — “two minutes”, “four minutes”, “one script”, “one sed”, “ninety seconds”.

B. An owner challenge (≈25% of BEFORE catches; the exclusive catcher of the most expensive ones)

Section titled “B. An owner challenge (≈25% of BEFORE catches; the exclusive catcher of the most expensive ones)”

Sessions where the owner was the only thing that caught it: S350, S355, S356, S379, S382, S423, S427 (×2), S429, S457, S475, S479, S509, S511, S513, S515 (×4), S521 (×2), S522, S523, S524, S525 (×3).

The mechanism is specific and worth naming precisely: the owner is non-technical, so his challenge is never a counter-analysis — it is a question about the frame.

  • S423: “did you fix sourcing or just suppress?”
  • S513: “verify with the /cocoindex skill” — a rider that reversed a ratified option.
  • S522: “with the new DR, is the manifest still required, or just the actions it pointed to?”“the highest-value question of the session”; turned a six-job spec into five and reframed the artefact.
  • S522: “is the miner able or told to produce the output in the format you expect, or are they saving it to their session transcript?” — four minutes to answer, recovered the session’s most valuable artefact, and reversed a recommendation to delete a working tool.
  • S525: “is it the case that the live files still referencing these are accurate and correct?”

S525 states the pattern outright: “The owner’s questions caught three things the session’s own process did not… All three were cases where I had written a conclusion from a document or a count rather than from the thing itself.

And S521 records the sharpest version: “A non-technical instinct beat a technical inference because the instinct was checked against the docs and the inference wasn’t.”

This directly confirms the task brief’s hypothesis. The owner does not ask “is X live?” He asks what X is for. That question has no existence-count answer.

C. Adversarial review — but only when briefed to attack the PREMISE (≈20%)

Section titled “C. Adversarial review — but only when briefed to attack the PREMISE (≈20%)”

Works: S379 (two-round adversarial review flipped a backwards verdict), S393 (pre-PR adversarial workflow caught 2 must-fixes that 13,909 tests + tsc + dupe-check all missed), S397, S407 (“a fresh-context skeptic agent overturned the delete-5 premise”), S416 (4-lens fidelity review), S440 (“critics explicitly licensed to challenge standing DRs found the CRITICAL curation-destruction blind spot the synthesis missed), S451, S479, S516 (“Briefed to attack the redaction and to treat a clean verdict as an acceptable result” — found a leak CI and a first review had passed), S521.

But it fails on exactly this class when the brief is diff-shaped:

S515: “The pre-PR review gate (agents push, never open PRs) is the right shape but did not catch this class. All four verdicts survived it. The gate checks diffs; these were errors of premise, which only a docs-grounded read surfaces.”

S475: “UI shipped functionally complete but visually UNSTYLED… THREE checker rounds (incl. a dedicated UI re-check) missed it; caught only by Liam eyeballing the viewer post-close.”

D. Sub-agent escalation on a premise defect (small but perfect record)

Section titled “D. Sub-agent escalation on a premise defect (small but perfect record)”

S357 (×3), S422, S429, S434, S455, S475, S481. Every recorded instance where an executor refused to proceed on a defective brief was correct, and the cost was near zero.

S357: “task-executor escalation-on-boundary-defect worked exactly as designed 3x: executors refused to silently expand scope and returned precise boundary-correction packets. High-signal, near-zero rework. The discipline of a tightly-scoped ALLOWED file set is what made the escalations crisp.

  1. One well-documented case where consumer-counting was used affirmatively and was CORRECT — S355’s content_items.summary_data: the product owner’s prior was that the column was intended-for-drop; investigation refuted it by enumerating live readers (generateSummary writer; change-reports / item-detail-brief / MCP / search readers). The retro treats that enumeration as the right answer, and the drop was rejected. So consumer-counting is not always invalid — it is invalid as evidence of correctness, not as evidence of breakage risk. The corpus supports the narrower claim.
  2. The absence half was also used affirmatively and correctly in S355“the drop-intent had zero supporting evidence in memory or specs” was a load-bearing part of the rejection. Same caveat.
  3. The class is not evenly distributed by domain. V1/V5 concentrate in retirement/deletion work; V2 in test/eval work; V3 in ledger-heavy coordination; V6 in infra. A general control would over-fit; the retirement gate is where the expensive ones live (S451 prod P0, S466’s 2,235 LOC).
  4. The corpus already contains the fix for the register problem, ratified. S499: Hazard guards live at the action site, not (only) the register — the agent about to run rename(dry_run: false) or git stash never reads the register at that moment.” DR-033/039/053 were re-homed on it. Two of them had survived only inside a quarantined never-swept tree, “i.e. effectively nowhere.”

Part 6 — Where in the workflow it enters

Section titled “Part 6 — Where in the workflow it enters”
Entry pointCountRepresentative
Session-start inheritance (continuation prompt / handoff premise)41S511: a false framing “carried verbatim through the S509 board, the id-396 mint, and this session’s brief, and nearly became the spec’s premise”
Verdict formation / task close (retire, orphan, residue sweep, AC tick)35S451: a residue-sweep PASS closed Task-131 and left prod /guide/[slug] down
Ledger / register write (the claim becomes the authority)30S498: {163.20} done in the ledger, code nowhere in git
Dispatch-brief composition26S513: a brief promoted an inherited count to “verified” and was 2.6× wrong
Tool output consumption19S472: memtrace reported 18 callers where GitNexus correctly reported 151
Deploy / runtime boundary10S481: Docker Up while the poller had never run

The hot spot is unambiguous: the handoff → session-start boundary. A verdict formed under hedge in session N is written into the prompt or the ledger, and in session N+1 it carries the authority of a recorded fact with none of its original qualification. The corpus names the mechanism twice:

S495: That a hedge survives the relay. The degradation happened at the relay, not the source. The checker wrote the build gate as UNPROVEN; the Coordinator restated it to the owner as settled fact with the hedge stripped.”

S522: A correction can itself be incomplete, and inherit false authority from being a correction. A prompt that has already caught one error reads as trustworthy on the next line, which is precisely when it should not.”

S522, again: “A hedge hardened between draft and commit — ‘AC-1 largely discharged’ became ‘both blockers spent’. Small, and it is the direction errors travel in a ledger: the commit outlives the draft.”

The handoff format has no field recording how a claim was established or what would falsify it. Every claim arrives flattened to the same confidence.


Part 7 — Are the catches getting earlier or later?

Section titled “Part 7 — Are the catches getting earlier or later?”

Split by era, classifying each occurrence’s catch point:

EraBEFOREAFTERPRODOwner-only catch
S264–S42426 (57%)16 (35%)4 (9%)6
S427–S48829 (57%)17 (33%)5 (10%)8
S490–S51421 (70%)8 (27%)1 (3%)5
S515–S52722 (65%)12 (35%)0 (0%)13

Two things move in opposite directions.

  1. Catches got earlier and much cheaper. PROD escapes fall from ~10% to zero in the last era. The most expensive single instances all sit in the middle band: S451 (live prod P0, /guide/[slug] down), S466 (2,235 LOC of the sole writer deleted by a dead-code gate), S472 (a fix dead for ~30 sessions), S483 (an eval comparing against zero rows for weeks), S496 (a machine-readable MIGRATIONS_FAILED signal nobody watched for six weeks), S500 (a 15-day integration-lane failure ending in live queue corruption). The last era has nothing of that magnitude.

  2. Catches did NOT get more autonomous. Owner-only catches rise from ~13% to 38%. In the last five sessions of the corpus (S521–S527) the owner personally caught roughly two premise errors per session. He is not the backstop; on this class he is the primary detector.

S515: “‘Live code references it, so it is correct.’ The session’s defining error, made four times and caught four times by the owner, never by CI or by review.”

That sentence is the health metric the owner should be watching, and it has not improved in the eight sessions since.


Part 8 — Root-cause candidates, ranked, each tied to the data

Section titled “Part 8 — Root-cause candidates, ranked, each tied to the data”

RC1 — The rule is not present at the moment of the act. (Strongest evidence; already ratified once and not implemented)

Section titled “RC1 — The rule is not present at the moment of the act. (Strongest evidence; already ratified once and not implemented)”

Evidence. DR-030 breached 3× the next session, with the retro naming the cause: “the guardrail is docs-only — no automated check exists.” DR-071 has a three-session zero-compliance record with the third retro writing “recording it twice changed nothing.” DR-104 was breached within four days by a session that had quoted it into its own brief the same day. S499 ratified the diagnosis verbatim — “hazard guards live at the action site, not (only) the register” — and found that two such rules had survived only inside a quarantined tree, “effectively nowhere.”

Positive control. The one thing that transferred with near-100% reliability across four consecutive sessions was a line in a dispatch-brief template (3/3 vs 6/6, 2/2 vs 3/3, 4/4, 4/4). The second was a tool-specific output gotcha anticipated rather than re-earned (S527).

What the data says the fix looks like: not another DR. A template line, a required brief field, a CI gate, or a hook that fires at the action site.

RC2 — The question the agent is asked invites the error. (Directly confirmed; most actionable)

Section titled “RC2 — The question the agent is asked invites the error. (Directly confirmed; most actionable)”

Evidence for the “is X live?” framing producing wrong verdicts: S407 (5 orphaned specs → 2 wrong), S443 (no surviving caller → prod P0), S449/S450/S451 (three consecutive sessions, each finding a NEW blind-spot axis from a single residue-sweep PASS — six axes catalogued by S451), S466 (dead-code gate deletes the sole writer), S484 (knip counts tests as consumers, five orphans invisible), S507 (~80 unwired columns → 135 are a deliberate convention), S515 (all four overturned verdicts).

Evidence for the reframe working:

  • S522: the owner asked “with the new DR, is the manifest still required, or just the actions it pointed to?”“the highest-value question of the session”; it converted a location register into an orphan-and-integrity register, and “part of its original job was compensating for scatter that no longer exists.”
  • S520: Ask what is IN the category, not just whether the category populated. One GROUP BY exposed a room that was 82% not-decisions.
  • S527: Measure the margin, not just the count. ‘413 files route at P3’ did not justify a second config pass; ‘52% of them are noise or coin-flips, and here is app/page.tsx filed under procurement on a 2-vs-2 tie’ did.”
  • S476: reframing “pre-existing” as “real-but-already-broken-on-staging, no live users” was “the load-bearing distinction” that stopped a real search regression reaching prod.

The retirement/deletion gate is where this bites hardest and where the reframe would pay most: six blind-spot axes catalogued across S449–S451, plus S466’s 2,235 LOC, all from “is this orphaned?” rather than “what requirement does this serve, and is that requirement still live?”

RC3 — Provenance laundering across the session boundary

Section titled “RC3 — Provenance laundering across the session boundary”

A hedged verdict becomes an unhedged premise; a correction inherits authority from being a correction. S495 (the hedge does not survive the relay), S511 (a false framing carried through a board, a mint and a brief), S507 (a hypothesis written into the prompt, then an owner ruling granted on it), S520, S512, S522 (twice, including “the commit outlives the draft”), S523.

41 of 161 occurrences enter here. The handoff carries no provenance field.

RC4 — Verification scope is invisible in the verdict

Section titled “RC4 — Verification scope is invisible in the verdict”

A PASS/green/ratified/verified token records the conclusion but not the coverage. S451 (residue-sweep PASS → prod P0), S453 (“attestation is not verification”), S478 (PASS with deferred debt), S481 (per-lane PASS ⇒ integration green, ×2), S499 (“generated hedge text is a work-list, never a record”), S518/S519 (“a green repair-status is not evidence the vector path works”), S500 (“two green gates shared one blind spot”), S516 (mutation testing found a guard that could not fail and a load-bearing function with zero coverage).

The implied fix is a field, not a rule: what did this check actually execute?

RC5 — Tool output is trusted at its edges, where it is systematically wrong

Section titled “RC5 — Tool output is trusted at its edges, where it is systematically wrong”

Not a reasoning failure — a calibration failure with a repeatable shape. grep, knip, gitnexus, memtrace, ast-dataflow, git grep, gh-axi, supabase migration list, repair-status and delete_by_source all return confident zeros or confident counts that are artefacts of the query. S393, S406, S409, S443, S455, S457, S472, S475, S484, S501, S507, S523, S524, S525.

S499: “‘Citation counts are facts.’ 250 → 193 → 21, each shift a counting-method change, not drift.”

Note the asymmetry: DR-071 got the ranking right (GitNexus over memtrace) and still failed, because ranking a tool does not cause it to be called. See RC1.

RC6 — Self-authorship confers false warrant. (Smallest bucket; explains why naming fails)

Section titled “RC6 — Self-authorship confers false warrant. (Smallest bucket; explains why naming fails)”

Seven instances, and they are the ones that matter:

  • S516: three of four review rounds found a regression in the fix for the round before — “all in code written and reviewed as correct minutes earlier.”
  • S521: DR-104 quoted verbatim into a dispatch brief in the same session that breached it twice. “Quoting a rule into a brief is not the same as applying it.”
  • S527: the config’s own header block documents the mechanism thirty lines above the keyword that violates it. “Documenting a mechanism is not the same as applying it. The guard was the measurement, not the understanding.
  • S525: “which is DR-104 restated and which this session violated while quoting it.”
  • S513: the coordinator reconstructing its own authored text from memory and attributing the discrepancy to a formatter hook.
  • S523: “I contradicted myself inside an owner-facing decision menu.”

This is the mechanism by which naming fails. Writing the rule down, quoting it into a brief, or documenting the mechanism in a header comment all produce the subjective sensation of having applied it. The corpus contains four sessions where the agent breached a rule it had authored or quoted within the same session. No amount of additional naming touches this, because the failure is that naming feels like doing.

The corollary S527 draws is the actionable one and is the single best line in the corpus for the owner’s purposes:

“The guard was the measurement, not the understanding.”


Recommendation implied by the data, in one paragraph

Section titled “Recommendation implied by the data, in one paragraph”

Stop writing rules about the error and start changing two things it cannot survive. First, the question: replace every “is X live / dead / orphaned / wired?” task framing with “what requirement does X serve, and is that requirement still live?” — the corpus has zero instances of the second framing producing this error and eight of the first producing an expensive one. Second, the artefact: require a projection, probe, or mutation before any retirement, rename, or ratification, because measurement is the only thing in 135 retros that has ever caught this class without the owner. Both fixes are already in the corpus as proven, cheap practices; neither is in a template, a brief field, or a gate — and that, per RC1, is exactly why they have not stuck.