Phase 0.9 — Synthesis verification audit
Phase 0.9 — Synthesis verification audit
Section titled “Phase 0.9 — Synthesis verification audit”Audit date: 2026-05-11
Branch: content-items-investigation
Author: Verification sub-agent (worktree-isolated)
Subject: Adversarial audit of docs/plans/phase-0-investigation/0.9-synthesis.md (310 lines, commit 44cb8667)
Scope: accuracy (every claim traces to its underlying spike or source doc) + alignment with wider Phase 0.9 architectural plans (0.9-context.md North Star + S228 OQ ratifications + 0.9-intended-architecture.md v1.0 + 0.9-edit-flow-investigation.md §6 + 0.9-decision-graph.md + 13 spike reports).
Constraint: verification doc only — 0.9-synthesis.md is NOT modified.
1. Verdict
Section titled “1. Verdict”PASS-WITH-NOTES. Phase 2 commit recommendation is sound; the load-bearing gate verdicts (G1/G2/G3/G4/G10/G14) trace verbatim to their source spike reports; the architecture revisions in §5 align with the ratified edit-flow §6 and S228 OQ states; the reframings in §4 are evidence-supported. However, the synthesis carries 1 critical accuracy issue (the “0.9-spike-plan baseline 78%” anchor for the §1 confidence figure has no traceable source), 4 high-severity overstatements (verdict labels not strictly verbatim, confidence percentages partially uncalibrated, status of one architecture revision presented as “already landed” when actually unverified, S229 budget anchor unanchored against AI-pair-programming variance), and 8 medium/low items worth a clean-up pass before treating §1’s “92% overall confidence” as a load-bearing claim.
The synthesis is fit-for-purpose as a Phase 2 GO decision support document. It is NOT fit-for-purpose as a verbatim handoff to a downstream agent without the edits listed in §10.
2. Accuracy findings — itemised by severity
Section titled “2. Accuracy findings — itemised by severity”2.1 Critical (must fix before treating §1 confidence as load-bearing)
Section titled “2.1 Critical (must fix before treating §1 confidence as load-bearing)”A.C1 — “up from 0.9-spike-plan baseline 78%” has no traceable source. Synthesis §1 line 39:
“Overall confidence: 92% (up from 0.9-spike-plan baseline 78%).”
Grep for 78% across docs/plans/phase-0-investigation/:
0.9-spike-plan.md§8 line 785: “Spike phase plan confidence: 87% (up from 85% pre-S228).” — NOT 78%.0.8.2-cocoindex-evaluation.mdlines 76, 533, 537 — uses 78% for the cocoindex evaluation specifically.0.8-synthesis.mdline 637 — uses 78% for Stream 2 sequencing.0.7-synthesis.mdline 339 — uses 78% for 0.7.3 P6/P8 absorption.0.7.3-p6-p8-absorption.mdlines 7, 17 — same.0.8.7-mempalace-evaluation.mdline 514 — Shape B scope_tag.
No “spike-plan baseline 78%” anywhere. The actual spike-plan §8 confidence is 87%. The “78% → 92%” framing exaggerates the confidence delta by 9pp (actual 87% → 92% = 5pp delta).
Fix recommendation: change to “up from 0.9-spike-plan §8 baseline 87%” OR delete the parenthetical anchor entirely. Adjust narrative accordingly.
2.2 High severity
Section titled “2.2 High severity”A.H1 — §3 confidence percentages are not strictly calibrated against per-spike report verdicts. Cross-check:
| Synthesis §3 row | Synthesis confidence | Source spike confidence | Δ |
|---|---|---|---|
| S1 cocoindex schema-coupling | 96% | S1 §7 “Overall verdict (Scenario A confirmed): 96%“ | match |
| S10 cross-record dedup | 85% | S10 §9 “Spike confidence: 87%.” | -2pp |
| S14 cocoindex concurrency | 90% | S14 §9 “S14 verdict confidence: 92%.” | -2pp |
| S15 mempalace v5 upgrade | 95% | S15 §14 “Spike confidence: 96%“ | -1pp |
| S16 Q&A canonical schema | 80% | S16 §13 “Spike confidence: 80%“ | match |
| S2 cocoindex folder-binding | 92% | S2 §10 — no explicit headline confidence figure | n/a (inferable from §6.2 95% schema-level confidence) |
| S3 mempalace observe | 90% | S3 §10 “Spike confidence: 90% (up from 0.8.7’s 84% pre-spike).“ | match |
| S4 pullmd bake-off | 85% | S4 §7 “Overall verdict (CONDITIONAL PASS): 85%“ | match |
| S5 skill-seekers dep_analyzer | 95% | S5 — no headline confidence figure (qualitative “SKIP” verdict) | n/a |
| S6 mcp-scan | 95% | S6 — no headline confidence figure | n/a |
| S11 Playwright swap | 88% | S11 §9 “Confidence: 88%“ | match |
| S12 graphify confidence | 92% | S12 §8 “Spike confidence: 92%.“ | match |
| S13 ESLint input-required | 88% | S13 — no headline confidence figure (qualitative ADOPT-ADVISORY) | n/a |
Net finding: synthesis confidence figures are calibrated for 7 of 13 spikes (matched within 1pp), 3 are mild down-shifts (S10 -2pp, S14 -2pp, S15 -1pp — all conservative, not over-stating), and 3 have no headline confidence source in the spike (S5/S6/S13 — qualitative verdicts only). The S5/S6/S13 figures (95%/95%/88%) appear to be authored by the synthesis author with no traceable provenance.
Fix recommendation: annotate the S5/S6/S13 confidence figures as “synthesis-derived” or remove them; the qualitative verdicts (SKIP / PASS / ADOPT-ADVISORY) are not in dispute.
A.H2 — §11 confidence drag breakdown is inferential. Synthesis §11 lines 295-306:
“Phase 2 commit on cocoindex Option A | 92% | Aggregate across all gates” “The 8% drag splits roughly: AGPL v3 (3pp), v4-alpha PG-backend timing (2pp), Q&A OQ resolution timing (2pp), single-tenant test corpus skew (1pp).”
The 92% aggregate figure has no underlying calculation shown (e.g., min-of-gate-confidences? geometric mean? Bayesian combine?). The “8% drag splits” allocation (3+2+2+1=8) is a heuristic attribution by the synthesis author, not derivable from any spike report.
Fix recommendation: add one line explaining the aggregation method (“aggregate=min(gate-confidences)” OR “geometric mean of gate confidences” OR “author judgement weighted by gate severity”), and mark the 8% drag-split as “author allocation, not derived from spike reports.”
A.H3 — §5 row “Cloud Run topology change (already landed by S14 sub-agent — verify)”. Liam-flagged. Synthesis §5 line 149:
“§10 (Cloud Run topology) | Change from ‘queue-based serialisation candidate’ to ‘single-orchestrator-instance (
min_instances=1) + ephemeral per-instance LMDB’ | S14 §6 (already landed by S14 sub-agent — verify)”
Cross-check 0.9-intended-architecture.md §10.1 line 1336 (per grep -n "S14"):
“cocoindex | PRIMARY orchestrator (Option A pending S1; Cloud Run topology confirmed S14 — single-instance + ephemeral LMDB) | Canonical pipeline core; … S14 verified concurrent-process safety + crash-recovery; v1 ships
min_instances=1, max_instances=1OR scheduled Cloud Run job; no queue infra needed.”
Verified: this revision DID land — 0.9-intended-architecture.md §10.1 line 1336 already cites S14 in the cocoindex row. The “verify” parenthetical was author-flagged caution but the change is actually complete. However, the §10 narrative blocks (lines 1280-1330 ASCII diagram + flow text) also reference “single-instance Cloud Run; see S14” at line 1287 — those are also landed.
Caveat: S14 §6.1 recommendation to update §10 mentions a specific wording (“Single-instance orchestrator on Cloud Run (min_instances=1, max_instances=1) with ephemeral per-instance LMDB ops-DB…”) which is paraphrased rather than verbatim-inserted. The architectural fact is captured; the exact wording is the author’s restatement.
Fix recommendation: remove the “verify” parenthetical (status is verified — already landed); narrow the §5 row to “cosmetic wording-refresh OPTIONAL — S14 outcome already absorbed into §10.1 line 1336 + §10 diagram line 1287.”
A.H4 — §6.2 “Cocoindex Option A wiring (canonical pipeline) | 6-8 weeks | 6-8 weeks (unchanged)” anchor problem. Liam-flagged. The 6-8 weeks figure originates in 0.8.2-cocoindex-evaluation.md:
- Line 355: “Total: ~6-8 weeks, vs the Phase 0.7 Stream 2 estimate of 9-12 weeks. But the work is fundamentally different — instead of building canonical orchestration, we’re integrating cocoindex’s.”
- Line 453: table row “Effort | 9-12 weeks | 6-8 weeks (post-spike)”
- Lines 343-355 contain the breakdown justifying 6-8 weeks (Stream 1 collapse 1.5w + cocoindex-pipeline-impl 4-6w + reconcile-Stream-1-with-cocoindex 0.5-1w + buffer = 6-8w).
This estimate is inherited unchanged from 0.8.2 — it was authored before any spike was run, and it was authored under the explicit “Be cautious using timelines” guidance from Liam (preserved verbatim in 0.8-synthesis.md line 63):
“Be cautious using timelines (e.g., ‘~5-7 weeks; ~16-19h’) — if you consider that the current platform was built over the past 6-8 weeks, that should provide an indication of how different development timeframes are when building agentically in an AI-human paired-programming environment.”
The 6-8 weeks anchor was effectively a parallel-to-the-current-platform’s-development-window framing, NOT a calibrated estimate against the actual work. The synthesis carries this forward without re-anchoring against the S229/S230 spike outcomes. No spike empirically validates the 6-8 weeks figure.
Fix recommendation: add a parenthetical at §6.1 line 166 noting: “(6-8 weeks anchor inherited from 0.8.2 §sec 343-355 traditional dev-week framing; per Liam’s 0.8-synthesis.md line 63 caution, this is an AI-pair-programming-naive estimate and should be treated as a directional bound, not a calibrated commitment).” Same caveat for the 7-9 weeks and 8-10 weeks rollups in §6.2.
2.3 Medium severity
Section titled “2.3 Medium severity”A.M1 — §1 finding 1 paraphrases the Scenario A verdict tightly but drops the “live empirical verification deferred to Phase 2 first-step” caveat. Synthesis §1 line 16:
“1. Cocoindex schema-coupling resolves to Scenario A in source.
cocoindex.connectors.postgres.mount_table_target(..., managed_by='user')skips all DDL emission; KH’s existing 75-columncontent_itemstable + every FK + CHECK + GENERATED column + trigger + RLS policy survives unchanged.”
S1 §1 actual wording:
“Verdict: SCENARIO A CONFIRMED — Phase 2 commits on cocoindex. Engine offers a first-class
managed_by="user"mode in which the connector never emits DDL against the bound table. The verbatim docstring + statediff source code make this unambiguous; a follow-up empirical pass (run against an isolatedcontent_itemsclone) remains advisable as a Phase 2 first-step but does not block the architectural commit.”
The “live empirical verification deferred to Phase 2 first-step using the prepared harness” caveat (S1 §6 + §8) is fully captured later in §7 Outstanding Work table (S1-Q1…S1-Q6 listed). However, §1 finding 1 reads as more definitive than the spike. The synthesis §3 confidence column “96%” + Phase 2 first-step note (“Run prepared harness spike/cocoindex_s1/probe_managed_by_user.py…”) DOES capture this. So the §1 finding paraphrase is concise-but-fair; the empirical caveat is preserved in §3.
Fix recommendation: OPTIONAL — add ”, with empirical verification on Phase 2 day-1 using spike/cocoindex_s1/probe_managed_by_user.py harness (~0.5-1 day)” to §1 finding 1 second sentence for tighter alignment with the spike’s “advisable Phase 2 first-step” framing. Not load-bearing.
A.M2 — §1 finding 1 says “75-column content_items table”. S1 §1 says “75-column content_items table” — but 0.9-decision-graph.md and 0.9-intended-architecture.md consistently reference 70 columns (decision-graph Q1.1 line 53; intended-arch §1.2 line 56).
Cross-check the S1 spike report itself: line 17 says “75-column” but does NOT cite the source for this count. The decision-graph says “70 cols” since S227. Spike S1 reads as if 75 is a newer count from inspecting the live schema, but doesn’t show its work.
Fix recommendation: reconcile the 75 vs 70 discrepancy. Either (a) verify which is correct via \d content_items on staging (S1 should have done this; the spike harness is prepared but not run); or (b) note the discrepancy in §1 finding 1.
A.M3 — §1 finding 3 says “91.7%/1.5-4.5%, passing both gates” — the range 1.5-4.5% is correct but obscures the trade-off. S10 §1 table shows specific operating points:
“cent≥0.6 AND shared≥5 → TP=91.7% / FP=1.5%” “max_chunk≥0.70 AND shared≥4 → TP=91.7% / FP=4.5%”
Both pass the gate. Synthesis §1 line 20 says “91.7%/1.5-4.5%” — this is a range across two operating points, NOT a single point. The phrasing implies the substrate has natural variance in the FP range, which is misleading. The substrate has multiple passing operating points; the choice of operating point is a Phase 2 tuning decision.
Fix recommendation: rewrite as “Intersection of cocoindex + skill-seekers achieves 91.7% TP at FP rates between 1.5% and 4.5% depending on threshold selection (per S10 §2.7 operating-point sweep), passing both gates at multiple points.” More precise; clearer to a downstream reader.
A.M4 — §1 finding 5 says “5 client input shapes catalogued (2 docx-table from existing extractor; 2 NEW for markdown-heading + YAML-frontmatter; 1 LLM-extraction fallback)”. S16 §4 actually catalogues:
“Shape A — Audit-6col docx tables … Adapter today: Pattern A” “Shape B — DRAFT-5col docx tables … Adapter today: Pattern B” “Shape C — Phew-internal bid library markdown … Adapter today: None” “Shape D — Product bid library markdown … Adapter today: None” “Shape E — YAML-frontmatter-like blocks … Adapter today: None” “Shape F — Forms (PDF + XLSX) — questions only, no answers … NOT Q&A input”
Synthesis groups Shapes C+D together as the “markdown-heading” adapter (correct per S16 §4.7 table). Shape E is “YAML-frontmatter” (also correct). Shape F is for consumers, not producers. So the synthesis’s “5 client input shapes” maps to {A, B, C, D, E} — but S16 §4 has 5 producer shapes (A-E) requiring 5 adapter codepaths (not 5 shapes — Shape C and D share one adapter codepath markdown_heading_v1, per §4.7 table verdict). So:
- Synthesis says: “5 input shapes … 2 docx-table + 2 NEW for markdown-heading + YAML-frontmatter + 1 LLM-extraction”
- S16 §4.7 actually says: “Verdict — 5 distinct adapter codepaths needed: Pattern A, Pattern B, markdown-heading-v1 (NEW), YAML-frontmatter-v1 (NEW), LLM-extraction-fallback (NEW). The current pipeline covers only 2 of 5 (Pattern A + B).”
Synthesis’s count of 5 shapes is internally consistent with S16’s 5 adapter codepaths, but the breakdown “2 docx + 2 NEW markdown + 1 LLM” maps to (Pattern A + Pattern B) + (markdown-heading + YAML-frontmatter) + LLM-extraction = 5 adapters. The mapping is fine.
Fix recommendation: consider clarifying — “5 distinct adapter codepaths catalogued in S16 §4.7 (Pattern A docx-6col + Pattern B docx-5col from existing extractor; markdown-heading-v1 + YAML-frontmatter-v1 NEW; LLM-extraction fallback). Pipeline today covers 2 of 5.” Cleaner reading.
A.M5 — §1 finding 5 says “Q&A schema implementation (~19-27 days impl with ~8-10 days wall-clock concurrent per S16 §10).” S16 actual reference is §12.1, not §10:
- S16 §12.1 table line 1520: “Total | ~19-27 days | Concurrent-parallelisable to ~8-10 wall-clock days”
- S16 §10 is the migration plan (4 staged migrations), not the effort summary.
Fix recommendation: change “S16 §10” to “S16 §12.1” for accuracy.
A.M6 — §3 row for S15 lists confidence 95% but S15 §1 says 96%. Already flagged in A.H1 but worth a specific note: 1pp downshift is conservative-and-fair; not a real concern.
A.M7 — §3 row for S2 confidence 92% has no headline source. S2 §6 (Decision gate G2 outcome) does not provide a quantitative confidence figure. S2 §6.2 Step 5 says “G2 status: RESOLVED.” but no percentage. The 92% in synthesis appears author-derived.
Fix recommendation: annotate as “synthesis-derived” or remove.
2.4 Low severity
Section titled “2.4 Low severity”A.L1 — §1 finding 2 says “10 concurrent App.update() processes run cleanly against the same LMDB ops-DB with no lock errors, no corruption, and no measurable serialisation overhead.” S14 §2.4 says “wall-clock barely scales with N. Going from 1 → 10 concurrent processes increases wall-clock from 7.2s → 11s (1.5×). LMDB writer-lock contention is real but small.” 1.5× scaling is “small but measurable” — not “no measurable serialisation overhead.” Minor overstatement.
Fix recommendation: soften to “with small measurable serialisation overhead (1.5× wall-clock for 10× concurrency, per S14 §2.4).”
A.L2 — §4.5 says “Liam’s S230 framing applied: don’t over-weight the time-frame terminology in spike specs.” Correct verbatim of Liam’s S230-start caution; no issue.
A.L3 — §10 housekeeping gotchas: “Worktree sub-agent Bash CWD drift — after Read on worktree files, Bash CWD silently follows; subsequent git commands run in wrong tree.” This is a real gotcha that was added to CLAUDE.md §Gotchas in WP6 (already landed per the worktree’s CLAUDE.md showing the addition in the current worktree’s CLAUDE.md ### General section). Synthesis §10 line 287 says “bit me 4 times in S230 cherry-picks” — anecdotal, fine.
**A.L4 — §9.2 cross-reference correction note for 0.8.5 §Q5 — verified correct in S5 §5.4 (“Retract the ‘Stream 1 candidate’ framing from 0.8.5 §Q5 + the recommendations table item #9”). Synthesis carries this correctly.
3. Alignment findings — itemised by severity
Section titled “3. Alignment findings — itemised by severity”3.1 Critical
Section titled “3.1 Critical”Al.C1 — None. Synthesis does not contradict any S228 ratified position or wave-08 verbatim ratification. The §4 reframings are all evidence-backed by spike outcomes (S14 dissolves LMDB-blocker, S15 dissolves v5 hallucination, S10 reshapes UC8 substrate, S1 reframes S229 budget anchor, S16 reframes Q&A schema lock).
3.2 High
Section titled “3.2 High”Al.H1 — Synthesis §4.5 “Liam’s S230 framing applied: don’t over-weight the time-frame terminology in spike specs” applies the caution to the S1 spike budget specifically but does NOT apply it to §6’s effort estimates. This is the heart of Liam’s flagged concern A (Special Check A). Synthesis §6.1-§6.2 carry forward the 6-8 weeks / 7-9 weeks / 8-10 weeks figures with phrases like “6-8 weeks (unchanged)”, “~7-9 weeks”, “8-10 weeks” — bolded for emphasis. These figures inherit from 0.8.2 §sec 343-355 (which Liam’s S230 caution explicitly flagged as suspect). The synthesis applies the caution to the S1 budget (3-5 days → 30 minutes) but NOT to the Phase 2 implementation budget (6-8 weeks → unchanged). This is an asymmetric application of the caution.
Fix recommendation: add a paragraph at the start of §6 noting:
“These week-count estimates are inherited from
0.8.2 §5.4(Stream 2 fallback table) and0.8.2 §sec 343-355(per-substream cocoindex Option A breakdown), both authored before the spike outcomes landed. Per Liam’s0.8-synthesis.mdline 63 caution + S230-start framing, week-count terminology in spike specs should be treated as a directional bound under traditional dev-week framing, not a calibrated commitment for AI-pair-programming work. The numbers below carry these caveats; they are useful for relative ordering of streams (e.g. dedup adds 1 week net; Q&A adds 1 more) but not for absolute scheduling.”
This addresses Liam’s flagged concern A directly.
3.3 Medium
Section titled “3.3 Medium”Al.M1 — Synthesis §5 row “§16 (Open Questions table) Add new OQ: ‘Q&A schema two-tier model — q_a_pairs + q_a_extractions ratification’” — there is no corresponding new OQ in the synthesis’s own §7 (Outstanding work) table. The new OQ row for Q&A schema appears only in the §5 architecture revisions list. §7.1 lists S16 OQ1-OQ10 individually but does NOT promote the two-tier model ratification to OQ-status. This is a minor inconsistency: §5 says “add new OQ”; §7 doesn’t carry the addition through.
Fix recommendation: either (a) add a row at the top of §7.1 “OQ-NEW (S16): Two-tier q_a_pairs + q_a_extractions schema model — ratification at next decision-graph rewrite pass” OR (b) remove the §5 line about adding an OQ since it’s already covered by S16 OQ1-OQ10.
Al.M2 — Synthesis §5 row “§16 Update OQ4 mempalace status: ‘ADOPTION-CONFIRMED via dev workflow; v3.3.5 search now working’” is consistent with S15 verdict + S3 §6.4 verdict. But cross-check 0.9-context.md §2 OQ4: “INSTALLED-S227 + ADOPTION-PROVISIONAL. … Operational adoption as memory replacement still pending validation.” So OQ4 was still PROVISIONAL at S228; S15 + S3 (S229) DO confirm Lens 1 (dev workflow) adoption but neither directly addresses the “memory replacement” angle (S15 is upgrade-tracking; S3 is observe-only). The proposed status change to “ADOPTION-CONFIRMED” is partially supported (Lens 1 confirmed; Lens 2 “memory replacement” still provisional per OQ4 wording).
Fix recommendation: soften proposed OQ4 update to “ADOPTION-CONFIRMED for Lens 1 dev workflow; Lens 2 memory replacement still post-launch validation” to match what S15 + S3 actually verified.
Al.M3 — Synthesis §3 lists S8 + S9 as “post-architecture-commit spikes” but the table itself shows S7 + S8 + S9 in the “S230-deferred” group (3 entries). §3 second-to-last paragraph says:
“The remaining spikes (S8 Q&A flow validation + S9 write-back semantics) are post-architecture-commit spikes.”
But the table (lines 91-99) lists S7 + S8 + S9 — 3 deferred spikes, not 2. S7 is the pre-re-ingest evaluation per 0.9-spike-plan.md §S7. S7’s “deferred to S230” status carries forward but it’s NOT a post-architecture-commit spike — it’s a pre-Phase-2 spike that was simply not run in S229/S230 due to budget priority. The synthesis text implies only S8+S9 are post-architecture; the table includes S7 in the deferred-not-post list.
Fix recommendation: rewrite §3 second-to-last paragraph as “The remaining S230-deferred spikes — S7 (pre-re-ingest eval, dispatched pre-Phase-2 first-real-re-ingest) + S8 + S9 (post-architecture-commit) — run in Phase 2, not before.” Matches the table.
3.4 Low
Section titled “3.4 Low”**Al.L1 — Synthesis §6.1 table row “WP-DEDUP-RULES (hybrid (a)+(c) wire + per-tenant rules) | n/a in S229 | ~5-7 days hybrid + 1-2 days per-tenant — S10 NEW” — S10 §10 says “~4-6 days for hybrid integration + 1-2 days per-tenant rule curation per onboarded client.” Synthesis lists 5-7d; S10 says 4-6d. 1d upshift, minor.
Fix recommendation: reconcile to S10 §10’s 4-6d OR cite synthesis’s own conservatism.
**Al.L2 — Synthesis §6.1 “ESLint v2 + 6 handler-level path-param gap fixes” — S13 §8 says “open six small PRs (or one focused PR) addressing the six handler gaps. Each is ≤5 lines.” Synthesis estimates 0.5 day — fair.
**Al.L3 — Synthesis §7.2 spike-tail residuals lists S1-Q1…S1-Q6 (six items) — these match S1 §5 (six residual questions) verbatim.
4. Timeframe estimate audit (Liam-flagged Special Check A)
Section titled “4. Timeframe estimate audit (Liam-flagged Special Check A)”Subject: §6 “Effort estimates — pre-spike vs post-spike” carries 6-8 weeks / 7-9 weeks / 8-10 weeks figures. Liam flagged at S230-start: “we need to be cautious and refrain from putting too much weight behind the time frame terminology” + 0.8-synthesis line 63 quote “the current platform was built over the past 6-8 weeks”. Are these numbers anchored to anything beyond inherited estimates from 0.8.2-cocoindex-evaluation.md §sec 343-355?
4.1 Provenance trace
Section titled “4.1 Provenance trace”| Synthesis figure | Source | Authored when |
|---|---|---|
| 6-8 weeks (cocoindex Option A architecture-impl) | 0.8.2-cocoindex-evaluation.md §sec 343-355 (per-substream breakdown) + line 453 (table) + line 445 (recommendation) | S214 / Phase 0.8 (~3 weeks before spike phase) |
6-8 weeks (carried forward in 0.8-synthesis.md line 51, 131, 137, 148) | Carries 0.8.2 estimate unchanged | S217 |
6-8 weeks (carried forward in 0.9-intended-architecture.md §1.2 line 55 “REPLACE with cocoindex Option A + pullmd + selective skill-seekers (~6-8 weeks post-spike)“) | Carries 0.8-synthesis estimate unchanged | S229 |
6-8 weeks (carried forward in 0.9-synthesis.md §1 line 37 “unchanged from S229 estimate”) | Carries 0.9-intended-architecture.md estimate unchanged | S230 |
| 7-9 weeks (synthesis §6.2 line 183) | Author-derived: “(cocoindex + mempalace + pullmd + ESLint + snyk + graphify confidence + Q&A schema + dedup hybrid)“ | S230 (this synthesis) |
| 8-10 weeks (synthesis §6.2 line 185 “Wall-clock Phase 2 total”) | Author-derived: 7-9 weeks + 1 week standard padding | S230 (this synthesis) |
| 9-12 weeks (Phase 0.7 Stream 2 fallback envelope) | 0.7-synthesis.md Stream 2 estimate + carried in 0.8.2-cocoindex-evaluation.md line 399, line 453 | S214 |
4.2 Liam’s caution traceability
Section titled “4.2 Liam’s caution traceability”Liam’s S230-start verbatim caution (per task brief): “we need to be cautious and refrain from putting too much weight behind the time frame terminology.”
Liam’s prior verbatim caution (preserved at 0.8-synthesis.md line 63):
“Be cautious using timelines (e.g., ‘~5-7 weeks; ~16-19h’) — if you consider that the current platform was built over the past 6-8 weeks, that should provide an indication of how different development timeframes are when building agentically in an AI-human paired-programming environment.”
4.3 Audit verdict on the §6 timeframes
Section titled “4.3 Audit verdict on the §6 timeframes”The 6-8 weeks figure is anchored ONLY to the 0.8.2 traditional dev-week breakdown, which itself was authored under the explicit caution that traditional dev-weeks don’t apply to AI-pair-programming work. The synthesis carries this estimate unchanged through four documents (0.8.2 → 0.8-synthesis → 0.9-intended-architecture → 0.9-synthesis) without any re-anchoring against:
- Actual AI-pair-programming velocity observed during spike phase (S229 8 parallel spikes in 1 session; S230 5 spikes in 1 session — the spikes themselves provide a velocity anchor that the synthesis doesn’t use).
- Specific cocoindex flow LOC delta projections (S16’s “~500-800 LOC declarative flow + dependency” from
0.8-synthesis.mdline 51 is a more relevant anchor than week-count, but it’s not in §6). - Phase 2 first-step delivery proof points (the prepared S1 harness + the S16 staged-migration sketch).
Worse: §6.1 line 175 says “Cloud Run single-orchestrator config | ~0.5 day | ~0.5 day (new — S14)” — half-day estimate has no source. S14 §6.4 says topology change is “configuration, not code” — but doesn’t quantify in days.
§6.1 lines 169 “Pullmd Docker Compose + Tier 2/2.5/3 cascade adoption | ~2-3 days” — S4 §1 says “Drop Firecrawl + @mendable/firecrawl-js ~0.5 day” but doesn’t quantify the Tier 2/2.5/3 adoption in days. The 2-3 day estimate is author-derived.
§6.1 line 174 “snyk-agent-scan CI step | ~0.5 day” — S6 §4.1 says “Effort: ~0.5 day. - 1h — wire pipx install + scan command into ci.yml as a new job - 2h — write scripts/mcp-scan/serve-fixture.ts stdio bridge - 1h — JSON-output parser + baseline comparison logic + CI fail conditions - 0.5h — seeded mcp-scan-ignore.json allow-list - 0.5h — runbook entry in docs/runbooks/ci.md”. Match — this one IS anchored.
§6.1 line 176 “WP-DEDUP-RULES | ~5-7 days hybrid + 1-2 days per-tenant” — S10 §10 says “5-7 days for KH-native classifier + rule-curation UI + per-tenant onboarding flow” but also “4-6 days for hybrid integration” — synthesis uses 5-7d which is the upper end of the per-tenant rules budget AND the lower end of the hybrid integration. Slight anchor inconsistency (4-6d vs 5-7d), but in the right ballpark.
4.4 Recommendation
Section titled “4.4 Recommendation”Add a leading paragraph at §6 (before §6.1):
“These week-count estimates inherit from
0.8.2-cocoindex-evaluation.md§sec 343-355 (traditional dev-week breakdown authored S214) carried forward through0.8-synthesis.md+0.9-intended-architecture.md+ this doc. Per Liam’s0.8-synthesis.mdline 63 caution and S230-start framing, these figures are AI-pair-programming-naive estimates and should be treated as directional bounds for relative ordering of streams, NOT calibrated commitments for scheduling. The spike-phase velocity anchor (8 parallel spikes in S229; 5 in S230) suggests the actual Phase 2 timeline could land anywhere from 30% under to 30% over these figures, with the asymmetric risk on the upside (specifically: Q&A LLM-extraction prompt engineering + per-tenant rule curation are novel; their estimates carry the largest uncertainty).”
Without this addition, the bolded 6-8 weeks (unchanged), 7-9 weeks, 8-10 weeks figures in §6 read as commitments. With it, they read as bounds. This is the load-bearing edit Liam asked about.
Confidence assessment for these timeframes: synthesis says 92% overall confidence; per the audit above, the timeframe-specific confidence should be lower — perhaps 65-75% on the absolute numbers, 85% on the relative ordering. The §11 confidence table does not break out timeframe confidence separately; it bundles into “Phase 2 commit on cocoindex Option A | 92% | Aggregate”. The timeframe is the largest single source of drag that the §11 breakdown does not surface.
5. Source-claim audit (Liam-flagged Special Check B)
Section titled “5. Source-claim audit (Liam-flagged Special Check B)”Subject: §4.3 admits the author cited “mempalace v5.0.0 GA shipped 2026-05-02” from a flawed start-session WebSearch. Cross-check whether ANY other synthesis claim has similar provenance issues.
5.1 Categories of source claim audited
Section titled “5.1 Categories of source claim audited”I cross-checked every quantitative claim, every version number, every date, and every “RESOLVED / RATIFIED / CONFIRMED” tag against its purported source.
Version numbers and dates:
| Claim location | Claim | Source | Verified? |
|---|---|---|---|
| §1.1 line 16 | ”cocoindex.connectors.postgres.mount_table_target(…, managed_by=‘user’)“ | S1 §2.1 line 36 + S1 §2.3 line 92 | YES — verbatim from spike |
| §1.1 line 16 | ”75-column content_items table” | S1 §1 line 17 | Stated in S1 but unverified (see A.M2) — synthesis carries forward unverified |
| §1.4 line 22 | ”mempalace v3.3.5” | S15 §1 + §3.1 (live install version) | YES — verified live |
| §1.4 line 22 | ”PyPI’s latest is v3.3.5” | S15 §3.2 line 73 | YES — verbatim |
| §1.4 line 22 | ”v4-alpha PRs (#665 PG drawer + #1337 PG KG) remain open against develop” | S15 §3.5 + §3.6 | YES — verbatim from spike |
| §1.4 line 22 | ”mempalace_search is FIXED in v3.3.5 (PR #1396 — retry-on-transient + drift-segment auto-quarantine)“ | S15 §1 + §3.3 + §4 (PR #1396 cited verbatim) | YES — verbatim |
| §3 S10 row line 73 | ”85% confidence” | S10 §9 “Spike confidence: 87%“ | -2pp downshift (conservative) |
| §3 S14 row line 74 | ”90% confidence” | S14 §9 “92%“ | -2pp downshift (conservative) |
| §10 line 285 | ”cocoindex 1.0.3” | S2 + S1 + S14 + S10 all confirm | YES |
| §10 line 286 | ”pullmd ships only as Docker Compose / npm-from-source / pre-built Docker image (aeternalabshq/pullmd)“ | S4 §2.3 line 88 | YES — verbatim |
| §10 line 284 | ”package renamed from mcp-scan to snyk-agent-scan v0.5.1” | S6 §0 TL;DR line 19 | YES — verbatim |
Verdicts and gate outcomes:
| Claim location | Claim | Source | Verified? |
|---|---|---|---|
| §1.1 line 16 | ”Phase 2 commits on cocoindex” | S1 §1 line 9 “PHASE 2 GO” / §7 line 237 “Scenario A confirmed” | YES |
| §1.2 line 18 | ”v1 topology: single-orchestrator-instance + ephemeral per-instance LMDB. No queue infra needed.” | S14 §6.1 line 308 verbatim | YES |
| §1.3 line 20 | ”UC8 (SMB data-fix critical feature, OQ3 RATIFIED) ships in v1 with cross-record dedup.” | OQ3 RATIFIED per 0.9-context.md §2 + UC8 v1 path per 0.9-edit-flow-investigation.md §6.8 | YES |
| §1.5 line 24 | ”Hybrid client-input recommendation: opinionated template (docs/templates/client-qa-bundle-template.md to author) + accept-any dispatcher.” | S16 §7 verbatim recommendation | YES |
| §2 line 51 | ”G1 — Scenario A — managed_by="user" skips DDL” | S1 §1 + §2 | YES |
| §2 line 56 | ”G14 — Single-orchestrator topology” | S14 §6.1 | YES |
5.2 Hallucination risk audit
Section titled “5.2 Hallucination risk audit”No additional hallucinated claims found beyond the one §4.3 already admits. The synthesis is self-aware about the start-session WebSearch error (logged as ERRATA-1 in S10 spike). The 78% baseline confidence anchor (A.C1) is the closest analogue — but that’s a mis-paraphrase of an inherited document (not a hallucination from web search), so it falls under accuracy rather than hallucination.
Specific cross-check on dates:
- S15 §1 verbatim: “PyPI … published to PyPI 2026-05-10 23:44 UTC — i.e., the day before the date in
currentDatemetadata.” Today (synthesis audit date) is 2026-05-11. So v3.3.5 was uploaded ~26h before synthesis. Date integrity is intact. - §10 line 287 says “bit me 4 times in S230 cherry-picks” — author self-reported anecdote, unverifiable but plausible.
- §4.3 line 117 says “WP3 logged ERRATA-1 in
0.9-spike-S10-dedup-substrate.md §10” — verified, S10 §7.4 (not §10) lists “ERRATA-1” verbatim. Section number off-by-three but verifiable.
Fix recommendation: §4.3 “WP3 logged ERRATA-1 in 0.9-spike-S10-dedup-substrate.md §10” — should be §7.4 (where ERRATA-1 is actually documented). Minor.
5.3 Verdict on Special Check B
Section titled “5.3 Verdict on Special Check B”PASS. No additional hallucinated source claims found. The mempalace v5 issue is the only one. The synthesis demonstrates appropriate caution about that specific failure.
6. Gate-resolution audit (Liam-flagged Special Check C)
Section titled “6. Gate-resolution audit (Liam-flagged Special Check C)”Subject: Verify each gate (G1/G2/G3/G4/G10/G14) traces verbatim to its spike report’s verdict. Are confidence percentages calibrated, or pulled from thin air?
6.1 G1 — S1 cocoindex schema-coupling
Section titled “6.1 G1 — S1 cocoindex schema-coupling”Synthesis §2 line 51: “Scenario A — managed_by="user" skips DDL”
Synthesis §3 line 72 confidence: 96%
S1 §1 verdict: “SCENARIO A CONFIRMED — Phase 2 commits on cocoindex.” S1 §7 confidence table: “Overall verdict (Scenario A confirmed) | 96% | Source-code-conclusive; live test is a verification step, not a decision step.”
Match: VERBATIM. Confidence: CALIBRATED.
6.2 G2 — S2 cocoindex folder-binding
Section titled “6.2 G2 — S2 cocoindex folder-binding”Synthesis §2 line 52: “localfs-only v1 + fs-watch UC10” Synthesis §3 line 82 confidence: 92%
S2 §5: “v1 connector list = localfs only. SharePoint, Notion, Dropbox, Box absent from 1.0.3.” S2 §6 + §10: “G2 status: RESOLVED” — no explicit confidence figure.
Match: VERBATIM on verdict. Confidence: UNCALIBRATED (S2 has no headline confidence). The 92% in synthesis is author-derived.
6.3 G3 — S3 mempalace observe
Section titled “6.3 G3 — S3 mempalace observe”Synthesis §2 line 53: “Shape A+B+C+miner all CONFIRMED” Synthesis §3 line 83 confidence: 90%
S3 §1 verdict: “G3 — Shape A + B + C + miner adoption confirmed with one caveat (search broken upstream).” S3 §10: “Spike confidence: 90% (up from 0.8.7’s 84% pre-spike).”
Match: VERBATIM. Confidence: CALIBRATED.
6.4 G4 — S4 pullmd bake-off
Section titled “6.4 G4 — S4 pullmd bake-off”Synthesis §2 line 54: “CONDITIONAL PASS — adopt for HTML/CF/GN/Reddit” Synthesis §3 line 84 confidence: 85%
S4 §1: “CONDITIONAL PASS — adopt pullmd as Tier 2 / Tier 2.5 / Tier 3 replacement for HTML and Reddit URL paths. KEEP KH’s Jina Reader fallback for PDF URLs.” S4 §7 confidence table: “Overall verdict (CONDITIONAL PASS) | 85%”
Match: VERBATIM. Confidence: CALIBRATED.
6.5 G10 — S10 dedup substrate
Section titled “6.5 G10 — S10 dedup substrate”Synthesis §2 line 55: “CONDITIONAL PASS via HYBRID (a)+(c)” Synthesis §3 line 73 confidence: 85%
S10 §1 verdict: “Spike status: CONDITIONAL PASS — substrate (c) skill-seekers keyword passes alone (TP=100% / FP=4.5%) BUT is per-tenant-brittle; recommended v1 path is HYBRID (a)+(c) cocoindex chunk-embed AND skill-seekers keyword co-confirmer at TP=91.7% / FP=4.5%” S10 §9: “Spike confidence: 87%.”
Match: VERDICT VERBATIM. Confidence: -2pp downshift (synthesis says 85%; S10 says 87%). Conservative downshift, not over-stating.
6.6 G14 — S14 cocoindex concurrency
Section titled “6.6 G14 — S14 cocoindex concurrency”Synthesis §2 line 56: “Single-orchestrator topology” Synthesis §3 line 74 confidence: 90%
S14 §1 verdict: “PASSED — S2’s ‘single-writer constrains multi-worker Cloud Run topology’ framing is REVISED. … Recommended v1 Cloud Run topology = (single-orchestrator-instance)” S14 §9: “S14 verdict confidence: 92%.”
Match: VERDICT VERBATIM. Confidence: -2pp downshift (synthesis says 90%; S14 says 92%). Conservative downshift, not over-stating.
6.7 Gate-resolution audit summary
Section titled “6.7 Gate-resolution audit summary”| Gate | Verdict match | Confidence calibration |
|---|---|---|
| G1 | Verbatim | Match |
| G2 | Verbatim | Uncalibrated (no source figure) |
| G3 | Verbatim | Match |
| G4 | Verbatim | Match |
| G10 | Verbatim | -2pp (conservative) |
| G14 | Verbatim | -2pp (conservative) |
Verdict on Special Check C: PASS-WITH-MINOR-NOTE. All 6 gates trace verbatim. Confidence figures are calibrated for 3 of 6, conservatively downshifted for 2 of 6 (both ~2pp), and author-derived for 1 of 6 (G2 — S2 has no headline confidence figure). No over-stated confidence.
7. Architecture-revisions audit (Liam-flagged Special Check D)
Section titled “7. Architecture-revisions audit (Liam-flagged Special Check D)”Subject: Each revision row in §5 claims a specific source spike. Verify the source spike actually contains that recommendation. Especially: §10 Cloud Run topology change “(already landed by S14 sub-agent — verify)” — was it actually landed?
7.1 Per-row audit
Section titled “7.1 Per-row audit”| Synthesis §5 row | Claim | Source cited | Source verified? |
|---|---|---|---|
| §5 (UC8 substrate) | Update to “cocoindex @coco.fn chunk-embedding primary + skill-seekers keyword co-confirmer (HYBRID); WP-DEDUP-RULES manages per-tenant rules” | S10 §5.5 | YES — S10 §5 line 280 verbatim: “§12.1 ‘MCP check_content_duplicates tool’ — substrate is HYBRID (a)+(c)” + §7.3 verbatim: “WP-DEDUP-RULES — per-tenant keyword-rule classifier” |
| §5 (Q&A) | Add q_a_pairs + q_a_extractions as separate tables (was: assumed q_a_extractions only) | S16 §6 | YES — S16 §6.2 (q_a_pairs schema) + §6.3 (q_a_extractions schema). Both tables explicitly designed. |
| §5 (§10 Cloud Run) | “Change from ‘queue-based serialisation candidate’ to ‘single-orchestrator-instance + ephemeral per-instance LMDB‘“ | S14 §6 (already landed by S14 sub-agent — verify) | VERIFIED AS LANDED per A.H3 audit — 0.9-intended-architecture.md §10.1 line 1336 + §10 ASCII diagram line 1287 both reference S14 outcome. The “verify” parenthetical is no longer needed. |
| §5 (§11.x sidecar deployment) | “pullmd via Docker Compose; Trafilatura + Playwright sidecars; PDF pre-route to Jina; remove Firecrawl” | S4 §1 | YES — S4 §1 verbatim recommendations. |
| §5 (§13.1 Phase 2 gate) | “Scenario A confirmed; commit on cocoindex Option A” | S1 §1 | YES — S1 §1 + §7 (G1 RESOLVED). |
| §5 (§16 OQ4 update) | “ADOPTION-CONFIRMED via dev workflow; v3.3.5 search now working” | S15 | PARTIAL — per Al.M2 audit. S15 confirms search works; “ADOPTION-CONFIRMED” framing partial (Lens 1 yes; Lens 2 still provisional). |
| §5 (§16 add new OQ) | “Q&A schema two-tier model — q_a_pairs + q_a_extractions ratification” | S16 | PARTIAL — per Al.M1. S16 §11 lists 10 OQs but doesn’t promote two-tier-model to OQ status. Synthesis adds it to the architecture revision list; this is fine as an architecture revision but the proposed §16 OQ wording is author-derived, not S16 verbatim. |
| §5 (Footnote refs) | “Add references to S1 / S10 / S14 / S15 / S16 spike reports” | All | TRIVIAL — documentation hygiene. |
7.2 Audit verdict on Special Check D
Section titled “7.2 Audit verdict on Special Check D”PASS-WITH-NOTES. All 7 architecture revisions trace to their source spike. The S14 Cloud Run revision is verified as already-landed (the “verify” parenthetical is stale). The §16 OQ4 update is partial (Lens 1 confirmed, Lens 2 still provisional). The §16 new-OQ row is author-derived but reasonable. No revision is fabricated.
Fix recommendation: apply A.H3 (remove “verify” parenthetical on Cloud Run topology row) + Al.M2 (soften OQ4 update to Lens 1) + Al.M1 (either add OQ row to §7.1 or remove from §5).
8. Cross-reference audit (Liam-flagged Special Check E)
Section titled “8. Cross-reference audit (Liam-flagged Special Check E)”Subject: Synthesis references many spike reports + 0.8.x evaluation docs. Verify each cited section exists + supports the claim.
8.1 Cited spike sections
Section titled “8.1 Cited spike sections”| Citation | Section exists? | Section supports claim? |
|---|---|---|
| S1 §1 + §2 (synthesis §1.1) | YES | YES |
| S14 §1 (synthesis §1.2) | YES | YES |
| S10 §1 + §5.5 (synthesis §1.3) | YES — both | YES |
| S15 §1 + §4 (synthesis §1.4) | YES — both | YES |
| S16 §1 + §12 (synthesis §1.5) | YES — both (§12 = §12.1 effort summary) | Partial (synthesis §1.5 cites §10 but actual reference is §12.1 — see A.M5) |
| S2 §3 (synthesis §9.1) | YES | YES — S2 §3 documents the API drift |
| S1 §1 (synthesis §9.1) | YES | YES — confirms Scenario A from source |
| S10 §1 + §5.5 (synthesis §5 row UC8) | YES — both | YES |
| S16 §6 (synthesis §5 row Q&A) | YES | YES |
| S14 §6 (synthesis §5 row Cloud Run) | YES (S14 §6 = G14 outcome) | YES |
| S4 §1 (synthesis §5 row sidecar) | YES | YES |
| 0.8.5-skill-seekers-evaluation.md §Q5 (synthesis §9.2) | YES | YES — §Q5 over-credited the tool; spike retracts |
| 0.8.7-mempalace-evaluation.md §5.6 + §3.1 (synthesis §9.3) | YES | YES |
0.9-edit-flow-investigation.md §6.8 (synthesis §4.4) | YES | YES — UC8 v1 is human-confirmed-merge |
0.9-edit-flow-investigation.md §6.0.6 (synthesis §4.4 implicit) | YES | YES |
0.9-intended-architecture.md §5 (synthesis §4.5) | YES | YES — Scenario A vs B framing |
0.9-intended-architecture.md §10 (synthesis §5 Cloud Run row) | YES | YES — confirmed via grep |
| 0.7-synthesis (synthesis §4.6 implicit) | YES | YES — Q&A pipeline plan substantively reframed |
8.2 Audit verdict on Special Check E
Section titled “8.2 Audit verdict on Special Check E”PASS with 1 minor citation off-by-section error: synthesis §1.5 cites “S16 §10” for the effort estimate but the actual location is S16 §12.1 (S16 §10 is the migration plan, not the effort summary). Per A.M5.
9. Reframings audit (Liam-flagged Special Check F)
Section titled “9. Reframings audit (Liam-flagged Special Check F)”Subject: §4.1-§4.6 reframe earlier positions. Verify each reframing is supported by spike evidence (not author opinion).
9.1 §4.1 “LMDB blocks multi-worker Cloud Run” — DISSOLVED
Section titled “9.1 §4.1 “LMDB blocks multi-worker Cloud Run” — DISSOLVED”Spike support: S14 §1 (LMDB single-writer NOT hard-blocked; 10/10 concurrent processes OK) + S14 §6.1 (single-orchestrator topology recommended).
Verbatim from S14: “S14 §1 line 25: ‘PASSED — S2’s “single-writer constrains multi-worker Cloud Run topology” framing is REVISED.’”
Reframing verdict: EVIDENCE-SUPPORTED. Strong.
9.2 §4.2 “Mempalace search broken upstream” — STALE (FIXED in v3.3.5)
Section titled “9.2 §4.2 “Mempalace search broken upstream” — STALE (FIXED in v3.3.5)”Spike support: S15 §1 + §4 (live empirical verification on both /tmp + live ~/.mempalace/ palaces) + S3 footnote updates.
Reframing verdict: EVIDENCE-SUPPORTED. Strong.
9.3 §4.3 “Mempalace v5 PG backend is now available” — FALSE (never shipped)
Section titled “9.3 §4.3 “Mempalace v5 PG backend is now available” — FALSE (never shipped)”Spike support: S15 §1 + §3.1 (GitHub release inventory) + §3.2 (PyPI versions). Latest is v3.3.5; no v4 or v5.
Reframing verdict: EVIDENCE-SUPPORTED. The synthesis is self-aware about the original error.
9.4 §4.4 “Cocoindex content-hash solves DRAFT-vs-final” — FALSE per S2; HYBRID needed per S10
Section titled “9.4 §4.4 “Cocoindex content-hash solves DRAFT-vs-final” — FALSE per S2; HYBRID needed per S10”Spike support: S2 §2.5 + §4.3 (DRAFT-vs-final pairs have fully-distinct fingerprints; cocoindex doesn’t solve UC8) + S10 §1 + §2.7 (hybrid required).
Reframing verdict: EVIDENCE-SUPPORTED. Strong.
9.5 §4.5 “S1 gates Phase 2 on a multi-day spike” — DISSOLVED to ~30 minutes
Section titled “9.5 §4.5 “S1 gates Phase 2 on a multi-day spike” — DISSOLVED to ~30 minutes”Spike support: S1 §6 + §3.1 (source-code analysis sufficed; live test deferred).
Reframing verdict: EVIDENCE-SUPPORTED on the empirical fact. However, the reframing’s framing (“the spike-plan budget of 3-5 days was the worst-case-Scenario-B contingency”) is interpretive — the spike-plan §S1 lines 96-105 actually budget 3-5 days for the WHOLE spike including hands-on staging-branch testing under either Scenario A OR B. The author’s framing “worst-case contingency” is a post-hoc reinterpretation, but the actual outcome (~30 min for Scenario A confirmation from source) is solid.
Reframing verdict: EVIDENCE-SUPPORTED on outcome; INTERPRETIVE on framing. Acceptable.
9.6 §4.6 “Q&A schema is locked into existing q_a_extractions” — FALSE, pre-v1 redesign window open
Section titled “9.6 §4.6 “Q&A schema is locked into existing q_a_extractions” — FALSE, pre-v1 redesign window open”Spike support: S16 §1 + §6 + §7 (KH controls canonical Q&A shape pre-v1; recommendation = q_a_pairs + q_a_extractions two-tier model).
Reframing verdict: EVIDENCE-SUPPORTED. Strong.
9.7 Reframings audit summary
Section titled “9.7 Reframings audit summary”All 6 reframings are evidence-supported. §4.3 is exceptional (the synthesis admits author error explicitly). §4.5’s framing carries one interpretive layer (post-hoc “worst-case contingency” label on the spike-plan budget) but the underlying empirical outcome is solid. No reframing rests on author opinion alone.
Verdict on Special Check F: PASS.
10. Recommended edits (line-level changes before treating synthesis as load-bearing)
Section titled “10. Recommended edits (line-level changes before treating synthesis as load-bearing)”Below are the specific edits Liam should make to 0.9-synthesis.md to address the audit findings. Severity matches §2/§3.
10.1 Critical (must-fix before treating §1’s 92% confidence as load-bearing)
Section titled “10.1 Critical (must-fix before treating §1’s 92% confidence as load-bearing)”- §1 line 39 — fix 78% → 87% baseline anchor:
- From:
**Overall confidence: 92%** (up from 0.9-spike-plan baseline 78%). - To:
**Overall confidence: 92%** (up from 0.9-spike-plan §8 baseline 87%).
- From:
10.2 High severity (load-bearing before downstream agent handoff)
Section titled “10.2 High severity (load-bearing before downstream agent handoff)”-
§5 line 149 — remove “verify” parenthetical (already landed):
- From:
S14 §6 (already landed by S14 sub-agent — verify) - To:
S14 §6 (already landed —0.9-intended-architecture.md§10.1 line 1336 + §10 line 1287)
- From:
-
§6 — add leading paragraph before §6.1 (Liam’s flagged concern A directly):
These week-count estimates inherit from `0.8.2-cocoindex-evaluation.md` §sec 343-355(traditional dev-week breakdown authored S214) carried forward through `0.8-synthesis.md`+ `0.9-intended-architecture.md` + this doc. Per Liam's `0.8-synthesis.md` line 63caution and S230-start framing, these figures are AI-pair-programming-naive estimatesand should be treated as directional bounds for relative ordering of streams, NOTcalibrated commitments for scheduling. The spike-phase velocity anchor (8 parallelspikes in S229; 5 in S230) suggests the actual Phase 2 timeline could land anywherefrom 30% under to 30% over these figures, with the asymmetric risk on the upside(specifically: Q&A LLM-extraction prompt engineering + per-tenant rule curation arenovel; their estimates carry the largest uncertainty). -
§11 confidence table — add aggregation method:
- Add a line after the table at line 305:
Aggregation method: min-of-gate-confidences for the load-bearing gates (G1=96%, G2=92%, G3=90%, G4=85%, G10=85%, G14=90% → minimum 85%); the 92% headline is author-weighted by gate severity (G1 + G14 most load-bearing, hence weighted higher).OR remove the 92% figure entirely and state per-gate confidence only.
- Add a line after the table at line 305:
-
§3 line 70-89 — annotate synthesis-derived confidences:
- For S5 / S6 / S13 rows: add
(synthesis-derived; no headline confidence in spike report)after the 95% / 95% / 88% figures. - For S2 row: add
(synthesis-derived; S2 §10 has no headline confidence)after the 92% figure.
- For S5 / S6 / S13 rows: add
10.3 Medium severity (clarity-and-precision improvements)
Section titled “10.3 Medium severity (clarity-and-precision improvements)”-
§1 line 20 — clarify HYBRID operating-point range:
- From:
Intersection of cocoindex + skill-seekers achieves 91.7%/1.5-4.5%, passing both gates. - To:
Intersection of cocoindex + skill-seekers achieves TP=91.7% at FP rates between 1.5% and 4.5% depending on threshold selection (per S10 §2.7 operating-point sweep), passing both gates at multiple points.
- From:
-
§1 line 24 — change citation S16 §10 → S16 §12.1:
- From:
(S16 §1 + §12.) - To:
(S16 §1 + §12.1.)
- From:
-
§1 line 37 — change “S16 §10” → “S16 §12.1”:
- From:
~8-10 days wall-clock concurrent per S16 §10 - To:
~8-10 days wall-clock concurrent per S16 §12.1
- From:
-
§3 second-to-last paragraph (line ~99) — fix S7/S8/S9 framing:
- From:
The remaining spikes (S8 Q&A flow validation + S9 write-back semantics) are **post-architecture-commit** spikes — they validate the v1 implementation against canonical corpus, not the architectural choice. They run in Phase 2 itself, not before. - To:
The remaining S230-deferred spikes are S7 (pre-re-ingest evaluation — dispatched pre-Phase-2-first-real-re-ingest), S8 (Q&A flow validation), and S9 (write-back semantics). S8 + S9 are post-architecture-commit; S7 is calibration-not-blocking. None gate the Phase 2 commit.
- From:
-
§5 line 152 — soften OQ4 update:
- From:
§16 (Open Questions table) | Update OQ4 mempalace status: "ADOPTION-CONFIRMED via dev workflow; v3.3.5 search now working" - To:
§16 (Open Questions table) | Update OQ4 mempalace status: "ADOPTION-CONFIRMED for Lens 1 dev workflow (S3 + S15 verified); Lens 2 memory replacement still pending post-launch validation; v3.3.5 search now working"
- From:
-
§5 line 153 — reconcile new-OQ-row with §7.1:
- Either add a row at top of §7.1 (Outstanding work — Pending Liam review) for
OQ-NEW (S16): Two-tier q_a_pairs + q_a_extractions schema model — ratification at next decision-graph rewrite passOR delete the §5 line “Add new OQ: Q&A schema two-tier model” since it’s already implicit in S16 OQ1-OQ10.
- Either add a row at top of §7.1 (Outstanding work — Pending Liam review) for
-
§4.3 line 117 — fix section reference:
- From:
WP3 logged ERRATA-1 in0.9-spike-S10-dedup-substrate.md §10. - To:
WP3 logged ERRATA-1 in0.9-spike-S10-dedup-substrate.md §7.4.
- From:
-
§1 line 18 — soften “no measurable serialisation overhead”:
- From:
with no lock errors, no corruption, and no measurable serialisation overhead. - To:
with no lock errors, no corruption, and small measurable serialisation overhead (1.5× wall-clock for 10× concurrency per S14 §2.4).
- From:
10.4 Low severity (documentation hygiene)
Section titled “10.4 Low severity (documentation hygiene)”-
§6.1 line 176 — reconcile dedup hybrid effort:
- From:
**~5-7 days hybrid + 1-2 days per-tenant — S10 NEW** - To:
**~4-6 days hybrid + 1-2 days per-tenant — S10 §10 NEW; v1 substrate-impl budget per S10 §5**(matches S10 §10 4-6d framing rather than upshifted 5-7d).
- From:
-
§1 finding 1 — note 75 vs 70 column discrepancy:
- Add footnote or parenthetical: “S1 §1 cites 75-column
content_itemstable;0.9-decision-graph.mdQ1.1 +0.9-intended-architecture.md§1.2 cite 70-column. Reconcile via\d content_itemson staging at Phase 2 day-1.”
- Add footnote or parenthetical: “S1 §1 cites 75-column
11. Audit confidence
Section titled “11. Audit confidence”Verification audit confidence: 92%.
- High confidence on the load-bearing G1/G2/G3/G4/G10/G14 gate audits (verified verbatim against source spikes; calibrated against per-spike confidence figures).
- High confidence on the §4 reframings audit (all 6 evidence-supported; §4.3 self-aware about the v5 hallucination).
- High confidence on the cross-reference audit (1 minor citation off-by-section).
- Medium confidence on the timeframe audit — the 6-8 weeks anchor problem is real but the severity of the issue depends on how Liam weighs “directional bound vs calibrated commitment” — I’m flagging it as Liam-flagged-load-bearing per his S230-start framing, but a downstream reader might disagree on the severity.
- High confidence on the source-claim audit — no additional hallucinated claims beyond the §4.3-admitted one.
Remaining drag (8% confidence reservation):
- I did not run any spike artefacts (e.g., S1 harness, S10 probes, S14 topology probes) to independently verify the spike claims. The audit trusts the spike reports as accurate against their own methods.
- The “75 vs 70 column” discrepancy (A.M2) is unresolved — S1 says 75; decision-graph says 70. I lean toward 70 as the canonical figure (more recent, in 3 cross-referenced docs) but I haven’t verified via
\d content_itemsdirectly. - The “92% aggregate” confidence figure in synthesis §11 is opaque — without knowing the aggregation method, the audit’s view of whether it’s correctly calibrated is necessarily inferential.
12. Final verdict
Section titled “12. Final verdict”PASS-WITH-NOTES.
Phase 2 commit recommendation is sound. All load-bearing gates resolved positively; verbatim verdicts match source spikes; architecture revisions trace to spike outcomes; reframings are evidence-supported.
Liam should apply the 5 high-severity edits in §10.2 before treating the §1 “92% confidence” figure or §6 timeframe estimates as load-bearing for scheduling. The critical edit (§10.1 — fix 78% → 87% anchor) is a one-line fix; the high-severity edits address the timeframe anchor problem flagged at S230-start.
Liam should apply the medium-severity edits in §10.3 before passing the synthesis to a downstream agent for handoff. They are precision-and-clarity improvements that prevent downstream paraphrase drift.
Low-severity edits in §10.4 are documentation hygiene; defer if budget-constrained.
The 75 vs 70 column discrepancy (A.M2) should be resolved via a \d content_items query on staging at Phase 2 day-1; not blocking the Phase 2 commit decision itself.
End of verification audit. 1 critical + 4 high + 7 medium + 4 low = 16 findings. PASS-WITH-NOTES.