id-396 RESEARCH — corpus inventory across the three fixture surfaces
{396.1-research} Corpus inventory — empirical ground (S511)
Section titled “{396.1-research} Corpus inventory — empirical ground (S511)”Produced by a read-only inventory agent (S511) + the S507 content-echo report and S440 R1–R8 anchors. TECH.md consumes this; do not re-derive.
Headline correction to the S507 framing
Section titled “Headline correction to the S507 framing”The three “structurally-unsatisfiable” tests are no longer unsatisfiable. Commit
10120201 (id-389, PR #146) landed the distinct-bytes rework S507 §8 item 2 recommended:
two byte-distinct md fixtures at docs/testing/test-data/entity-variants/
(certification-variant-space.md = “ISO 27001”, certification-variant-nospace.md =
“ISO27001”), and cross-document-dedup, stage5-canonical-name-freshness,
cross-workspace-isolation were retargeted onto them in the same commit. Each staging
lands its own source_documents row; expect(items.length).toBeGreaterThanOrEqual(2)
(cross-document-dedup:98, freshness:130) is satisfiable. The fixtures’ docstrings name
their three consuming tests.
S507 §8 item 1 (rel_path-seeded em/er PKs — the F4 gap, id-398) is unaddressed and applies to a different population: the same-bytes re-staging tests below, plus every run-over-run re-staging (census #40’s 423 UniqueViolations).
Surface 1 — cocoindex-nightly forms corpus
Section titled “Surface 1 — cocoindex-nightly forms corpus”Staging chain (.github/workflows/cocoindex-nightly.yml): :624 docker cp
docs/testing/. → :634 verify_driver --fixtures templates (loopback POST /stage) →
:648 POST /walk → :709 background walk pump (30s re-POST) → :746 Vitest.
docs/testing/test-data/ — 11 files, all byte-distinct (SHA-1 verified), no
byte-identical duplicates:
| Fixture | Format | Consumers |
|---|---|---|
templates/csp-cloud-security-principles/Cloud Security Principles Checklist V5_3.xlsx | xlsx | 16 tests (Inv-4/5/6/8/9/10/11/12/17/20 families, stage-5 suite, pair-resolver-determinism, unresolved-mention-retains-canonical) |
templates/itt-services-efa/evaluation-matrix-itt-vol8.xlsx | xlsx | extract-memoisation; verify_driver templates (VERIFY-ITT-EFA) |
templates/sq-standard-selection-questionnaire/standard-selection-questionnaire-ppn-03-24.pdf | sidecar-cold-start, sidecar-mime-coverage; verify_driver (VERIFY-SQ-SSQ) | |
templates/rfp-british-council/annex_2_supplier_response.docx | docx | sidecar-mime-coverage; verify_driver (VERIFY-RFP-BC) |
templates/rfp-british-council/annex_3_pricing_approach.xlsx | xlsx | sidecar-mime-coverage; scripts/tests/test_form_extractors.py |
templates/rfp-british-council/rfp_-_learning_partners_osch.doc | .doc | extract-contract-honour (entity_mention branch) |
templates/itt-services-charnwood/ITT Services.docx | docx | extract-contract-honour (q_a_form branch) |
templates/itt-services-charnwood/ITT Evaluation Matrix.xls | .xls | ORPHAN — zero references |
templates/rfp-british-council/rfp_onlinetdcops.doc | .doc | ORPHAN — zero references |
entity-variants/certification-variant-space.md | md | the three dedup/freshness/isolation tests (W1) |
entity-variants/certification-variant-nospace.md | md | same three (W2) |
Second tree: __tests__/fixtures/cocoindex-chunking/ — short-clause.md (756 B;
Inv-2/13/15/16, sidecar-mime-coverage, sidecar-version-metadata) and long-terms.md
(5,120 B; chunking C-10/C-11/C-13/C-31, stage-topology).
Staleness: workflow :202 says “~1.5MB / 8 files” — it is 11.
verify_driver.py:97 defines only ONE fixture set (templates, 3 tuples); the other 8
fixtures reach the corpus per-test via stageFixture (hence the walk pump).
Same-bytes re-staging population (stageFixture does no in-byte injection —
__tests__/integration/cocoindex/_helpers/fixture-staging.ts:104-110):
admin-merge-coexistence + op-id-scoping (RUN_A/RUN_B), per-doc-canonicalisation
(-reingest), pair-resolver-determinism (?fullReprocess=1), extract-memoisation +
stage-5-row-counter (same DEST_PATH ×2), stage-5-op-id-memo (×3), chunking C-31 (×2).
This population is what id-398’s PK-seed fix unblocks.
Retention: corpus dir ${RUNNER_TEMP}/cocoindex-state/corpus (:444) — ephemeral.
Staging DB never resets: no post-run cleanup step; 19 of 41 cocoindex specs never call
dropFixture (audit-log-shipping, extract-memoisation, extractor-version-cross-ref,
faiss-pin, file-change-detection, health-probe, latency-budget, memo-hit-pipeline-run,
merge-entities-typed-return, nested-corpus, no-partial-row-writes, non-pipeline-write,
persistent-failure-dlq, sidecar-cold-start, sidecar-mime-coverage,
sidecar-version-metadata, stage-topology, transient-retry, url-landing-set).
entity_pair_resolutions caches by (name_a, name_b, entity_type) across runs AND tiers
(S507 §8.4) — mock verdicts replay into later real-LLM runs.
Surface 2 — Platform synthetic corpus (two things share the name)
Section titled “Surface 2 — Platform synthetic corpus (two things share the name)”(a) Vendored in-repo tree — scripts/cocoindex_pipeline/fixtures/platform-corpus/,
10 files: content/ (8 incl. the 5 ID-132.30 grain files), qa/synthetic-qa-pairs.md,
edge/synthetic-sparse-edge.md. Sole consumer:
__tests__/integration/cocoindex/platform-corpus-shape.test.ts — runs in every
bun run test; read-only guard (exact 10-entry set, magic bytes, synthetic- prefix, no
forms/ tree per DR-014, no reserved __qa__/). Filenames are load-bearing:
scripts/cocoindex_pipeline/sources/l_records.py:94-99 filename gates for
case_study/company/certification grains. Deploy artefact ({134.4} BI-6). Never staged.
(b) ID-127 {127.4} design-stage corpus — never built as specced.
local-fs-platform/corpus/ does not exist; scripts/seed-synthetic-corpus.ts retired
(ID-145.25); the folder→workspace manifest premise retired wholesale (ID-127.37,
DR-038/056/061). Cite specs/id-127-platform-pipeline/notes/id127-corpus-structure-research.md
(§1 one-record/M2M principle, §3 OKF-aligned tree, §4.4 the one-record-many-views gap) and
specs/synthetic-platform-corpus/PROPOSAL.md (S462 banner: forms half DR-014-retired;
only the DB-seed/win-rate half survives).
Surface 3 — e2e / eval seeds
Section titled “Surface 3 — e2e / eval seeds”id-392’s S508 content_chunks retarget landed in exactly TWO places (both seed a
position-0 content_chunks row): scripts/seed-e2e-users.ts:376
(ensurePublicationReviewFixtureChunk, [E2E-PUB-REVIEW-FIXTURE], consumed by
e2e/tests/review-publication-tab.spec.ts; idempotent, backfills pre-retarget envs) and
e2e/tests/publication-bulk-action.e2e.spec.ts:226.
GAP: e2e/fixtures/test-data-fixture.ts — the worker-scoped fixture seeding the BULK
of e2e source_documents (:240, :516, :735) — writes no body at all (no
extracted_text, no content_chunks; placeholder mime/size/hash per {128.14}). Readers
composing via lib/source-documents/body.ts see nothing. The retarget covered the two
deterministic seeders and missed the highest-volume one.
Retention: per-worker delete-by-id cleanup (test-data-fixture.ts:809+);
e2e/global-teardown.ts is a safety sweep (title-prefix deletes; content_chunks rides
ON DELETE CASCADE, 20260628200000_id131_extract_reparent.sql). [E2E-PUB-REVIEW-FIXTURE]
is declared never-deleted (documented cross-shard race).
Eval seeds never touch content_chunks: lib/eval/fixtures.ts → 4 public gold
standards in __tests__/fixtures/eval-gold/; PRIVATE_FIXTURES empty post-{114.8};
baselines in __tests__/fixtures/eval-baselines/; results accumulate in eval_runs.
eval-nightly.yml:64 forbids mock tier.
Anchors
Section titled “Anchors”- R1–R8: S441 ratification, text in
specs/id-138-corpus-durable-home/notes/s440-corpus-durable-home-decision.md:130-175— esp. R4 (ingest-once vs keep-and-watch, assigned at the binding gate), R3/DR-026 (promotion boundary), item 7 (“what survives the engine” contract per record class). - DR-025 (corpus reframe), DR-023 (durable home), DR-026, DR-093 (no backfill — delete), DR-014 (forms manual-upload), DR-024 (identity: content_hash-first, rel_path mutable), DR-050, DR-036, DR-032.
- Content-echo report:
specs/id-389-llm-tier/notes/S507-content-echo-report.md§5 §6 §8. - F4 live proof: census #40 (id-128 journal, run 30453466563) — 423 UniqueViolations, all
content-hash-derived mock names vs non-NULL source_document_id under rel_path-seeded PKs
(
flow.py:2393-2396,2467-2470; self-flaggedflow.py:3178-3181;ingest_oncealready seeds correctly:3347-3350,:3389-3391). - id-370 owner rulings: S494 data posture (all synthetic, no backfill, delete freely);
S503 observation (
reference_items.source_document_id— live FK vs residue, adjudicate).
Three structural findings for the model
Section titled “Three structural findings for the model”- Three unrelated trees, no shared contract — only
platform-corpus/has a rot guard; the other two drift silently; two fixtures already orphaned. - Cleanup is opt-in and half-adopted — 19/41 specs never clean up; nightly adds no
sweep; the shared staging DB is the accumulation point; cross-tier
entity_pair_resolutionscache compounds it. - Body-seeding retarget incomplete —
test-data-fixture.tsproduces body-less rows at the highest volume.