Skip to content

ID-114 {114.1} RESEARCH — relocate product functionality out of the private docs-site

ID-114 {114.1} RESEARCH — relocate product functionality out of the private docs-site (de-ID + re-home)

Section titled “ID-114 {114.1} RESEARCH — relocate product functionality out of the private docs-site (de-ID + re-home)”

Date: 15/06/2026 Author: Task Planner (fresh dispatch; {114.1} RESEARCH only) Tier: RESEARCH (domain-complexity warranted — cross-repo coupling, live-runtime dependency, client-identity de-ID, ID-68/ID-95 ownership entanglement). Recommend the parent Task open as PRODUCT+PLAN or TECH+PLAN tier (de-ID + relocation is behaviourally-and-mechanically specified; the per-item homes here are near-spec-ready). Promoted from: bl-322. Not flip-blocking — see §7.

Governing principle (verbatim, prior audit 2026-06-15): “the private docs-site must host docs ONLY — no product functionality.” This Task closes the gaps where product / coupled content currently sits in the private docs-site (knowledge-hub-docs-site) as an ID-68 Option S-B (sensitivity-primary) split workaround.


1. Context — what this RESEARCH does NOT re-derive

Section titled “1. Context — what this RESEARCH does NOT re-derive”

Per the dispatch brief, this item “has been raised several times.” A prior session (2026-06-15, sub-agent agent-a1bf3ef5ebdd2d8f3, parent session cf00f896-8599-4927-8e57-e500594a927f) already produced a complete “Docs-Site Product-Content Audit + Relocation Map” (14 rows, per-item provenance + commit hashes + recommended homes). This RESEARCH adopts that map as its baseline and corrects it against the current tree state (the map is ~1 week old; two of its premises have since shifted). It does NOT re-audit from scratch.

The prior map is corroborated by:

  • RELOCATION-STATUS.md (docs-site root) — the canonical ID-68 {68.32} layout doc.
  • The ID-95 routing appendix (per {95.2} PRODUCT / {95.3} TECH, docs-site 40830d51), which named only four items (harness eval imports → ID-68; docs-site root vitest → docs-site repo; MCP-plugin seven-κ → ID-71; PI-9 denylist-lint residual → ID-68) and left the rest uncovered.

1.1 Prereq-1 disposition — classification-prompt.md prior decisions (REQUIRED review, done)

Section titled “1.1 Prereq-1 disposition — classification-prompt.md prior decisions (REQUIRED review, done)”
  • decision-graph.md (themes/canonical-pipeline/reference/decision-graph.md) and 07-collapse-list.md (themes/canonical-pipeline/intended-architecture/) were read read-only. Finding: these are the canonical-pipeline collapse register (schema retires/renames, Q&A restructuring, CV-driven retires). They carry NO disposition for classification-prompt.md relocation — that file is not a collapse-list row. They are context, not the prior decision for this Task.
  • MemPalace (searched WITHOUT wing filter per the known wing-filter bug; filtered client-side): the strongest hit is the prior audit map (§1). Its verbatim closing line: “MemPalace corroborates the ID-68/S-B/PC-7/PC-31 provenance strongly for harness/eval-fixtures/denylist; it has no hits on the classification-prompt.md relocation. So: no prior decision exists for a skeleton/tenant_config split of classification-prompt.md. The only recorded disposition (prior-audit Row 7) is “GAP — highest concern. Relocate to MAIN repo (de-ID’d worked-examples) … owning task ID-68.” The skeleton + tenant_config.classificationDisambiguation split named in the dispatch brief is a NEW design — and, critically, §3 below shows it is already largely implemented under ID-95.5.

1.2 Prereq-2 disposition — eval-fixtures wider eval work (REQUIRED review, done)

Section titled “1.2 Prereq-2 disposition — eval-fixtures wider eval work (REQUIRED review, done)”

Ledger slice-reads (bun scripts/ledger-cli.ts get task <id> …):

  • ID-104 (eval-engine, in_progress): bottom-up agent-eval engine — “the long-term one” per Liam S354. {104.4} PLAN ratified S357 → 16 impl Subtasks {104.5}{104.20}. W1 migrations staging-first; W3 runner shares functional-correctness.ts with {71.14}. ID-104 OWNS the AgentEvalContract, eval-runner, nightly lane, Claude-as-judge rubric. It is a clean-sheet rebuild (NOT lift-and-extend of legacy L1/L3/L4 + lib/eval/baseline.ts).
  • ID-71 (AI-tooling, in_progress): rationalises the AI-consumption layer. Lane {71.14} = “fixtures + evals + inventory” (rename strategy, lockstep with plugin bundle + fixtures + client-guide refresh). W2/W3 gated on the ID-104 contract.

Implication for the eval-fixtures home: the gold-standards are active eval inputs, not inert files. Their de-ID + re-home must not collide with ID-104’s runner build or ID-71 {71.14}’s fixture rename. Recommended sequencing in §5 routes the physical relocation + de-ID to this Task (ID-114) but defers the consuming-path rewiring to land in lockstep with whichever of ID-104/{71.14} touches the fixtures next. This is an OQ for Liam (§8 OQ-3): own the fixture move in ID-114, or fold it into {71.14}?

1.3 Code-intelligence orientation (REQUIRED — cited verbatim)

Section titled “1.3 Code-intelligence orientation (REQUIRED — cited verbatim)”
  • gitnexus_query({query: 'classification prompt classify pipeline runtime load', repo: 'knowledge-hub'}) → returned NO execution processes ("processes": []); only test-file definitions (scripts/tests/test_cocoindex_prompts.py, test_cocoindex_extractors.py) and lib/intelligence/pipeline.ts:getActivePrompt / runPipeline. The classifier-prompt load path is not a graphed runtime flow — it is build-time codegen (see §3).
  • classify.py is Python and is GONE. grep sweep (gitnexus/ast-dataflow cover TypeScript only): scripts/kb_pipeline/ is an empty directory (only __pycache__/ + .DS_Store; no .py source). grep -rn classification-prompt scripts --include=*.py finds only a test assertion. The dispatch brief’s premise — “scripts/kb_pipeline/classify.py loads ops/classification-prompt.md at RUNTIME” — is STALE. The classifier moved to TypeScript (lib/ai/classify.ts); see §3 for the corrected coupling.
  • The actual consumers of ops/classification-prompt.md (grep, TS): scripts/generate-classification-prompt-taxonomy.ts, scripts/bundle-plugin.ts, scripts/lib/taxonomy-parser.ts — all build-time, all resolving the path via resolvePrivateDocsDir() / 'ops' / 'classification-prompt.md' (the KH_PRIVATE_DOCS_DIR bridge, fail-loud, Inv 28/29, relocated {68.23} e00c522f).

2. Per-item relocation table (item → proper home → de-ID requirement → owning Task)

Section titled “2. Per-item relocation table (item → proper home → de-ID requirement → owning Task)”

Adopts the prior-audit 14-row map; rows 2 and 7 corrected against current tree state. Paths are docs-site-root-relative (the audit targets live at the repo root, NOT under src/content/docs/).

#Item (docs-site path)What it isClient-ID?Proper home + actionDe-ID requirementOwning Task
1harness/ (52 files; knowledge-hub-internal pkg)Private vitest/eval product test infra; own package.json/tsconfig/vitest.config; CI runs itIndirect (via fixtures)Relocate to MAIN __tests__/eval/ (private-gated). Product test infra, not docsNone directly (fixtures = rows 4/6)ID-68 (extend) or ID-114
2harness/scripts/eval-{classification,entity-classification}.ts + harness/lib/ fork (anthropic.ts, ai/pricing.ts, eval/*)Eval runners + forked copies of main libsNoRelocate to MAIN + DELETE the fork → imports resolve to canonical @/lib/*. Fixes latent breakNoneID-68 or ID-114
3eval-fixtures/*.json (root: procurement-drafting + summarisation, ~1171 lines)Eval gold-standardsYES — saturated (“Phew Design Ltd” ×20+, “Telehouse South”, ICO/ISO/CREST prose, real KB UUIDs)Relocate to MAIN private-gated fixtures AFTER de-ID (or de-ID in place if consumed-path rewiring deferred)HARD — verbatim client prose + entity names + real content UUIDs → Example Client Limited swapID-114 (move/de-ID); consuming rewire → coordinate ID-104 / {71.14} (OQ-3)
4harness/__tests__/fixtures/*.json + eval-baselines/*.baseline.jsonClassification/entity gold-standards + metric baselinesLikely (labelled from real client content; UUIDs)Relocate to MAIN as private-gated fixtures (with row 1)De-ID review per Inv 7ID-68 or ID-114
5ops/classification-prompt.md (1877 lines — not ~881)Build-time codegen source + reference doc; TAXONOMY block (L33–255) injected to plugin; NOT the runtime classifier prompt (see §3)YES — Phew worked-examples L194–809 (Phew Design Ltd, Phew Audit System, CREST/DBS prose)SPLIT (already mostly done — §3): (a) TAXONOMY block is DB-generated → keep codegen source in MAIN, de-ID’d; (b) client worked-examples → tenant_config.classificationDisambiguation (ID-95.5, already modelled); (c) retire the docs-site copy + the KH_PRIVATE_DOCS_DIR build-time bridgeHARD — strip Phew worked-examples; replace with {CLIENT_*} placeholders or move to tenant_configID-114 (de-ID + retire bridge); validate vs ID-95 §3
6.config/ip-denylist.txt (381 B; “phew”)2nd denylist → MAIN hook ip-leak-filename-guard.shYESConsolidate with row 7; proper home = private secret/config store, not docs-siten/a (it IS the denylist)ID-68/ID-114
7ops/identity-denylist.json (PC-31, S321-ratified incl. ICO)Canonical client-name denylist; synced → public KH_CLIENT_NAME_DENYLIST secret + sweep scriptsYES by definitionBORDERLINE — legitimately private; but docs-site ≠ governance-config host. Interim KEEP acceptable; proper home = dedicated private secret/config storen/a (it IS the denylist)ID-68/ID-114
8scripts/check-token-parity.tsReads ../../app/globals.css (MAIN path) vs docs-site CSS — cross-repo path couplingNoTENSION — resolved in §6. Recommend: relocate to MAIN OR a shared composite action (NOT keep as-is). The “keep” lean (owner correction) is reconcilable only by re-homing it so the coupling points into docs-site, not outNoneID-114 (or shared action)
9src/content/docs/workflow-evaluation/sessions/** (41 dirs)Workflow-evaluator runtime telemetry (events.jsonl/final_report.yaml/meta.json — embeds MAIN paths)Indirect (infra paths)KEEP (owner correction — OUR dev-workflow-evaluation content). It is ID-48 data; not docs-corpus but legitimately archived hereOptional: scrub meta.json absolute MAIN pathsKEEP (ID-48 owns the data)
10.claude/{workflows,agents/workflow-evaluator,skills/evaluate-*}Workflow-evaluator operators (read-only, no MAIN import)NoKEEP — operate ON the docs-site corpus; properly homedNoneKEEP
11Root __tests__/ (10 tests inc. ledgers-integrity) + vitest.config.tsDocs-site’s OWN tests — unwired in CI (ci.yml runs harness/ vitest only)NoKEEP, but WIRE INTO docs-site CINoneKEEP (docs-site repo)
12.github/workflows/{docubot,sync-source-docs}.yml + docubot action + harness/scripts/{docubot,skills}/*Docubot lane — generates docs FROM public PRs (cross-repo by design)NoKEEP — the one class of “functionality” properly homed in docs-siteNoneKEEP

Counts the table touches: RELOCATE → rows 1–6, 8 (de-ID required on 3, 5, and the two denylists by-definition). KEEP → rows 9–12.


3. classification-prompt.md split design (validated vs prereq-1)

Section titled “3. classification-prompt.md split design (validated vs prereq-1)”

The brief’s recommended split is ALREADY substantially implemented under ID-95.5. The prior-audit Row 7 (“relocate wholesale to MAIN, de-ID worked-examples”) is superseded by what ID-95 already shipped. Verified facts:

  1. The runtime classifier prompt is NOT ops/classification-prompt.md. It is lib/ai/skills/classification.md (MAIN repo, 53 KB), loaded by lib/ai/classify.ts:1177 via loadSkill('classification'). The docs-site file itself says so (L11–14): “the [classifier] does not read this file. The content between TAXONOMY_START and TAXONOMY_END markers is used by the plugin sync scripts.”
  2. The skeleton already uses placeholders. lib/ai/skills/classification.md carries {TAXONOMY} (L20), {CLIENT_DISAMBIGUATION} (L283), {CLIENT_PRODUCT_NAME}, {CLIENT_ORGANISATION_NAME}, {CLIENT_ORGANISATION_SHORT}, {CLIENT_PRODUCT_SHORT}. classify.ts:1183–1213 resolves them via a .replaceAll chain + buildDisambiguationBlock().
  3. The per-client data home already exists. classification_disambiguation_rules is modelled in lib/client-config.ts (BrandingConfigSchema.classificationDisambiguation: entityExamples[] + selfReferenceRules[]), seeded in lib/branding/clients/phew.json + default.json, and lives in DB tenant_config.config (migration 20260613090000_id95_5_tenant_config.sql — single-row, service-role-only, set out-of-band via re-seed manifest, NEVER committed, PI-10). This is exactly the brand→DB precedent the brief cites.
  4. What ops/classification-prompt.md actually still does: build-time only. (a) its TAXONOMY_START…END block is the read source AND codegen target for generate-classification-prompt-taxonomy.ts (DB taxonomy → injected back); (b) the same block is the canonical taxonomy bundle-plugin.ts:validate() checks the plugin SKILL.md against; (c) the rest of the file (rules + Phew worked-examples L194–809) is now a reference duplicate of the skeleton + tenant_config — stale and client-identity-bearing.

Recommended split design for ID-114 (validated, corrected):

  • (S1) De-ID the worked-examples region (L194–809): replace hardcoded Phew Design Limited / Phew Audit System etc. with the {CLIENT_*} placeholders the skeleton already uses, OR delete them (they duplicate tenant_config rules). HARD de-ID gate.
  • (S2) Decide the canonical home of the TAXONOMY codegen source. Options: (a) keep ops/classification-prompt.md purely as the codegen scratch artefact but move it into MAIN so the build-time KH_PRIVATE_DOCS_DIR bridge (bundle-plugin.ts:173, generate-classification-prompt-taxonomy.ts:27) is retired — removing docs-site as load-bearing build infra; OR (b) regenerate the taxonomy block directly from DB at build time with no markdown intermediary (the taxonomy is already DB-canonical). (b) is the cleaner end-state and aligns with the “docs-site hosts docs only” principle.
  • (S3) NO new runtime work — the runtime skeleton + tenant_config path is done. ID-114 is a de-ID + bridge-retirement task here, not a re-architecture.

Divergence flagged (OQ-1): the brief frames this as a fresh split into tenant_config.classificationDisambiguation; that home already exists and is wired. The real residual is de-ID + retiring the docs-site build-time bridge. Confirm with Liam that ID-114’s scope on this item is de-ID + bridge retirement, not duplicating ID-95.5.


4. eval-fixtures home (validated vs prereq-2)

Section titled “4. eval-fixtures home (validated vs prereq-2)”
  • Prior-audit decision (Row 6) + bottom-line: root eval-fixtures/*.jsonMAIN repo private-gated fixtures, after de-ID, owning ID-68.
  • Current-tree confirmation: both files verified client-identity-saturated (head of procurement-drafting-eval-gold-standard.json: “Phew Design Ltd”, ICO/DPO/GDPR prose, ISO 27001 + Cyber Essentials Plus, real expected_kb_items_used UUIDs). HARD de-ID.
  • Wider-eval coordination (prereq-2): these are live eval inputs. ID-104 is building the eval-runner (W3 shares functional-correctness.ts with {71.14}); ID-71 {71.14} is the “fixtures + evals + inventory” lane. Recommended home: MAIN repo, private-gated (__tests__/eval/fixtures/ alongside the relocated harness), de-ID’d to Example Client Limited. Sequencing (OQ-3): physically move + de-ID under ID-114; rewire the consuming paths in lockstep with whichever of ID-104/{71.14} next touches them — do NOT rewire blind while those lanes are mid-flight. Real KB UUIDs in expected_kb_items_used may need re-keying to a de-ID’d seed corpus — flag to ID-104 (OQ-4).

5. Sequencing notes (ID-68-owned vs ID-95-owned vs new ID-114)

Section titled “5. Sequencing notes (ID-68-owned vs ID-95-owned vs new ID-114)”
  • ID-68 (repo visibility / IP separation): owns the original Option S-B split + de-ID discipline (Inv 7 gold-standard moves, denylist mechanics, identity-guard). The prior audit routed most rows to “ID-68 (extend remit) or a new harness re-home task” — ID-114 IS that new task. Recommend ID-114 owns the physical relocation + de-ID for rows 1–6, 8; ID-68 remains the policy owner (de-ID review gate).
  • ID-95 (per-client topology): owns tenant_config + the classificationDisambiguation model + the brand→DB pattern. ID-114 must not re-implement what ID-95.5 shipped (§3) — it consumes it.
  • ID-104 / ID-71 {71.14}: own the eval-runner + fixture-rename lanes. ID-114 sequences the eval-fixtures move to avoid colliding (§4, OQ-3).
  • Within ID-114 (proposed wave order):
    1. Denylists (rows 6, 7) — low-coupling consolidation decision; unblocks nothing else.
    2. Harness + fork delete + fixture move (rows 1, 2, 4, 6/eval-fixtures) — atomic; the fork-delete fixes the latent lib/ai/classify --live import break in the same move.
    3. classification-prompt.md de-ID + bridge retirement (row 5) — depends on the §3 S2 decision (OQ-2); de-ID is independent and can land first.
    4. check-token-parity re-home (row 8) — independent; resolve §6 first.
    • These are sibling-only within ID-114 — no cross-Task Subtask deps. Cross-Task coordination (ID-104/{71.14} for fixture rewire) is handled at Task level (OQ-3), not as a Subtask dependency.

6. The check-token-parity.ts relocate-vs-keep tension (resolved)

Section titled “6. The check-token-parity.ts relocate-vs-keep tension (resolved)”

The dispatch brief flags that scripts/check-token-parity.ts appears in both the relocate list (prior-audit Row 13: “Move to MAIN or a shared composite action”) and the keep list (owner correction: “OUR dev-workflow-evaluation content”). Resolution:

  • The script’s purpose is legitimately docs-site-adjacent (token-parity between the docs-site app’s CSS and MAIN app/globals.css) — so the owner’s “keep” instinct is right about the function.
  • But the current implementation is the workaround — it reads ../../app/globals.css, escaping the docs-site repo and assuming MAIN is the parent directory (runtime-fragile, same anti-pattern as the classification-prompt bridge).
  • Recommendation (resolves the tension): KEEP the token-parity function, but re-home the coupling. Either (a) move the guard to MAIN (where it reads MAIN CSS directly and fetches docs-site tokens via the existing GitHub-App-token bridge, same pattern as docubot/sync-source-docs), or (b) make it a shared composite GitHub action invoked from both repos. Do not leave the ../../ parent-dir escape. This honours both lists: keep the capability, relocate the cross-repo coupling out of a fragile relative path. OQ-5 for Liam: (a) MAIN-homed guard vs (b) shared composite action.

None of these relocations block the public-repo flip. The identity-guard required check scans the public knowledge-hub repo only — it does not scan the private docs-site. Client identity sitting in the private docs-site (ops/classification-prompt.md worked-examples, eval-fixtures/*.json, the denylists) is not a public-repo leak and does not trip the flip gate. This Task is hygiene + correctness + de-fragilising cross-repo coupling, not a flip prerequisite. (Caveat: any item that gets relocated into MAIN must be de-ID’d before it lands there, because at that point identity-guard will scan it — the de-ID gate is the relocation precondition, not a flip precondition.)


8. Open Questions for Liam (blocking marked [BLOCKING])

Section titled “8. Open Questions for Liam (blocking marked [BLOCKING])”
  • OQ-1 [BLOCKING]classification-prompt.md scope. §3 shows the skeleton + tenant_config.classificationDisambiguation split is already implemented (ID-95.5). Confirm ID-114’s scope on this item is (de-ID the worked-examples) + (retire the docs-site build-time bridge) — NOT re-building the split. (If Liam expected a net-new split, that’s a misconception to correct before PRODUCT/TECH.)
  • OQ-2 [BLOCKING]TAXONOMY codegen source end-state. §3 S2: keep ops/classification-prompt.md as a codegen artefact relocated into MAIN (bridge retired), OR regenerate the taxonomy block directly from DB with no markdown intermediary? The latter fully removes docs-site as build infra.
  • OQ-3 [BLOCKING]eval-fixtures consuming-path ownership. Move + de-ID the fixtures under ID-114, but who rewires the consuming paths — ID-114, or fold into ID-71 {71.14} / coordinate with ID-104’s runner build to avoid mid-flight collision?
  • OQ-4real KB UUIDs in expected_kb_items_used. Do these need re-keying to a de-ID’d seed corpus when the fixtures move to MAIN? (Coordinate ID-104.)
  • OQ-5check-token-parity re-home (§6): MAIN-homed guard vs shared composite action?
  • OQ-6denylist canonical home (rows 6/7): interim-keep-in-docs-site acceptable, or stand up a dedicated private secret/config store now? And: consolidate the two denylists (.config/ip-denylist.txt + ops/identity-denylist.json) into one source?
  • OQ-7ID-68 vs ID-114 ownership boundary: should ID-114 own the physical relocation/de-ID (with ID-68 as policy owner), or does this fold back into ID-68’s remit?

  • Prior audit map (verbatim, 14 rows) — MemPalace drawer agent-a1bf3ef5ebdd2d8f3 (2026-06-15), recovered full from transcript.
  • RELOCATION-STATUS.md (docs-site root) — ID-68 {68.32} layout.
  • Current tree: ops/classification-prompt.md (1877 L), eval-fixtures/*.json, harness/{lib,scripts,__tests__}, .config/ip-denylist.txt, ops/identity-denylist.json, scripts/check-token-parity.ts.
  • MAIN repo: lib/ai/classify.ts, lib/ai/skills/classification.md, lib/client-config.ts, scripts/{generate-classification-prompt-taxonomy,bundle-plugin}.ts, scripts/lib/taxonomy-parser.ts, supabase/migrations/20260613090000_id95_5_tenant_config.sql.
  • Ledger slice-reads: ID-104, ID-71 (descriptions + status_notes).
  • decision-graph.md + 07-collapse-list.md (prereq-1 read; no classification-prompt disposition).
  • gitnexus_query (repo: knowledge-hub) — no graphed runtime flow; grep for the Python load path (classify.py absent).