Skip to content

ID-71 {71.1} RESEARCH — AI tooling surface rationalisation

ID-71 {71.1} RESEARCH — AI tooling surface rationalisation

Section titled “ID-71 {71.1} RESEARCH — AI tooling surface rationalisation”

Status: {71.1} RESEARCH exit artefact. Folds Lane A (outcome/workflow research, ratified at the 2026-06-10 workshop) and Lane B (technical + ecosystem sweep) against the settled lane-a-workshop-outcomes.md decision record (WS-1…WS-14). Every WS-numbered decision is treated as settled input. Verdicts attach to outcomes, not tool names (WS-11 / SYNTHESIS §1.1). Date: 2026-06-10. British English throughout.

Exit contract: verdict table covers 100% of the current surface inventory (§2); target-surface options presented before {71.2} PRODUCT (§3). Open [LIAM] rows and the queued /idea-refine headless session are the only items deliberately left for PRODUCT (§7).


The substance lives in the cited input docs; this section folds, it does not re-narrate.

ID-71 is the umbrella over the whole AI-consumption layer: MCP tools + resources + prompts + MCP Apps + plugin commands + plugin skills + inline AI touchpoints (SYNTHESIS.md §1.1). The canonical surface counts are asserted by the inventory-parser integration test (__tests__/scripts/mcp-inventory-parser.test.ts — “verify 58 tools, 12 resources, 7 prompts”), parsed from real source under lib/mcp/. Counts are not hard-coded in prose here — the test is the source of truth, per the Tier-0 rule that counts drift everywhere (SYNTHESIS.md §1.4). The four MCP Apps’ trigger tools (show_coverage_matrix, show_procurement_dashboard, show_reorient_me, show_intelligence_feedlib/mcp/tools/apps.ts) are themselves counted inside the tool total; plugin commands = 8 (.claude/plugins/knowledge-hub/1.0.0/commands/) and plugin skills = 9 (.../skills/; the docs-vs-bundle disagreement in SYNTHESIS.md §1.1 resolves to 9 — bundle is truth).

Outcome ranking holds with O2 (revenue documents) above O3 (trust maintenance) (WS-1). O3 is largely engineered at the ingestion point by the cocoindex pivot (controlled local-fs source, automatic change pickup, ETL additions), so the tooling surface inherits trust from the pipeline rather than retrofitting it. Personas are hats, not heads (WS-3): Maya/David/Priya/Emma (...strawman.md §A) are the analytic lenses; marketing exists as a person at Phew, finance/commercial is a hat at Phew but a person at future clients (lane-a-workshop-feedback.md validation points).

The structural finding from the surface mapping (...strawman.md §D obs. 1): the 58 tools concentrate in curation (Content management ≈ half the surface) while the ranked outcomes concentrate in consumption and output (O1, O2, O4, O5). The surface is inverted relative to the outcome ranking. The two biggest gaps — O9 onboarding and the O5 consumption half — are also the two biggest differentiation claims (under-4-hours; signal-to-action), so gap-filling outranks consolidation in client-visible value (...strawman.md §D obs. 3).

1.3 Settled consolidation decisions feeding §2

Section titled “1.3 Settled consolidation decisions feeding §2”
  • WS-7 — single outcome-shaped find entry (type/scope/granularity parameterised) replaces the search trio + find_similar_items, unless workflow evidence argues otherwise for a named case.
  • WS-8 — exposure/queue is over-complicated today; collapse to two outcomes (“where are we exposed”, “what’s in my queue”) presented via the WS-4 layers framing, with resolution first-class (gaps/issues surface with suggested resolutions in-platform and via MCP App — “Draft content for X”, “Discuss options for Y”) and flexible AI involvement per task (background HA for sign-off, or two-way thinking-partner via Skills).
  • WS-9completing-forms is the carried-forward concept (procurement = first form type); W2.4 questionnaire-fill folds in.
  • WS-4 — headless-complete baseline at launch: O1/O4/O6 reads + W5.6 + W9.3, no publication gates; O4 widens beyond KH (reorient the person), O6 simplifies via the layers presentation. “Headless” includes managed agents; MCP + Skills consumed from Claude Desktop/claude.ai/Cowork is itself a headless-agent form (MCP = ability, Skill = system-prompt-like expertise). Ease of connectivity (incoming and outgoing) is a first-class headless requirement.
  • WS-11no hard numeric ceiling as design law; structured tool-justification + progressive utility as trust increases on quality metrics; surface shape follows the outcome architecture (many headless agents + few human entry points beats many human-facing tools).

1.4 Stale surface that must not be inherited (concept-level)

Section titled “1.4 Stale surface that must not be inherited (concept-level)”

Bid-domain naming/description drift is now fresh drift against the renamed DB (bid_responsesform_responses, get_bid_*get_form_*, cite_content arg → form_response_id; SYNTHESIS.md §1.4). get_bid_question, kb://bids/{id}, ui://bid-dashboard, the bid_briefing/bid_pipeline_review prompts, the bid-writing plugin skill and /kb:bid-* commands all carry MCP-layer naming drift to resolve under WS-9’s completing-forms generalisation. Phase-0 hygiene is done ({71.5} live on canonical via ID-103.1): list_user_workspaces now accepts legacy 'bid' as an alias for procurement (lib/mcp/tools/workspaces.ts), and eval:* package-script rot + extract-questions dead model id are fixed.

1.5 Lane B grounding (verified 2026-06-09; see §5)

Section titled “1.5 Lane B grounding (verified 2026-06-09; see §5)”

@anthropic-ai/sdk is on 0.96.0; structured outputs are GA via output_config.format; strict tool use (strict: true + recursive additionalProperties: false) is GA; citations × structured-outputs remain architecturally incompatible (verbatim 400-warning quoted in ...lane-b-structured-outputs-citations-currency.md §4), which is why drafting stays a 3-pass split. These are the standards seeds for §5. No re-run was needed for this docs-only RESEARCH: Lane B’s currency doc is VERIFIED-from-source against the pinned SDK and is cited rather than re-derived.


Every tool, resource, prompt, app, command, plugin skill, and inline AI touchpoint is mapped to its outcome/workflow (O1–O9 / W-numbering, ...strawman.md) with a concept-level verdict: keep-concept / refine / consolidate-into-X / retire / gap. Settled decisions (WS-7/8/9/4) are applied. [LIAM] marks the few rows where a verdict genuinely still needs Liam. Tool/app/prompt names locate the concept — the verdict attaches to the outcome.

Surface (concept)Outcome / workflowVerdict
search_knowledge_baseO1/O7 W1.1, W7.1consolidate-into find (WS-7)
search_qa_libraryO1 W1.2 prior-answer matchconsolidate-into find (WS-7; preserve corpus-level q_a_pairs + scope_tag semantics as a scope/type param)
search_content_chunksO1 W1.1 verbatim fetchconsolidate-into find (WS-7; the granularity param = chunk vs item)
find_similar_itemsO1 W1.x, dedup adjacencyconsolidate-into find (WS-7)
(grounded answering) W1.3 entity/ontology-grounded answerO1 W1.3gap — no deliberate affordance today (A21 ontology grounding); seed in §5

2.2 MCP tools — Dashboard & orientation / exposure

Section titled “2.2 MCP tools — Dashboard & orientation / exposure”
Surface (concept)Outcome / workflowVerdict
list_user_workspacessubstrate (resolve workspace)keep-concept, refine (alias fixed in {71.5}; description still bid-tinged)
get_dashboard_summaryO4 W4.1 “what needs me”consolidate-into what’s-in-my-queue (WS-8)
get_reorientationO4 W4.3 re-orientationkeep-concept, refine — widen beyond KH state (WS-4: reorient the person)
get_expiring_contentO6 W6.2 freshness/expiryconsolidate-into where-are-we-exposed (WS-8)
get_freshness_reportO6 W6.2consolidate-into where-are-we-exposed (WS-8)
get_coverage_gapsO6 W6.1 coverage-gapconsolidate-into where-are-we-exposed (WS-8; resolution loop attaches here)
get_certification_statusO6 W6.2 (Priya certs)consolidate-into where-are-we-exposed (WS-8)
suggest_content_creationO6 resolution sidekeep-concept, refine — becomes the resolution affordance of the exposure outcome (WS-8 “Draft content for X”)
audit_contentO6 W6.3 quality actionsconsolidate-into where-are-we-exposed (WS-8)
get_quality_actionsO6 W6.3consolidate-into where-are-we-exposed (WS-8)
get_quality_briefingO6 W6.3consolidate-into where-are-we-exposed (WS-8)
get_quality_summaryO6 W6.3consolidate-into where-are-we-exposed (WS-8)
get_review_queueO3 W3.4 doc-controlconsolidate-into what’s-in-my-queue (WS-8)
get_assignments_for_userO4 W4.1 / queueconsolidate-into what’s-in-my-queue (WS-8)
get_governance_queueO3 W3.4 governanceconsolidate-into what’s-in-my-queue (WS-8) — [LIAM] keep content-review vs governance as one queue concept with a facet, or two? (CLAUDE.md flags them as separate workflows)
Surface (concept)Outcome / workflowVerdict
list_active_procurementO2 W2.1keep-concept, refine under completing-forms (WS-9)
get_procurement_detailO2 W2.1keep-concept, refine under completing-forms (WS-9)
get_bid_questionO2 W2.1refine + rename → form-question (fresh DB drift, §1.4; WS-9)
get_content_effectivenessO3 W3.5 flywheel signalkeep-concept
cite_contentO1/O2 citation writekeep-concept, refine (arg already form_response_id post-rename; align description)
list_templatesO2 W2.4 template coverageconsolidate-into completing-forms coverage (WS-9)
get_template_coverageO2 W2.4consolidate-into completing-forms coverage (WS-9)
get_template_gapsO2 W2.4consolidate-into completing-forms coverage (WS-9)

2.4 MCP tools — Content management (curation; densest group)

Section titled “2.4 MCP tools — Content management (curation; densest group)”
Surface (concept)Outcome / workflowVerdict
get_workspace_itemssubstratekeep-concept
get_content_item / get_content_itemsO1/O7 fetchconsolidate single+batch into one get (one-or-many param)
get_document_diff / get_document_versionsO3 W3.4 doc-controlkeep-concept (lineage substrate; A20)
assign_content_owner / bulk_assign_ownerO3 W3.4 ownershipconsolidate single+batch into one assign
create_content_itemingest-adjacentrefine — track external-folder-canonical reality (ID-101/ID-45), not the retired inline classify+embed model (SYNTHESIS.md §1.4) [LIAM]
update_content_itemcuration writekeep-concept
delete_content_itemcuration writekeep-concept
update_publication_statusO3 publication gatekeep-concept (human-gated per WS-5)
classify_contentinline AI touchpointkeep-concept, refine — apply §5 grounding standard (strict tools)
generate_summaryinline AI touchpointkeep-concept, refineoutput_config.format (§5)
get_entity_relationshipsO1 W1.3 groundingkeep-concept, refine — promote to ontology-grounding affordance (A21); mis-grouped today (belongs with answering)
find_duplicate_candidates / find_all_duplicatesdedupconsolidate dup-detection into one (scope param); de-overlap with import dedup
create_review_assignmentO3 W3.4consolidate-into what’s-in-my-queue write side
review_governance_item / update_governance_statusO3 W3.4 governancekeep-concept (human-gated)
get_change_reportO4 W4.3 / change feedkeep-concept, refine (feeds reorientation + briefing)
supersede_content_itemO3 W3.2 supersessionkeep-concept (A20; human-confirmed)

2.5 MCP tools — Intelligence, guides, apps trigger tools

Section titled “2.5 MCP tools — Intelligence, guides, apps trigger tools”
Surface (concept)Outcome / workflowVerdict
get_intelligence_summaryO4 W4.2 / O5keep-concept, refine — becomes the consumption read for the per-persona “so-what” layer
trigger_intelligence_pollO5 W5.1 adminkeep-concept
list_guides / get_guideO8 W8.1 readkeep-concept
create_guide / update_guideO8 W8.1 writekeep-concept
show_coverage_matrix (app tool)O6 densitykeep-concept (visual-density layer; WS-8 resolution surface)
show_procurement_dashboard (app tool)O2 densitykeep-concept, refine → forms framing (WS-9)
show_reorient_me (app tool)O4 W4.3 densitykeep-concept, refine (widen beyond KH; WS-4)
show_intelligence_feed (app tool)O5 densitykeep-concept
ResourceOutcomeVerdict
kb://items/{id}O1 fetchkeep-concept
kb://bids/{id}O2refine + rename → forms (fresh DB drift; WS-9)
kb://qa/{id}O1 W1.2keep-concept
kb://coverageO6keep-concept (feeds exposure)
kb://dashboardO4keep-concept (feeds queue)
kb://taxonomysubstratekeep-concept (absorbs bl-52 taxonomy resource — D5)
kb://entitiesO1 W1.3 groundingkeep-concept, refine (A21)
kb://quality-briefingO6consolidate into exposure framing (WS-8)
ui://coverage-matrix/app.htmlO6keep-concept
ui://bid-dashboard/app.htmlO2refine + rename → forms
ui://reorient-me/app.htmlO4keep-concept, refine
ui://intelligence-feed/app.htmlO5keep-concept
PromptOutcomeVerdict
reorientO4 W4.3keep-concept, refine (widen beyond KH; WS-4)
bid_briefingO2/O4refine + rename → form/procurement briefing (WS-9)
coverage_analysisO6consolidate into exposure framing (WS-8)
draft_responseO2 W2.1keep-concept, refine under completing-forms (WS-9)
review_itemO3 W3.4keep-concept
sector_briefingO4 W4.2keep-concept, refine (per-persona so-what layer)
bid_pipeline_reviewO2refine + rename → form/procurement pipeline (WS-9)

2.8 Plugin commands (8) and plugin skills (9)

Section titled “2.8 Plugin commands (8) and plugin skills (9)”
SurfaceOutcomeVerdict
/kb:searchO1/O7keep-concept, refine → thin orchestrator over the single find entry (WS-7)
/kb:briefingO4 W4.1keep-concept
/kb:sector-briefingO4 W4.2keep-concept
/kb:coverageO6keep-concept, refine (exposure framing)
/kb:change-reportO4keep-concept
/kb:bid-pipeline-reviewO2refine + rename → forms (WS-9)
/kb:bid-statusO2refine + rename → forms (WS-9)
/kb:draft-responseO2 W2.1keep-concept, refine under completing-forms
plugin skill search-strategyO1keep-concept (Skill = expertise layer)
plugin skill knowledge-synthesisO1keep-concept
plugin skill bid-writingO2refine → completing-forms (WS-9; the named generalisation)
plugin skill classificationO3 inlinekeep-concept, refine (eval-born; §5)
plugin skill content-creationO3keep-concept
plugin skill content-governanceO3 W3.4keep-concept
plugin skill governance-reviewO3 W3.4keep-concept[LIAM] merge with content-governance, or keep distinct (mirrors the two-queue question, §2.2)?
plugin skill daily-briefingO4 W4.1keep-concept
plugin skill guide-builderO8 W8.1keep-concept
TouchpointOutcomeVerdict
classify.ts (2-pass forced tool use)O3 ingestkeep-concept, refine — strict tools + additionalProperties:false (§5; bl-50)
draft.ts 3-pass (analysis / cited drafting / quality)O2 W2.1keep-concept — 3-pass split stays forced (§5; citations×SO incompatible)
quality-check.ts (Pass 3)O2keep-concept (already output_config.format)
extract-questions.tsO2 W2.xkeep-concept, refine — model-id fixed {71.5}; add strict schema (§5)
summarisation / intelligence scoring / guide generation / cronsO4/O5/O8keep-concept, refine — each declares an eval requirement (§5; ID-104)

2.10 Outcome-level gaps (concept = gap; no current surface)

Section titled “2.10 Outcome-level gaps (concept = gap; no current surface)”
GapOutcome / workflowSource
Ontology-grounded answeringO1 W1.3 (A21)...strawman.md §C, §D
Source connection + discovery + monitoringO9 W9.1–W9.6 (A16, A22)SYNTHESIS.md §1.7; own Task (WS-6)
Fact-check on demandO3 W3.1 (A17)client-promised, unbuilt
Feature-ingest with supersessionO3 W3.2 (A20)client-promised, unbuilt
Case-study pushO3 W3.3client-promised, unbuilt
Sales trigger → outreach draftO5 W5.2 (A23)unserved; first consumption build (WS-10)
Marketing content pipelineO5 W5.3unserved (U13)
Roadmap / competitor signalO5 W5.4, W5.5underserved
Proposal assembly + exportO2 W2.2 (A18)design-only; Phase 3 (WS-10)
Renewal pack assemblyO2 W2.3unspecced; own Task candidate
Scheduled/triggered push deliveryO4/O5/O9 (A23)every briefing is pull-only today

Coverage check: all current-surface rows (tools incl. app-trigger tools, resources, prompts, commands, plugin skills, inline touchpoints) carry a verdict above; gaps are enumerated separately. Exit contract §2 met.


Honouring WS-11 (no hard numeric ceiling as design law; structured tool-justification

  • progressive utility; surface shape follows the outcome architecture). The layer rules (SYNTHESIS.md §1.3) place: tools = ability, skills = expertise, apps = visual density, commands = thin orchestrators; duplicate across layers only where the layer adds genuine value (the Skill Question).

Option A — Single outcome-grouped server (one surface, parameterised entries)

Section titled “Option A — Single outcome-grouped server (one surface, parameterised entries)”

One MCP server. Curation collapses into a small set of outcome-shaped entries: one find (WS-7), one where-are-we-exposed + one whats-in-my-queue (WS-8), one completing-forms surface (WS-9), one get (one-or-many), one dedup, one assign. The 4 MCP Apps stay as the visual-density layer for exposure/queue/forms/intelligence; prompts

  • plugin commands become thin orchestrators over the entries; plugin skills carry expertise. Headless agents (the many) read/propose via the same entries.
  • Pros: lowest cognitive overhead for Claude and users; matches “you don’t need to remember the tools — just ask”; one client-contract break; simplest eval surface.
  • Cons: consumption and curation share one auth/role surface; a writer-heavy admin set sits next to read-only consumption entries (relies on checkMcpRole/RLS to separate, not server boundary).

Option B — Split by audience (consumption server + curation/admin server)

Section titled “Option B — Split by audience (consumption server + curation/admin server)”

Two servers: a consumption server (find, briefings, intelligence, exposure read, forms read, grounded answering — the headless-complete WS-4 set) and a curation/admin server (create/update/delete, governance, assignment, supersession, classify/summarise writes). Apps + skills split to match.

  • Pros: clean audience separation; the headless-complete consumption set is a self-contained, easily-evaluable connector; admin risk surface is isolated; aligns with the ratified “split into multiple servers rather than defer-load” instinct (now a shape tool, not a ceiling mandate, per WS-11).
  • Cons: two connectors to provision/auth; some workflows (resolution loops in WS-8) cross the boundary; double the bundle/inventory/eval lockstep.
Section titled “Option C — Outcome-grouped core + headless-agent fleet (recommended)”

Option A’s single outcome-grouped human-facing server, plus a deliberately small set of human entry points backed by a fleet of headless agents doing the work in the background (WS-11’s “50 headless agents + 1 insight/action tool beats 58 user-facing tools”; lane-a-workshop-feedback.md Q10). The human surface is the insight/action layer (find + the two attention outcomes + completing-forms + briefings, surfaced through apps with resolution affordances); the agents (sweeps, watch-triggers, pre-assembly, discovery/monitoring) read + propose under WS-5 discipline and graduate to per-workflow auto-apply as ID-104 quality metrics earn it (WS-5).

  • Pros: directly implements the ratified outcome architecture and the progressive-utility thesis; minimises the surface a human/Claude must “get to grips with”; the agent fleet is where O5/O9/O3 gaps get served without inflating the human tool count; eval-and-trust loop (ID-104 + Raindrop-style metric surfacing, WS-5/WS-13) is the graduation mechanism.
  • Cons: depends on a headless-agent runtime KH does not yet have (the WS-4 /idea-refine session must land the requirement first, §7); more eval contracts to author up front (per-agent), which ID-104 must absorb.

Recommendation: Option C, layered on Option A’s single human-facing server for v1 (defer Option B’s audience-split to the point where the admin/consumption auth surfaces diverge enough to warrant a second connector). Rationale: it is the only option that honours WS-11’s “shape follows outcome architecture” literally, serves the O5/O9/O3 gaps through the agent fleet rather than new human tools, and uses ID-104 as the progressive-utility motor (WS-5). The single human server keeps the one-coordinated-break migration property of Option A; the agent fleet is additive and can grow per-Task.

Layer placement (all three options): prompts and plugin commands stay thin orchestrators over the consolidated entries; the 4 MCP Apps remain the visual-density layer and the WS-8 resolution surface (“Draft content for X” / “Discuss options for Y”); plugin skills remain the expertise layer (completing-forms, knowledge-synthesis, guide-builder, etc.); bl-26 outputSchema is the forward standard for new tools, not retrofitted onto retiring ones (§5).


ID-71 is the structuring task — it defines the approach, standards, and target surface (refine + define). Net-new product capabilities graduate to their own Tasks with their own spec chains, scheduled on a prioritisation basis in logical groupings (Liam’s Q1 answer, lane-a-workshop-feedback.md). ID-71 reserves the surface concepts (e.g. A16 source-connection, A22 propose-confirm) but is not the delivery vehicle.

Workflow familyGraduates to own Task?Notes / priority grouping
O9 onboarding / content-gathering (W9.1–W9.6)Yes — own Task NOW, local-firstRATIFIED WS-6. v1 gated: local file server, KH-defined directory structure, ETL additions; first client = testbed; NO SharePoint/Notion connector until pipeline confidence earned. ID-71 reserves A16/A22 concepts. Group 1 (differentiation) — day-one is where “this is different” forms.
O5 W5.2 sales trigger → outreachYes — own Task (recommended; not inside sales-proposal-workspaces)First consumption build (WS-10). Account data lives in HubSpot, already connected to the client’s Claude Cowork via MCP — account-matching rides that connector, not KB entity tables. See §4.1 — the existing sales-proposal spec is NOT its home. Group 1 (differentiation).
O5 W5.3 marketing content pipelineYes — own Task, sequenced after W5.2 (WS-10)U13, named-never-specced. Group 2.
O2 W2.3 renewal packsYes — own TaskEntirely unspecced; net-new product capability (assembly + per-client sector context). Emma persona. Group 2.
O3 W3.1–W3.3 trust trio (fact-check, feature-ingest, case-study push)Yes — own Task(s)All client-promised, all unbuilt. Candidates to group as one “trust maintenance” Task or split fact-check (A17) from supersession (A20) from case-study push. Group 2/3 — O3 partly inherited from the pipeline (WS-1), so urgency is lower than O5/O9.
O2 W2.2 proposal assembly + exportPhase-3 build (existing workstream)Sales-proposal-workspaces Phase-3 owns the composer/export (see §4.1); not an ID-71 deliverable. Group 3.
Search/exposure/queue/forms consolidation (WS-7/8/9)Stays in ID-71 as implementation wavesSurface rework, not new product capability — the core of {71.2}/{71.3}.
Inline-touchpoint standards rollout (§5)Stays in ID-71Standards work; per-touchpoint eval requirement.
Bid→forms rename lockstep (§1.4)Stays in ID-71One coordinated client-contract break (code + fixtures + bundle + inventory + evals + client guide), ideally aligned with the prod DDL cutover wave.

4.1 WS-10 assessment — W5.2’s home and the sales-proposal drafts’ suitability

Section titled “4.1 WS-10 assessment — W5.2’s home and the sales-proposal drafts’ suitability”

The existing sales-proposal-workspaces spec (.../specs/sales-proposal-workspaces/PRODUCT.md + TECH.md, ratified S243) is a Phase-1 reserved-DB-seat-only spec: it lands the sales_proposal_workspaces satellite (PK + FK + RLS + grants, inheriting RWS S-1..S-8) and explicitly defers ALL behaviour — composer, workflow state machine, win-loss, export, proposal-specific extraction — to a Phase-3 build (S-6, S-7). Its own §S-8 gap flag records that the four substrate drafts (themes/workspaces/sales-proposals/sales-proposals-{reuse-audit,workspace-plan}-arm-{a,b}.md) are unratified, unreviewed by Liam, and pre-date the canonical-pipeline pivot — exactly the suitability concern WS-10 raised.

Finding: W5.2 (sales-trigger detection → outreach draft) does not belong inside the sales-proposal workstream. Two distinct reasons:

  1. Different outcome. W5.2 is an O5 intelligence-consumption workflow (detect a trigger-shaped intelligence item → match to accounts → draft an outreach hook). The sales-proposal spec is an O2 revenue-document-assembly capability (compose a full proposal). They share the David persona but not the data flow or the surface.
  2. Different data home. W5.2’s account-matching rides the HubSpot↔Cowork MCP connector (WS-10), not the sales_proposal_workspaces satellite or KB entity tables.

Recommendation: open W5.2 as its own Task (per WS-2), sequenced first in the O5 consumption group. Treat the sales-proposal-workspaces Phase-3 build as a separate future Task whose substrate (the four arm-a/arm-b drafts) still needs Liam’s ratification + canonical-pipeline reconciliation before adoption — carry proposal-writer (the in-repo skill, WS-12) as a named input/exemplar for that W2.2 thread, not for W5.2. proposal-writer was flagged by Liam at sales-proposal kickoff and missed in earlier findings; it is recorded here as the exemplar for the proposal-assembly thread.


Grounded in Lane B currency (...lane-b-structured-outputs-citations-currency.md, VERIFIED 2026-06-09, SDK 0.96.0).

5.1 Grounding standard — pick exactly one of three shapes per touchpoint

Section titled “5.1 Grounding standard — pick exactly one of three shapes per touchpoint”
  1. Structured data, no source attributionoutput_config.format with json_schema (GA since 2026-01-29). Default for analysis, quality checks, metadata extraction, summaries-as-data. Validate parsed output with the existing zod schema (z.infer remains canonical).
  2. Model-decides / forced-tool extraction → tool use with strict: true + recursive additionalProperties: false on every object. Pre-flight the new strict limits (≤20 strict tools, ≤24 optional params, ≤16 union-typed params per request); classify’s type: ['string','null'] unions must be counted (the only genuine risk in bl-50).
  3. Output must be traceable to KB sources → citations over search_result content blocks (GA, no beta header). These calls must not include output_config.format / output_format (architectural 400 — quoted verbatim in Lane B §4).

Needs both structure and citations → split into two calls. The 3-pass draft.ts pipeline is the canonical instance and stays forced (do not merge). Record as a standing constraint with a changelog watch; revisit only if Anthropic ships cited structured outputs.

Mandatory hardening at every structured touchpoint: handle stop_reason: "refusal" and "max_tokens" explicitly; no silent fallback defaults (replace draft.ts Pass 1’s bare try/catch — log + surface); never use assistant prefills (400 on all 4.6+ models); keep schemas static per call site (grammar-cache discipline — name/description changes are free, schema-structure changes recompile). Schema enforcement ≠ semantic correctness — evals remain the quality gate (the durable S195/OPS-30 lesson).

bl-26 outputSchema is the forward standard for new MCP tooling, NOT retrofitted onto retiring/changing tools (Liam’s note, supporting-ai-tooling-notes.md; D5). New outcome-shaped entries (§3) declare an outputSchema where it fits the use case.

5.3 Per-touchpoint eval requirement (contracts live in ID-104)

Section titled “5.3 Per-touchpoint eval requirement (contracts live in ID-104)”

Every tool / prompt / plugin skill / inline touchpoint / headless agent declares an eval requirement. The contracts — eval-runner, severity/variance thresholds, baseline lifecycle (promoteBaseline, history), Claude-as-judge tool-description rubric — live in Task ID-104 (NOT ID-102; older docs saying ID-102 are stale). Today only tools (behaviourally, L1/L3/L4) and 4 inline touchpoints (baselines) are evaluable; the 7 prompts, 12 resources, plugin skills and MCP Apps are eval-blind (SYNTHESIS.md §1.5). {71.2} PRODUCT must require each new/refined touchpoint to ship born-evaluable against an ID-104 contract.

5.4 Forcing-function hooks (tooling change ⇒ skill + eval/fixture update)

Section titled “5.4 Forcing-function hooks (tooling change ⇒ skill + eval/fixture update)”

TECH must specify hooks so that any tooling change forces create-skill/update-skill invocation and eval/fixture updates — mirroring the dev-workflow skill-update hooks (Liam’s note). The enforcement-test pattern (the recordAiCall() grep-guard generalises) plus the existing guard tests (mcp-fixture-sync.test.ts, pipeline-parity.test.ts, inventory-parser test) extended to prompts/skills are the mechanism. The WS-12 skills are the working material: create-skill/agent-development/mcp-builder carry built-in eval mechanisms (test-prompt loops, variance benchmarking, the 10-question MCP-eval process via mcp-builder/reference/evaluation.md), and the eval-process trio (llm-evaluation, prompt-engineering-patterns, context-engineering-collection) adds the methodology layer. WS-12 doctrine: use what exists, adjust only where we absolutely have to, adopt patterns where they benefit users — TECH wires these in rather than building bespoke eval scaffolding.


6. Third-party adoption plan (WS-13) + memory position (WS-14)

Section titled “6. Third-party adoption plan (WS-13) + memory position (WS-14)”

Verdicts ratified at the workshop (lane-a-workshop-outcomes.md WS-13/WS-14); sources in ...lane-b-third-party-sweep.md (all checked 2026-06-09).

ToolVerdictScope / trigger
MCPJam (MCPJam/inspector, Apache-2.0 core)AdoptLocal npx inspector for the interactive gap (JSON-RPC traces, OAuth debugging, MCP-Apps rendering outside Claude Desktop). Avoid the commercial evals module — KH’s L1/L3/L4 stays the CI source of truth. Schedule after {71.1} lands.
Raindrop Workshop (raindrop-ai/workshop, MIT)Adopt (dev-workflow)Local per-run trace inspection + agent-authored evals. Keep usage local; avoid coupling eval artefacts to Raindrop-proprietary formats. Candidate for the WS-5 metric-surfacing that earns auto-apply. Hosted Raindrop for in-platform agents = deferred.
watchmen (firstbatchxyz/watchmen, MIT)Pilot (one week)Cheap trial against the archived session corpus; could feed the evaluator-efficiency-sweep’s redundant-dispatch findings with concrete skill candidates. Revisit after consolidation lands.
Claude for Small Business pluginPattern sourceMap its 15 workflows against KH domains (below). Mirror its structure: small named workflow set, owner-initiated approval gates, skills as the unit of reuse.
nebula-graph posts, osiris, rowboat, html-anything, iiiReject (stand)Vendor pitch / wrong-shape / re-architecture; KG stage stays in Postgres via ID-101.
knowhere, mirageWatchRe-check at a concrete hard-PDF failure class (knowhere) or when KH scopes an in-platform agent runtime (mirage).

Claude SMB plugin — 15-workflow → KH-domain mapping (pattern source, not dependency): the SMB plugin connects to tools (QuickBooks/HubSpot/Slack/…); KH’s differentiation is governed company knowledge feeding such workflows, so KH’s remote MCP server is the knowledge-side counterpart. Workflows whose KH-data-grounded analogues are direct ID-71 inputs: morning business brief → O4 W4.1/W4.2 briefings; sales campaign execution → O5 W5.2/W5.3 (HubSpot connector already live for the client — WS-10); month-end / finance → Priya/O6 W6.2 numeric-traceability framing. The owner-initiated-approval-gate pattern is the concrete model for WS-5 propose-only-now / earned-auto-apply-later. (The remaining workflows — payroll, bookkeeping close, etc. — are connector-to-tool patterns with no governed-knowledge grounding and are out of KH scope; recorded as market signal.)

Memory (WS-14): MemPalace stays for dev-workflow — local-first, already wired into session lifecycle, no data-exposure risk. supermemory (MIT shell, closed engine) and memanto (MIT client, proprietary Moorcheh cloud core) offer no adoption-grade benefit over MemPalace and both route data through third-party cloud — disqualifying for SMB client data. Mine patterns only for the post-launch platform user-memory direct-on-Supabase build: memanto’s typed memory categories + information-theoretic retrieval (arXiv:2604.22085), supermemory’s temporal contradiction handling (KH already has lib/entities/temporal-reconciliation.ts pointing the same way). The MemPalace direct pattern stays ratified; wrapped pattern + user-memory tools stay post-launch-deferred.


Ratified 2026-06-14 (Liam). Dispositions: (1) RESOLVED — see headless-requirement-refinement.md (WS-4 /idea-refine outcome; invariants HC-1…HC-6). (2 two-queue) RESOLVED — one queue concept with a facet, not two outcomes. (2 create_content_item) carried to PRODUCT (depends on canonical-pipeline cutover state). (3 renewal packs) PRODUCT to confirm horizon + grouping. (4 trust-trio) PRODUCT to confirm grouping + sequence. (5) RESOLVED — Option C on Option A confirmed; Option B audience-split deferred. {71.1} flips done on this ratification.

  1. Queued /idea-refine headless-requirement session (WS-4)RESOLVED, see headless-requirement-refinement.md. The headless-complete set is baselined (O1/O4/O6 reads + W5.6 + W9.3) but must be refined in a dedicated interactive /idea-refine session with Liam before PRODUCT fixes the surface-#4 guarantee — including the “headless = managed agents” and incoming/outgoing connectivity clarifications. This session also lands the Option C headless-agent-fleet requirement (§3). Blocks the headless-completeness invariants in {71.2}.
  2. [LIAM] verdict rows carried from §2:
    • Two-queue vs one-queue — content-review queue vs governance queue (§2.2 get_governance_queue; §2.8 governance-review skill). CLAUDE.md treats them as separate workflows; WS-8 collapses to “what’s in my queue”. One queue concept with a facet, or two distinct outcomes?
    • create_content_item ingest shape (§2.4) — must track the external-folder- canonical reality (ID-101/ID-45), not the retired inline classify+embed model. Final shape depends on the canonical-pipeline cutover state at {71.2} time.
  3. Renewal packs (W2.3) horizon — confirmed own-Task candidate (§4); PRODUCT to confirm whether v1-adjacent or later, and whether it groups with the O2 proposal- assembly Phase-3 build.
  4. O3 trust-trio grouping (§4) — one “trust maintenance” Task or split fact-check/supersession/case-study-push? Priority is lower than O5/O9 because O3 is partly inherited from the pipeline (WS-1) — PRODUCT to confirm the grouping and sequence.
  5. Single-server (Option C on Option A) vs eventual audience-split (Option B) — §3 recommends Option C now; PRODUCT to confirm the deferral of the consumption/admin server split and the trigger condition for revisiting it.

End of {71.1} RESEARCH. Verdict table (§2) covers 100% of the current surface; target-surface options (§3) and Task-split map (§4) presented for ratification before {71.2} PRODUCT.