Skip to content

01 — Vision

Last verified: 28/07/2026 (S504 R7 re-verify — §5 workspace-era language rewritten to activity terms per DR-038, stale docs/… path refs re-pointed to their docs-site homes, heritage-table paths refreshed. Prior: 10/06/2026 ID-71 lane-a workshop outcomes — v1 canonical-store correction, headless-agent widening, AI-tooling direction §2.5, tool-ceiling supersession; per specs/id-71-ai-tooling/lane-a-workshop-outcomes.md. Scope: Mission, vision, user model, and the application-types-as-applications framing for Canonical v1. Status: [CURRENT-CANONICAL] — first of the 9-way split superseding the Phase-0.9 0.9-intended-architecture.md (fully superseded; not mirrored on the docs-site — history in the knowledge-hub-archive repo). Pattern ratified for Wave 1+ sub-doc dispatch. Layer: 1 (no upstream sub-doc deps) per specs/core-docs-pathway-assessment/INV-architecture-split-readiness.md §4. Companion sub-docs: 02-data-flow.md, 03-tech-stack.md, 04-workspace-types.md (describes the retired workspace tier — historical per DR-038), 05-qa-flow.md, 06-mcp-tooling.md, 07-collapse-list.md, 08-new-features.md, 09-diagrams.md.


Canonical (formerly Knowledge Hub) helps UK SMBs organise their data so AI agents — Claude in particular — can do meaningful work on their behalf. The blocker to SMB AI adoption is not the AI; it is the data. Email threads, shared drives, half-finished SharePoint sites, and the knowledge that lives in two senior heads cannot be handed to an AI assistant and produce reliable answers. Canonical exists to fix that side of the problem: build a clean, structured, well-governed knowledge base once, and everything downstream gets easier — bid responses, onboarding, compliance evidence, sales qualification, training.

The Canonical is the product. Procurement is the first application built on top of it. Sector Intelligence is the second. Sales Proposals is the next. More applications will follow, each consuming the same canonical corpus through different workspace and workflow surfaces.

1.1 The ontology pipeline as the first step

Section titled “1.1 The ontology pipeline as the first step”

A KB that AI can use is not just a pile of documents with embeddings on top. It is a corpus with a stable, evolvable shape — controlled vocabularies for content type, layer, lifecycle, scope, provenance, and the rest of the Layer-1 register; consistent application-type and form-type framings; and machine-readable structure so the same record can be served as a brief summary, a detailed explanation, or a reference-level appendix without the AI having to re-derive structure on every read. The ontology pipeline (WP6 markdown-as-source-of-truth scaffold, the parity-tested registry, and the cocoindex-fed classification flow) is one of the first concrete steps Canonical takes toward making SMB data AI-ready. Subsequent steps — application_types-as-applications (§4), per-application satellites (§4.2), corpus-level q_a_pairs with scope_tag-driven relevance, and the wider Phase 0.9 ratifications — extend that foundation.

The Wikipedia Principle — “one record, many views” — is inseparable from “the KB is the product”. Each fact has one canonical home in the corpus; that record is presented at multiple depths (summary / detail / reference) and through multiple surfaces (Claude in Word, Claude Desktop, the Canonical web UI, headless agents), but there is never a duplicate to keep in sync. Data quality is paramount because Claude consumes the corpus directly through MCP — there is no human translation layer between corpus and AI consumer, and any inconsistency surfaces immediately as a bad answer in the user’s workflow.

This is the inversion of the traditional knowledge-base shape:

  • A traditional KB is a destination — users come to the KB UI to find content.
  • Canonical is infrastructure — content flows to wherever Claude is operating, and the Canonical UI exists to govern, curate, and audit the corpus.

The platform reframes “AI features” as a category. AI processing — classification, embedding, entity extraction, summarisation, freshness scoring, deduplication — is invisible platform infrastructure. The product surfaces AI-derived outputs (Quality scores, summaries, change reports, classifications) as native platform capabilities, not as labelled AI features. Claude itself appears only as a destination (“Open in Claude”, “Continue in Claude”) — never as a badge on platform output. The AI Visibility Policy (reference/ai-visibility-policy.md) and the AI Integration Strategy (reference/ai-integration-strategy.md §1 “Vision and Strategic Positioning” — [PARTIALLY-SUPERSEDED] per §6.1) give the prior-art for this stance.


The platform is a curated corpus plus a set of governance, visualisation, and onboarding surfaces, fronted by an MCP server that Claude (in any of its clients) can use as a knowledge source.

Three convergent reframes from Phase 0.9 set the shape:

  1. Evidence-binding model (restated S441 — DR-025). The client’s sources are evidence streams connected through per-binding retention classes (keep-and-watch / ingest-once / live-connected / external-referenced); the canonical layer is the client’s database (promoted records) plus the OKF concept bundle — authority is earned at the knowledge-admission gate (promotion + ontology linter), never inherited from a folder. In v1 the primary binding is a controlled local file server: Canonical defines the directory / organisation structure, new content is added ETL-style, and changes to underlying content are picked up automatically (ID-71 WS-1, WS-6 per specs/id-71-ai-tooling/lane-a-workshop-outcomes.md). For genuinely living documents the file stays authoritative at source — edits happen in the source application (Word via the Claude extension, markdown editor, etc.); Canonical detects the change via the binding and re-ingests, with already-promoted records changing only via proposals (DR-026). SharePoint / Notion-style connectors remain the growth path of the same connect-don’t-upload vision, gated on pipeline confidence built with the first client as testbed (ID-71 WS-6). (Historical note: this reframe originally read “the client’s source files are the canonical content store” — true in the pre-concept-layer era, when files were the only place authority could live; retired by DR-025.)
  2. AI-consumer-first. The primary user is Claude with Canonical connected: a remote MCP server (the single reusable core), plugins (Cowork + Claude Code), and per-domain skill files. Headless includes managed agents, not only MCP-completable tooling — MCP + Skills from Claude Desktop / claude.ai / Cowork is itself a headless-agent form (MCP = connector, Skill = expertise); connectivity, incoming and outgoing, is a first-class headless requirement (ID-71 WS-4).
  3. Application types as applications. Six baseline application types (procurement, intelligence, sales_proposal, product_guide, competitor_research, training_onboarding), each realised as ACTIVITIES carrying their own ids (a form, a proposal, a guide, …) bound to an application_types row — one architectural pattern, not bespoke verticals. The earlier “each a workspace + satellite table” framing is retired (DR-038, S452): no new *_workspaces tables ever; structured data scopes to the activity.

Five principles run through every architecture decision. These are an architecture-adapted restatement of the Platform Overview’s customer-facing version — the customer doc is the canonical phrasing; this list is the architecture-doc lens for the same principles.

  1. One record, many views. Every piece of knowledge lives in exactly one entry. That entry carries a brief summary, a detailed explanation, a reference-level appendix, version history, classification, provenance, and lifecycle state — all attached to a single canonical record. Editing one never leaves another stale. Search returns one result, not three. (Architectural consequences: corpus-level q_a_pairs; no per-workspace content duplication; scope_tag overlap drives workspace relevance.)
  2. Helping users organise, not extracting their value. The platform’s job is to help SMBs structure their own knowledge so AI can use it — not to learn from their data, train models on it, or accumulate centralised insight. Each tenant’s corpus stays the tenant’s corpus; provenance is recorded; the platform does not aggregate across tenants. (Architectural consequences: per-tenant RLS; provenance enum on every Layer-1 vocabulary table; no cross-tenant analytics surfaces.)
  3. Observe and intervene, not prevent and approve. Content can be updated freely. Every change is versioned. Significant changes are surfaced for review through change reports and the review queue. This keeps the library moving without forcing every edit through a gatekeeper. (Architectural consequences: optimistic write paths; change reports as a post-hoc surface; review queue as a workflow, not a write gate.)
  4. Programmatic where it can be, AI where it must be. Deterministic work (deduplication, freshness states, filtering, link resolution, governance scheduling) uses ordinary code. AI is reserved for tasks only it can do well — classification, summarisation, entity extraction, drafting. (Architectural consequences: cocoindex flow stages for AI work; deterministic functions for deterministic tasks.)
  5. The library and the applications feed each other. Applications (procurement, intelligence, sales proposals) pull from the KB to draft / surface / respond. The artefacts they produce — won bid responses, vetted summaries, validated proposals — get curated back into the library. The corpus that wins bid #20 is meaningfully better than the one that wrote bid #1. (Architectural consequences: bidirectional flow between application workspaces and the corpus; bid-feedback ingestion pipeline; corpus-level q_a_pair shape that absorbs application-derived content cleanly.)

The Wikipedia / one-golden-record principle gives the platform integrity (single canonical home, single attribution, single freshness state). The AI-consumer-first lens makes the corpus useful where work is actually happening (in Claude clients, not in a separate destination UI). The application-types-as-applications framing keeps the platform extensible without re-architecting per vertical. These three together — corpus integrity, surface flexibility, structural extensibility — are the v1 shape.

Three ratified directions govern how the AI-consumption layer (MCP tools / resources / prompts, plugins, skills, apps, inline AI touchpoints) evolves, per specs/id-71-ai-tooling/lane-a-workshop-outcomes.md (10/06/2026):

  1. Outcome-first tooling. Verdicts attach to outcomes, not tool names (concept-over-artefact). The surface is designed by working backwards from the most valuable client workflows (ID-71 WS-11; ranked outcomes per WS-1).
  2. Eval-everything. Every AI touchpoint is born-evaluable; eval infrastructure is Task ID-104, and ongoing feedback loops are the mechanism by which tooling earns expanded autonomy (ID-71 WS-5).
  3. Progressive utility / trust graduation. Agents propose; humans gate publication and outcome writes at launch. Auto-apply is earned per-workflow when its quality metrics justify it, with an audit trail required (ID-71 WS-5, WS-11).

The ~30–40 tool ceiling carried by reference/ai-integration-strategy.md is superseded: no hard numeric ceiling operates as design law. It is replaced by structured justification per tool — whether a tool is required and how best to implement it — with the surface shape following the outcome architecture (e.g. many headless agents behind one insight/action MCP tool can beat a wide user-facing tool set) (ID-71 WS-11; see §6.1). The MCP surface detail remains in 06-mcp-tooling.md.


Canonical’s primary user is a user of Claude who has Canonical connected as a knowledge source. The connection runs through three vehicles, each carrying a different surface contract:

  • Remote MCP server. Tools, resources, prompts, and MCP Apps exposed at the platform’s MCP endpoint. The single reusable core — every Claude client below consumes it.
  • Plugins. Canonical ships a Cowork plugin (commands + skills for Claude’s collaboration product) and a Claude Code plugin (developer-focused KB access). Plugins extend MCP-server tool access with surface-specific commands.
  • Skill files. Per-domain skills (procurement-writing, UK procurement context, etc.) that teach Claude how to use Canonical tools in context, distinct from the raw tool definitions.

Cross-ref: reference/ai-integration-strategy.md §3-§9 carries the four-layer detail (MCP Server / MCP Apps / Cowork Plugin / Claude Code Plugin) — [PARTIALLY-SUPERSEDED] per §6.1; §1 + the layer descriptions remain load-bearing.

Primary surfaces (Claude clients, ranked by likely usage):

  1. Claude-native products: Claude Desktop, claude.ai, Cowork. Chat-style interaction. Claude has access to the Canonical corpus via the remote MCP server plus the Cowork plugin where applicable. User asks: “What is our policy on data protection?”; Claude searches the corpus, returns a cited answer.
  2. Claude in Word / Excel / PowerPoint. Claude embedded in Office. The user edits a document; Claude has Canonical access via the MCP server; relevant pre-curated content drops into the document. Edits to source files happen here — not in Canonical UI in v1.
  3. Claude Code. For SMBs building their own automation. The Canonical MCP server is one of several MCP servers the user has connected; Canonical also ships a dedicated Claude Code plugin.
  4. Headless agents. Automated workflows (e.g. nightly procurement-response drafting). “Headless” includes managed agents, not only MCP-completable tooling; MCP + Skills consumed from Claude Desktop / claude.ai / Cowork is itself a headless-agent form (MCP = connector, Skill = system-prompt-like expertise). Core workflows remain completable by MCP alone, and ease of connectivity — incoming and outgoing — is a first-class headless requirement as the world becomes more agentic (ID-71 WS-4, widening the earlier MCP-only framing).

The Canonical web UI is the platform’s admin / governance / visualisation tier. It is intentionally scoped narrower than a typical “knowledge base UI”:

  • Onboarding. Workspace creation, content-folder connection, taxonomy review, role assignment.
  • Governance. Review queue, freshness dashboards, expiry tracking, content-owner management.
  • Visualisation. Coverage matrix, gap analysis, intelligence feed, change reports.
  • Admin. Workspace settings, user roles, taxonomy / tag / scope_tag administration (where v1.1 extends this to client-managed vocabularies — see backlog).
  • Power-user curation. Curators reviewing bulk data; rich-text editing of long-form Q&A entries; metadata sidebar for batch attribute edits.

Canonical supports several applications — coherent product surfaces, each with their own goals and workflows — but each application is an instance of the same architectural pattern. An application_types instance table holds the registered applications.

Six core-provenance application types ship in v1:

Application typeDomainFirst-client status
procurementBid / RFP / PQQ / ITT / framework / DPS / G-Cloud responsesFirst domain application
intelligenceSector intelligence feeds, news monitoring, market signal captureSecond domain application
sales_proposalSales proposal drafting and reuseNext-after-intelligence
product_guideProduct / capability guidesBaseline, v1 instance
competitor_researchCompetitor research and positioningBaseline, v1 instance
training_onboardingTraining and onboarding materialBaseline, v1 instance

The term “Workspace” was used previously, and was the wrong abstraction for scoping application-type data: each activity carries its own id (form id, proposal id, guide id, …) and structured data scopes to the ACTIVITY. The workspaces table still exists as legacy infrastructure (a back-pointer during migration only); no new *_workspaces tables are ever minted, and workspace-keyed procurement surfaces are legacy-shaped, being retired under ID-145. workspace was never a synonym for client/tenant.


Procurement is the first application Canonical serves. It was previously known as and scoped only to “bids” - this was then updated, with Procurement becoming the umbrella term for various form types, including “bids”.

The Procurement application enables a form to be completed, by a user, or AI. The application includes:

  • A library of past responses (Q&A pairs derived from past forms, sales proposals, and corpus content; corpus-level + scope_tag-driven per 00-synthesis-v2.md §3.6).
  • A response composer that assembles a draft from Q&A pairs + corpus content.
  • A coverage / gap surface for form templates, showing which question types are missing for a specific form template.
  • A workflow state machine (PROCUREMENT_WORKFLOW_STATES) tracking each response from intake → drafting → review → submission → outcome capture.

The procurement application supports several form_type values within one application shape — see the procurement subset of ontology/26-form-type.md baseline_values (source-of-truth). A client’s procurement application runs multiple form types over its lifetime, each form carrying its own id (DR-038); per-form-type behaviour is code-driven in v1 (data-driven extension deferred to v2 per Q-OQR1-14).

Sector intelligence is the second first-domain application. The intelligence application tracks RSS feeds, news monitoring, and market signal capture, with the corpus annotated by domain / subtopic / scope_tag classification to make signals retrievable by Claude during downstream work (e.g. when drafting a procurement response that needs current market positioning).

Sector intelligence shares the corpus with procurement — there is one knowledge base, not separate ones per application.

Sales proposals is the next-after-procurement application. The shape is parallel to procurement: a library of reusable proposal content, a composer that draws from the library and corpus, a coverage surface, and a state machine.

The Wikipedia Principle is the load-bearing piece here: a fact captured in a procurement response, a sales proposal, and a product guide is one record in the corpus, not three.

product_guide, competitor_research, and training_onboarding are baseline application types — they have schema seats in v1 but are not the headline applications.


This sub-doc is the first of nine that supersede the Phase-0.9 0.9-intended-architecture.md (2015 lines, S229; fully superseded — not mirrored on the docs-site, history in the knowledge-hub-archive repo). The source doc predates S233-S237 ratifications; its central frame is at variance with the ratified state. Full audit trail of the 10 superseded items + ratifying doc per row: specs/core-docs-pathway-assessment/INV-architecture-split-readiness.md §2. Canonical-state sources downstream sub-docs cite: 00-synthesis-v2.md §3 + §4 (Q-OQR1-01..17, WP8 N7/N9/audit_log RLS, S237 CV resolutions); 0.9-decision-graph.md §11 (consolidated ONT / COCO / combined-PR rows).

Per construction guide §4.1 — three-tier status taxonomy ([CURRENT-CANONICAL] / [PARTIALLY-SUPERSEDED] / [FULLY-SUPERSEDED]); table shape mirrored by every subsequent sub-doc against its own heritage set per §4.2.

DocDateStatusUseful for
docs/client-documentation/Knowledge Hub — Platform Overview.md23/04/2026[PARTIALLY-SUPERSEDED] — vision + Five Guiding Principles + AI-as-plumbing framing hold; bid_workspaces framing stale per Q-OQR1-02.”Data is the blocker” framing; AI-as-plumbing positioning; Five Guiding Principles cited in §1 + §2.2.
docs/client-documentation/Knowledge Hub — Claude Integration Guide.md(companion to Platform Overview)[PARTIALLY-SUPERSEDED] for Claude integration onboarding narrative.Claude-via-MCP user-journey framing; integration positioning to end users.
reference/ai-integration-strategy.md11/03/2026 (verified 28/04/2026)[PARTIALLY-SUPERSEDED] — §1 “Vision + Strategic Positioning” + four-layer architecture (MCP server / MCP Apps / Cowork plugin / Claude Code plugin) hold; §17 + §18 build status + effort estimates pre-Phase-0.9; tool-count ceiling (~30–40) superseded per ID-71 WS-11 (10/06/2026) — structured justification + progressive utility replace any hard numeric ceiling (§2.5).”Working within an LLM” framing (§1.2); SMB differentiator framing (§1.3); four-layer integration model cited in §3.1.
reports/product-differentiation-audit.md (re-filed from reference/ S504)19/03/2026[PARTIALLY-SUPERSEDED] — competitive positioning + “structured KB beats AI-on-top-of-pile” argument hold; “AI features” framing shifted post-S237 CV resolutions.Competitive context (Notion / Loopio / Glean); five-differentiators framing.

End of sub-doc. (Construction-era dispatch footer, historical: Wave 1 parallel 03-tech-stack.md + 07-collapse-list.md; Wave 2 main 04-workspace-types.md.)