corpus¶
Semantic search and structured query over your own content.
corpus is a knowledge base for your own content — markdown, documents, or structured records. It classifies each item, embeds it through an OpenAI-compatible endpoint, and stores the vectors alongside structured metadata in Postgres/pgvector. Email lands first (IMAP and Gmail); anything that yields records fits the same pipeline.
A raw markdown layer keeps each item verbatim, and corpus derives a configurable set of access-controlled storage tiers from it — you set each tier's projection, the tool that exposes it, and who may read it, human or agent.
corpus answers three kinds of question, over a REST API and an MCP server:
- Find by meaning — semantic search: vector similarity with optional metadata filters.
- Filter by fact — structured query: exhaustive metadata queries that return every match (a given tag, an account, a time window) as plain SQL.
- Ask what matters — the enrichment priority signal: what needs an action, what's due or time-sensitive, what you're waiting on, what's happening in a domain like banking, health, or work.
The full surface (search, structured query, whole-document fetch, stats) is open to callers you trust with the raw items; the priority signal, free of raw content, is what a lower-trust caller — a cloud-model agent — reads on the sanitized tier.
Where to go next¶
- Architecture — the rendered module diagram of the pipeline.
- Configuration — every
CORPUS_setting. - Sources — the fetcher protocol, plus IMAP and Gmail.
- Database — the pgvector store and sync state.
- Enrichment — structured per-document metadata and the secret audit.
- Enrichment evaluation — compare enrichment models on a labelled synthetic set.
- Observability — OpenTelemetry traces and metrics.
- Development — tests, the pre-PR gate, and builds.