Skip to content

Configuration

Every setting is an environment variable prefixed CORPUS_. The full set lives in src/corpus/config.py; the common ones:

Variable Purpose
CORPUS_DATABASE_URL Postgres DSN (pgvector)
CORPUS_DB_SCHEMA schema for the document table (default corpus)
CORPUS_OPENAI_API_BASE OpenAI-compatible embedding endpoint (…/v1)
CORPUS_OPENAI_API_KEY key for that endpoint
CORPUS_EMBEDDING_MODEL embedding model name (default local-embed)
CORPUS_EMBEDDING_DIMENSIONS vector dimension (default 1024)
CORPUS_ENRICH_MODEL model for batch enrichment (empty disables it)
CORPUS_AUDIT_MODEL model for the secret audit (default: the enrichment model); keep it local when enrichment runs remotely
CORPUS_ENRICH_CONCURRENCY enrichment requests in flight (default 8)
CORPUS_ENRICH_MAX_INPUT_CHARS characters sent per document for enrichment, 0 = unlimited (default)
CORPUS_MODEL_OPTIONS per-model request options as JSON (see Enrichment)
CORPUS_ENRICH_RETRIES attempts per enrichment request (default 10)
CORPUS_ENRICH_RETRY_MAX_WAIT cap on the backoff between them, seconds (default 60)

The embedding endpoint is any OpenAI-compatible API. Point it at a local model and nothing leaves your network; point it at a hosted provider and the pipeline is unchanged.

Riding out a busy enrichment endpoint

A model server saturated by a concurrent backfill sheds load as 5xx and drops connections, then recovers within minutes. Rather than abort the pass, the enricher retries such a failure in-process with exponential backoff — capped per wait at CORPUS_ENRICH_RETRY_MAX_WAIT and jittered, so concurrent workers do not return in one burst — over CORPUS_ENRICH_RETRIES attempts; the defaults wait out roughly four minutes of outage before giving up with EnrichUnavailableError. A 4xx is never retried: it is the request's own fault, and the record is skipped instead.

Raise CORPUS_ENRICH_RETRIES where the endpoint shares a GPU with other work and can be gone for longer; enrichment is resumable either way, so a run that does give up continues from enriched_ids on the next pass.

Riding a 5xx out hides it, so each retry is logged with the response's attribution detail — see diagnosing a shedding endpoint for reading those lines before changing CORPUS_ENRICH_CONCURRENCY.

Capping enrichment input

Enrichment sends the subject plus the whole body. On a KV-cache-bound backend — a single consumer GPU at concurrency 16–32 — long bodies inflate prefill and cache use until the server preempts, which caps throughput. Setting CORPUS_ENRICH_MAX_INPUT_CHARS bounds each prompt: the head is kept (the subject and opening lines, where the classification and summary signal concentrates) and the dropped tail is replaced by a [truncated] marker. The subject is not privileged, only first, so a limit shorter than the subject cuts the subject itself — keep the cap comfortably above your longest subject. It trades fidelity on long bodies for throughput, so it is off by default; a backend with cache headroom should leave it at 0.

The cap applies only to enrichment. The secret audit always sees the full text — a credential can sit past the cap, and its candidates come from a full-body scan. A text longer than the audit model's context window therefore fails its audit; the enrichment is kept and the audit is counted as audit_failed.

Per-source variables are namespaced by fetcher name — see IMAP and Gmail.

scan-gate

corpus scan-gate applies the egress policy (corpus.egress.policy) to LLM request bodies before they leave the network. The policy is gateway-neutral; a protocol adapter connects it to a gateway. The only adapter today is envoy-ext-proc, Envoy's external-processing gRPC protocol, which any Envoy-based proxy can call.

Variable Purpose
CORPUS_SCAN_GATE_ADAPTER gateway protocol adapter (default envoy-ext-proc)
CORPUS_SCAN_GATE_PORT listen port (default 9002)
CORPUS_SCAN_GATE_WORKERS concurrent request streams (default 8)
CORPUS_SCAN_GATE_FAIL_OPEN pass a body that is not JSON (e.g. multipart audio) through unchanged; the default refuses it
CORPUS_SCAN_GATE_BLOCK_TYPES comma-separated secret types refused with a 403 instead of redacted (default private_key)
CORPUS_SCAN_GATE_SKIP_MODELS comma-separated model names passed unscanned because they are served locally (default none)
CORPUS_SCAN_GATE_BATCH_CLIENTS comma-separated client ids of batch jobs (default none)
CORPUS_SCAN_GATE_BATCH_MAX_BYTES largest body a batch client may send to a scanned model; larger gets a 413 (default 65536)
CORPUS_SCAN_GATE_GRPC_MAX_MESSAGE_BYTES envoy-ext-proc: largest gRPC message accepted and returned (default 50 MiB)

For each body, in order: a body that is not a JSON object is refused (or passed, when failing open); a model in SKIP_MODELS passes unscanned; a batch client's oversized body gets a 413; otherwise every message's text is redacted, and a finding in BLOCK_TYPES gets a 403. Text is read from chat message content (including nested tool results and tool-call arguments), Anthropic system, embeddings input and completions prompt. An error inside the policy is refused (or passed, when failing open) rather than left to the proxy. Refusals carry an OpenAI-style JSON error. SKIP_MODELS takes exact names, not patterns, so a typo can only make the gate scan more. Skipping trusts the body's model field, so it is sound only where the gateway routes on that same field: a request naming a local model must not be able to reach a cloud backend. Do not list a model if any route sends that name off the network, and keep path-routed cloud endpoints off the skip path. The client id is the caller's x-client-id request header, which the gateway sets after authenticating the API key.

The envoy-ext-proc adapter expects the request body in Buffered processing mode, where the proxy sends the whole body as one gRPC message, so CORPUS_SCAN_GATE_GRPC_MAX_MESSAGE_BYTES must be at least the proxy's body buffer limit. Request headers must be sent too (the default), so the adapter sees x-client-id. Only the text of each message and content part is redacted; image and audio parts pass through byte-identical.

Sanitized tier

corpus sync projects enriched documents into a separate, trust-downgraded database: summaries and the priority signal, never raw content, subject or sender. corpus index serves that database to cloud-side consumers over MCP.

Variable Purpose
CORPUS_SANITIZED_DATABASE_URL DSN the sync writes to
CORPUS_INDEX_DATABASE_URL DSN the index server reads with (a read-only role)
CORPUS_SANITIZED_DB_SCHEMA schema of the sanitized messages table; empty means CORPUS_DB_SCHEMA
CORPUS_INDEX_SENSITIVITY_GATE sensitivity at which free-text summaries are withheld (default high)
CORPUS_TIER_ACCESS JSON map of tier name to the principals allowed to reach its MCP surface, e.g. {"sanitized": ["orchestrator"]}; a tier without an entry grants no one

Set CORPUS_SANITIZED_DB_SCHEMA when several raw schemas feed one sanitized view. For example, one raw schema per mailbox (mbx_kasserar, mbx_post, …) can each run corpus sync into a shared kbl schema. Record ids include the source (imap:<mailbox>::…), so rows from different mailboxes never collide.