@brain-core

Memory and sync

Vector memory, embedding, and the sync pipeline.

brain-core keeps a vector index of your sources so agents can search by meaning, not just by filename. The index is multi-tenant — every point carries the owning projectId, userId, and (for memory entries) agentId — and is backed by Qdrant in production, with an in-memory substitute for tests and degraded operation. This page explains how content becomes searchable, when sync runs, and how to keep the index healthy.

The embedding pipeline

Embeddings are produced by a declarative model registry: every supported model is an explicit record with its provider, protocol, default dimension, supported dimensions, and credential id. Supported providers are Ollama (local), OpenAI, Gemini, Qwen, and Voyage, plus a deterministic offline embedder for tests. Two carried models are multimodal — Gemini Embedding 2 (images, PDF pages, audio and video) and Voyage's multimodal model (images and PDF pages) — while every other carried model is text-only. Models that tell queries and documents apart are told which one each call is.

Embedding has its own credentials, separate from the chat keys: embedding.openai, embedding.gemini, embedding.qwen and embedding.voyage, kept in the credential store — packages never read embedding secrets from the process environment. They are read at the platform (global) scope only: indexing is background work with no user in context, and one collection serves every project, so a member's own value for the same id is never consulted. Because they are not LLM keys, the per-credential use limits count them, and the platform meters embedding spend in dollars per call: embeddingDailyUsdCap stops embedding for the rest of the UTC day once today's priced spend reaches it (sync pauses with a resumable error, semantic search falls back to full-text; 0, the default, means no cap).

Any server that speaks the OpenAI /v1/embeddings protocol can be added as an embedding endpoint — from the admin Vector section or during setup: its URL, its models with the dimension each returns (and optionally a price), its authentication, and its address class. A keyed endpoint's key is the credential embedding.endpoint.<id>, never part of the entry, and its models appear in the model list as endpoint:<id>:<model>. An endpoint whose model is the active one, the target, or a running rebuild's target cannot be removed, disabled or renamed (409 endpoint_in_use, nothing written) — switch the index away from it first. A public endpoint is address-checked and IP-pinned on every call, exactly like the cloud providers; private is the operator's explicit statement that the address is inside their own network. The two cloud base-URL overrides (openaiBaseUrl, qwenEmbeddingBaseUrl) accept public addresses only — a proxy or gateway inside your network is added as a private embedding endpoint instead.

Many models expose MRL (Matryoshka) selectable dimensions: their supportedDimensions list (or range) lets you pick a smaller vector — trading a little quality for less storage and faster search.

Three properties make the pipeline safe to operate:

  • One active model, recorded. The model that built the collection is recorded with it and stays the active model: every write and every search embeds with it, and every vector written is stamped with the model that made it, so vectors from two models never mix in one index. The configured embeddingModelId / embeddingDimension is only the target. Changing it starts nothing: the admin Vector section reports the switch as pending — how many points sit on the active model, the target, and on request a cost and time estimate computed from a sample (a range, not a quote) — and waits.
  • A model switch is an explicit rebuild. An owner starts it from the Vector section (or the rebuild-generation collection action below). The target model embeds eligible points again from stored text, or from source bytes when inline content is unavailable or the input is media, into a new collection while the old one keeps serving. Retained recovery content, including brain:// content and its version history, is preserved. An already recovery-only record whose input is unavailable stays preserved without vectors in the new collection, including during a source outage; verified input restores its searchable embedding. Unmarked missing input still pauses the rebuild. Writes and deletes that happen meanwhile are applied to both, and when the new collection is complete the two are swapped in one step; the dimension may change freely. A rebuild can be cancelled (the new collection is dropped, the active one is untouched), and one interrupted by a restart waits, paused, until it is resumed. The previous collection is dropped after the swap unless vectorKeepPreviousGeneration keeps it as a rollback copy (at double the storage). While it runs, storage is doubled and every new write is embedded twice.
  • No silent fallback. A missing credential raises EMBEDDING_CREDENTIAL_MISSING naming the credential id and its scope, and an unsupported dimension raises UNSUPPORTED_EMBEDDING_DIMENSION. The platform never swaps in a different provider or another key behind your back. Nothing is embedded without the active model's key; semantic search then answers with full-text results and a warning that names the missing credential, rather than failing the call.

pnpm reset:vector (and the admin Vector section's confirmed reset) is not the way to change models: it drops the collections and the model record, and with them every brain:// file and its version history, which live only in the vector store. Connector-backed sources come back on their next sync, embedded again at provider cost. Keep it as a last resort.

Each indexed file produces a parent document point plus content chunks, so semantic search can rerank chunk hits against document hits. The parent keeps only the file's first 200k indexed characters; the chunks hold the rest, which is where every search mode reaches it. Chunks are skipped by content hash when unchanged, which keeps re-syncs cheap.

That skip is also why the per-source controls are three different things. Rescan changes (sources/reindex) walks the whole source again, but a file whose content did not change keeps its vectors. Forget sync state (sources/reset) only clears what the last sync saw and cancels a running job, so the next sync compares every file — it deletes no file and no vector. Re-embed… is the one that spends: it reads every file in its scope — a whole source, a folder or one file — and embeds it again with the active model, changed or not. It shows a sample-based cost and time estimate first and starts only on confirmation; it is offered in the Sources panel, on the Files tree's context menu, and to agents through the manage-sources skill, and it needs drive.sync like every sync control. Re-embed is not offered on brain://, whose vectors are the content itself.

Media. With a multimodal active model, sync also embeds media from the file's own bytes: images, PDFs (one embedding per rendered page, up to brainSyncMaxPdfPages, default 20), and — only where you switch them on — video and audio. Each source chooses its kinds in its sync settings (sync.media {images, pdf, video, audio}; images and PDFs default on, video and audio off), and a file larger than brainSyncMaxMediaMb (default 8 MB) is skipped. A kind is embedded only when the ACTIVE model supports it, so on a text-only model nothing is sent and the Sources panel says how many media files were not embedded and why. Structurally invalid PNGs are skipped locally before provider work and counted as invalid when the media lane is invoked; existing indexed and recovery content is preserved. This is a structural preflight, not a pixel-decoder guarantee. PDF pages are rendered from local-disk sources; embedding endpoints (the OpenAI /v1/embeddings protocol has no media input) embed text only. A media hit in search carries its kind and, for a PDF, its page.

When content is indexed

Writes that go through brain-core (fs_write, uploads) update the vector index inline — the file is searchable as soon as the call returns. The background sync exists for everything else: files changed on disk by shells, skills, external editors, or other processes.

Inline writes honor the same exclude/include filters as sync: a file an agent creates under an excluded path (say inside node_modules) is written to disk and tracked for review, but it is not embedded — exactly as the sync walk would skip it. So what is searchable is decided by one consistent rule, no matter how a file arrives. Excluded files are still fully readable and editable through the connector; they simply stay out of the meaning index.

Each source declares its sync trigger in its configuration:

TriggerBehavior
manualSync only when explicitly triggered (UI button, route, or skill)
autoEvent-driven: syncs on the first run, when the filter configuration changes, and on a periodic full-reconcile cadence (reconcileEvery, default 24h) — never a full tree walk on every tick
interval:<N><s|m|h|d>Sync at most once per interval (for example interval:15m)

The legacy on_change trigger (which re-walked the whole tree every scheduler tick) was replaced by auto; existing configurations migrate automatically when read.

auto sources are watched at the operating-system level: any change — an agent write, a shell command, a git checkout, an external editor — marks the touched files dirty, and after a short debounce a targeted sync reconciles exactly those files (read, re-embed if changed, tombstone if deleted, move detection for renames). Excluded trees are ignored at the watcher level too, so they cost no OS watch resources. If the watcher ever loses events, the source falls back to a full reconcile rather than trusting the stream.

Filters are two lists. Exclude patterns (gitignore syntax) always apply on top of the built-in defaults. An optional include allowlist narrows the source further: when non-empty, only files matching an include pattern sync — and an exclude match still wins, so an excluded path can never be re-included. Either edit changes the configuration hash and forces a reconcile pass on the next run, and every full sync removes the index search data of files the filters no longer cover. Pattern/media exclusions remove obsolete parent/chunk search data, including modern parentless chunks in the same project and source. Chunk cleanup has its own count and refreshes Files without inflating removed-file counts; legacy chunks without source ancestry need explicit health maintenance. Recovery-bearing parents — pending/deleted records, saved baselines and sole-copy brain:// content — retain their complete payload with recovery_only:true, no embedding stamp and no dense or BM25 vectors. Retrieval, browsing and review remain policy-gated; search, enrichment and source-package discovery exclude retained recovery points. Re-inclusion produces a real new embedding before clearing the marker. Enabled/auto inactivity never purges an index, and an immutable READ floor alone never authorizes cleanup. Generation rebuilding classifies current source policy before provider work and again at its config-coordinated swap: unknown provenance or unavailable/invalid configuration pauses without swapping. The dollar estimate samples only eligible content and exposes protected/excluded/unresolved counts; it remains a range, not a hard spending bound. A new project's data source starts with the platform's numbers-only machine-written files (usage counters) already excluded; an existing project's filters are never rewritten. Sources also offers these defaults when enabling sync on its own data source with no sync block; Save applies the editable proposal. Custom sources and explicit-empty configured lists keep their choices.

Every source walk — for indexing and for the displayed file count alike — stops at an effective scan cap: a per-source Max scan files override, falling back to a platform-wide default (200 000), so a very large source never walks unbounded. The displayed connector count applies the same exclude set as the sync, so it counts the files actually eligible for indexing rather than the whole tree. When a walk reaches the cap the count is shown as N+ (the source genuinely holds more than the cap allows) — raise the cap to index and count beyond it.

A platform-wide auto-sync mode (admin setting) controls the background scheduler: on (default), off (manual sync only), or packages-only — in which only sources pinned always active keep auto-syncing (package-content sources are pinned by default). Manual sync is available in every mode.

A background scheduler polls source configurations every 60 seconds (tunable via BRAIN_SYNC_POLL_INTERVAL_MS) and enqueues due jobs on a shared sync queue. The queue is the single execution path for every sync-shaped operation — scheduled runs, the UI Sync button, reindex, and delete-time purges all go through it. It runs a bounded number of jobs at a time (default two), rotates fairly across projects so one large source can never starve another project's sync, deduplicates repeat requests per source, and gives every job its own cancellation handle: triggering a sync returns a job reference immediately (HTTP 202), and cancelling stops the run within moments — mid-walk or mid-embedding — instead of waiting for it to finish. A cancelled run is not an error; everything indexed before the cancel stays valid and the next run converges. The runner walks the connector's changed-files-since-checkpoint delta (the checkpoint token from the previous successful run is persisted and passed back, so nothing slips through the gap between scan and completion), embeds new and changed files, removes deleted points, and detects renames by exact content-hash match: a moved file keeps its index entry and every chunk embedding, and only its file-level vector is refreshed when the name changed. An append that leaves the first 200k characters as they were re-embeds only the new tail. Default excludes keep node_modules, package-manager content stores (.pnpm-store, .yarn/cache), build output, lockfiles, .env files, and key material out of the index, and per-file caps (5 MB, 200k indexed characters) bound the embedding cost. Append-shaped files (.jsonl, .ndjson — conversation transcripts, event logs) are treated differently, because they only ever grow at the end: they have their own, higher size limit (brainSyncMaxAppendFileMb, 20 MB, never above the read limit), they are never cut at the per-file chunk cap, and an append embeds only the new tail — the platform checks that the stored file is an exact prefix of the new one before it trusts that, and re-reads every chunk hash otherwise.

Some of those patterns are more than a default. A runtime sync floor is unioned under every source's own exclude list on every check, so it also covers sources that were attached before a pattern shipped and an include rule can never bring a floored path back. Like every exclude pattern, it also clears what was already embedded under it on the next full sync. It holds the dotfile credential stores and browser profiles — .ssh, .gnupg, .aws, .azure, .kube, .docker, .netrc, .git-credentials, .npmrc, .pypirc, the login keyrings, the NSS certificate and private-key store in both of its locations (.pki and .local/share/pki), the gcloud / gh / rclone config trees, and the Chromium, Chrome and Firefox profiles — written as path shapes rather than as one connector's paths, because a home directory synced through a local source and a sandbox /config synced through a virtual desktop are the same exposure.

What cannot be read cannot be embedded. Beyond that pattern list, indexing consults the source's immutable uri-policy floors directly: a path whose floor denies read is never newly embedded, on every source, with no exclude entry to add and none that can be edited away. The two inputs are different kinds of thing on purpose — the exclude list is a preference you own and can change at any time, while a floor is a security declaration that belongs to the package owning the connector. That is why there is no package-declared "sync-exclude" contract: deriving the index rule from the read floor covers every floored path automatically.

Two boundaries worth knowing. A per-role denial does not keep a file out of the index — somebody can read it, so it stays searchable and each caller's own results are filtered for them; this is why an owner still finds conversation content a member cannot. Pattern cleanup preserves recovery-bearing history while clearing search vectors; an immutable READ floor hides content and blocks new indexing without alone purging old records.

Conversation summaries are a deliberate part of these defaults. A conversation's active SUMMARY.md is indexed like any other file, so an agent can search what it learned by meaning. The immutable summary revision archive behind it, however, is on the same floor — historical revisions are never embedded or searchable, on new and pre-existing projects alike.

After each successful run the source stores an absolute index snapshot (indexed, unindexed, and total file counts) that the UI and the agent's runtime stack read cheaply — no live directory walk is needed to display counts.

Contribution categories at sync time

The indexer assigns every file a category payload field. Beyond the runtime-data categories (conversations, usage logs), files whose path matches the shared contribution layout get a first-class contribution category — the same vocabulary packages use:

Path patternCategory
SKILL.md (any directory) + commands/*.md (legacy single-file form)skill
instructions/**/*.mdinstruction
AGENTS.md, CLAUDE.md, CLAUDE.local.md, GEMINI.md, .github/copilot-instructions.mdinstruction (identity form)
rules/**/*.md, .cursor/rules/*.mdc, .cursorrules, .windsurfrulesrule
agents/**/*.md, *.agent.mdagent
docs/**/*.md (READMEs and templates excluded)docs
package.json, neuralis.package.json, known plugin manifestspackage
.mcp.json, hooks.jsonconnector / hook (indexed only)

These patterns also match inside well-known tool directories (.claude, .cursor, .codex, .gemini, .github, and friends). A markdown keeps its contribution category only when it is admitted: the filename-identity forms (SKILL.md, commands/*.md, the identity basenames, .cursorrules, .mdc) always are; a markdown under agents/, rules/, instructions/ or docs/ and a *.agent.md need frontmatter id or name, otherwise the file is indexed as a plain file. A free-form category a caller sets on a file (note, config) is kept as given; a contribution category is always derived from the path and the content, never taken from a request, and a stale one is re-derived on the next full sync. Admitted markdown contribution files additionally carry a whitelisted frontmatter capture (id, name, title, description, required features, credential ids), manifests carry a usability flag when their content declares a recognized package shape, and files under an agent's private agent-core/<agentId>/ subtree carry an owner marker. Unchanged files pick these fields up through a payload-only refresh on the next sync pass — no re-embedding. This category layer is the discovery substrate for source packages — skills, rules, instructions, agents, and docs loaded from synced sources.

Search semantics

fs_search has three modes, chosen by search_mode:

  • semantic — meaning-based vector search over the index. Requires a warm index; results carry relevance scores. On Qdrant the query is hybrid: the vector ranking is fused with a keyword leg over the same content, ranked by BM25. A fresh install creates its collection with BM25 ranking from the first file (each indexed write also runs one BM25 encoding on the vector server; this needs Qdrant 1.15.2 or later); an older collection gets it when an owner enables text ranking (below). A file not yet covered by it still matches through the plain keyword filter. A search whose scope — your files in a project, or one folder of them — holds at most vectorExactSearchMaxPoints indexed points (20 000 by default) compares the query with every one of them instead of walking the approximate index, so a small folder finds everything it holds; each search counts its scope first to decide.
  • filter — metadata and keyword queries: category, tags, status, and creation-date ranges, plus query as a tokenized, case-insensitive keyword over indexed content — the file's chunks included, one result per file — (results ranked by term frequency) and glob as a filename keyword match. No embedding involved.
  • grep — literal text search (ripgrep, git grep, or grep — whichever is available) over live connector content. Bypasses the index entirely, so it always reflects the current state of disk.

Across all three modes, fs_search enforces the per-URI read policy on every result, not just the drive.search feature gate — a read:false path (another agent's private memory subtree, for example) is silently absent from results, and a fully-denied folder_uri returns a 403, giving search the same path protection as fs_read and fs_write. This makes grep the uri-policy-read-gated counterpart to shell grep: a caller holding drive.search plus read policy on a path can grep it without any execute / shell access.

Reads and listings are connector-first: fs_read and fs_list take live content and tree structure from the connector even when the index is cold, then reconcile index presence per listed file. Presence states distinguish vector-only entries (brain), files where index and connector agree (synced), on-disk files not yet indexed (unindexed), and indexed files no longer present on the connector (stale_indexed).

brain:// memory queries are isolated per agent inside a project: an agent searches its own memory unless sibling inclusion is explicitly requested, and cross-agent observation requires the corresponding feature grant.

Vector health and degraded mode

If Qdrant is unreachable when the package initializes, brain-core wires an in-memory index instead and the platform stays up: connector-first reads, listings, and grep search keep working, while semantic search and sync operate against the in-memory substitute. On Qdrant-backed deployments a background watch re-probes connectivity every 30 seconds, logs loss and recovery transitions, and reports { backend, qdrantReachable, lastCheckedAt } into the host's health endpoint. The live backend is not hot-swapped — after Qdrant recovers, a restart re-enables Qdrant-backed search.

The vector health service is the project-scoped janitor behind the Files UI health panel. It separates cheap from expensive work so the panel never slows the rest of the UI:

OperationCostEffect
StatsCheap (exact counts, no scan)Total / document / chunk / deleted counts, per source — computed from indexed fields, so they stay instant at millions of points
Scan for problemsExpensive (walks every point)Orphan-chunk and duplicate-document detection — explicit, on-demand only
Garbage collectDestructiveDrops records marked deleted that are older than the retention window, except those a pending change still needs for its revert
Repair orphansDestructiveDrops chunks whose parent document is missing or flagged deleted
Find duplicatesRead-onlyReport of documents sharing a content hash

The stats load when you open the panel; the integrity scan runs only when you ask. Garbage collection and repair are destructive and support a dry-run flag; the duplicates report never removes anything.

Garbage collection also runs on its own, once a day by default (vectorGcIntervalHours, 0 turns the schedule off), for every project as explicit system work — never under a user's identity, and whether or not automatic sync is on — and every run is logged with who started it. The schedule's clock starts when the server starts, so the first automatic run comes one interval after a restart, never at boot; running it by hand is always immediate. Deleted records a still-unresolved pending change refers to are kept, so reverting that change keeps working. Garbage collection and orphan repair never run twice at once for the same project: a second request while one is running answers 409 action_in_progress, and an automatic run that meets a manual one skips that project and tries again on the next tick.

The index also watches its own writes. Every write is counted per file over a rolling hour, within a bounded list per project, and a file rewritten vectorWriteHotUriThreshold times or more in an hour is logged as a hot file — the signature of an automated writer that rewrites instead of appending. Members who may change files (drive.write) can list the hottest files through GET health/hot-uris; the list is filtered for the caller exactly like a file listing, so it names only files that caller may read.

One vector collection serves every project, so the work on the collection itself belongs to the owner, in a maintenance window: health/collection requires the platform feature platform.vector.maintain, which no role holds by default. GET reports what the server has — missing or differently configured indexes, quantization, text-ranking coverage — and POST runs one action: rebuild the payload indexes with their declared settings (only the fields vector searches filter on get extra graph structure, which keeps an index rebuild cheap), apply the segment settings (vectorDefaultSegmentNumber, vectorMaxSegmentSizeKb, vectorMaxIndexingThreads), apply vectorQuantizationMode, or add the BM25 text-ranking vector to a collection created before it was the default and backfill it in bounded, resumable batches. Each of these rebuilds index structures on the server; none runs at startup. The same surface carries the model switch: GET also reports the active model, the target, the point counts, a running rebuild's progress, the last estimate (dropped once the target model changes) and today's embedding spend, and the actions rebuild-estimate (a sample-based estimate, changes nothing), rebuild-generation (starts the rebuild, or resumes a paused one) and rebuild-cancel run it — the admin Vector section is the same surface with buttons. Startup creates indexes that are missing, rebuilds a keyword-search index only when its tokenizer settings are wrong, and otherwise just checks and logs what differs. Startup also compares the Qdrant server version with the client the platform ships and warns when they are more than one minor version apart — the admin Vector tab shows the same.

Qdrant does not log its own background optimization, so the platform reports it: every 30 seconds it reads the server's list of optimizer runs and writes one vector.optimizer.completed line per finished run (kind, points, segments and the duration the server measured), plus a warning for a run still going after ten minutes — in the container log as well as the vector log. The Qdrant the setup runs — the generated container or the downloaded binary — also has its anonymous usage telemetry turned off.

The pending-change journal

Every revertible filesystem change an agent makes is recorded in a durable, append-only journal (with a baseline copy of the prior content alongside it). This journal is the source of truth behind the Changes tab: listing, approving, and reverting a change all consult it directly and survive a restart, including files that were never indexed. Listing can fall back to the journal during an index outage; review mutations refuse backend errors rather than treating them as missing state.

Listing pending changes reads two sources in parallel — the durable journal and a query for files awaiting review in the vector index — and merges them, preferring the journal. If the vector index is unreachable, the list falls back to the journal alone rather than failing. Review uses a fresh POST /pending {action:"preview",id} receipt, separating current connector bytes from the recorded snapshot and saved baseline. The preview reports missing/unavailable/binary/too-large content and missing/stored/stale/recovery-only index states within pendingPreviewMaxBytes (default 262144). Approve (POST /approve) and Revert/Dismiss (POST /pending) carry expectedFingerprint; states requiring acknowledgement additionally need acknowledgeCurrentState:true after review. Freshness, tenant, owner, source scope and URI READ/WRITE gates run again at apply. Missing state is never fabricated, backend errors never become absence, and acknowledgement grants no permission. Approve/Dismiss preserve the stored snapshot and index state, including a genuinely absent index; recovery-only metadata uses explicit vector-clearing retention. Revert uses the saved baseline, refuses missing baselines and occupied destinations, and gates both move coordinates. Legacy unsafe pending /write revert:true requests refuse; ordinary non-pending one-level undo remains available.

The journal is never vectorized: it lives in an excluded path, so sync skips it, and it is readable only by owners and admins. Each entry is re-checked against per-file path policy and per-user ownership when it is listed, so a user only ever sees their own changes (unless their role carries the scoped-observation feature).

Source delete and vector purge

Deleting a source removes its configuration and connector binding but leaves its vector points in place by default — they become orphans whose parent files are no longer synced. Pass the purge option on delete to remove the source's points in the same operation, or run the orphan repair afterwards. The full source-deletion contract is on sources and connectors.

Memory entries are files

brain:// is a source like any other: agents write memory with fs_write create, refine it with an edit, and retrieve it with fs_search in semantic mode. There is no separate memory API to learn — the URI vocabulary covers durable notes, insights, and working state.

On this page