@brain-core

Sources and connectors

Source kinds, the URI vocabulary, and how files resolve across environments.

A source is a mounted data location: a project's data folder, an extra host directory, the vector memory, a machine sandbox. Each source has a slug, and that slug is a URI scheme — every file in the platform is addressed as <source>://<path>. A connector is the driver behind a source kind; the runtime contract is described in the connector port. This page covers the vocabulary brain-core ships and how source configuration behaves in practice.

Govern external data before connecting it

Sources can hold an organisation's shared corpus or one user's private data. Before copying or syncing, establish the dataset owner and approved location, intended audience and source scope, sensitivity and retention, credential scope, copy permission, and provenance/refresh record. If copying is not authorised, do not create or index the source. Public availability does not by itself permit collecting sensitive personal data, and external text is data rather than trusted instruction.

Corpus content belongs only behind a configured source/connector with explicit policy and sync lifecycle. A writable project-data folder is not automatically a governed corpus. Agent identity and durable memory must not contain dataset rows or shadow copies; memory can retain a compact canonical source-URI pointer with owner/scope, provenance, permitted use, refresh date and governing retention/policy reference. Packages add skills, deterministic tools and bounded workflows around those source URIs instead of embedding the data in package prose.

The coordinate four-tuple

Every file artifact lives at the coordinate (uri, connector, os_uri?, container_dir?):

  • uri — <source>://<path>, the stable identity used by tools, routes, the UI, and the vector index.
  • connector — the connector instance bound to that source slug at boot.
  • os_uri — optional full host-OS path, produced by connector.resolveOsUri(uri). When the platform runs inside a container and no host mirror path is configured, the connector throws a structured OsUriUnresolvableError instead of fabricating a path — the UI and the vector index persist os_uri: undefined rather than a stale guess.
  • container_dir — optional container-side path counterpart, used when a shell or skill script needs to execute against the file.

The host-versus-container split is why the local connector's config has two fields: root (the path the server actually reads, required) and hostRoot (the host-side mirror when the server runs in Docker, optional). The source listing shows these paths only to someone who could attach that source or can read its root, and the last sync's error text only to someone who may run syncs; everyone else sees the connector kind and a generic "the last sync reported errors".

URI vocabulary

SchemeConnector kindWhat it addresses
data://localThe project's data zone — the default working area for agents and uploads
brain://brainVector-only memory backed by the index, agent-scoped by default
packages://localThe project's installed package directories
app://localThe platform app zone — a privileged mount (drive.mount.privileged); agent writes are denied
machine://webtopA machine sandbox, contributed by machine-core
custom slugsanyEvery additional source you attach — the slug you choose becomes the scheme (repo://, notes-alex://)

Connector kinds shipped by brain-core

local (LocalDiskConnector) is the full-capability on-disk source: list, read, stat, write, delete, move, mkdir, rmdir, scanDelta, walk, count, peekCount, scanContent, subscribeEvents, resolveOsUri — every capability in the port except exec and its execDetached companion. stat answers "is this still there" without transferring the file. It blocks symlink traversal and .. escape and refuses explicit access to the platform app zone. Code execution is NOT a connector capability — shell/skill scripts run through the execute tool's shellRunner, gated per-path by the uri-policy exec permission, not by the connector. Sync excludes keep secrets, lockfiles, .env files, key material, dotfile credential stores, browser profiles, and build artifacts out of the index — the credential shapes on a runtime floor that also covers sources attached earlier (for what gets indexed next — it does not remove what was embedded before it shipped), the rest as an editable per-source default.

brain (BrainInternalConnector) is the Qdrant-backed internal memory. It is discoverable like any other kind, but reads and writes are serviced by the filesystem service against the vector index — there is no on-disk tree behind it, and no os_uri. Viewer roles are read-only on brain:// by default.

Other packages contribute further kinds through their manifests — for example the webtop kind from machine-core (addressed as machine://). The registry, capability checks, and the trust gate are the same for every kind.

Source scope and naming

A source belongs to a project, a user, or an agent. Each connector kind declares which scopes it allows (allowedScopes) and the default; project-scoped sources are visible to every project member, while user- and agent-scoped sources require ownership or a scoped-observation feature grant.

Source slugs must be unique because they are URI schemes. The Files UI suggests scope-qualified keys when you attach a source: a project source named after its folder (repo), a user source qualified by the user (repo-alex), an agent source qualified by the agent (repo-nova). Never reuse one slug across scopes — the scheme is the identity.

Declarative sources and package attribution

Packages don't just register connector kinds — they can also declare the concrete source instances they own, right in their manifest. A declaration names the slug, the connector kind, the root, the scope, and a required one-line description, and it marks the source as either auto (created in every project at initialization) or discoverable (surfaced for on-demand attachment). This is how the project data zone, the package directory, and the vector memory exist in a fresh project without anyone attaching them by hand, and how the platform app zone is offered as a privileged add — gated by drive.mount.privileged, which no role below the admin tier holds by default.

Because the instance is declared, every source carries a contract-backed owning package — its origin_package. The platform resolves it by matching a live source against the declarations, and the match is deliberately unspoofable: a rooted on-disk source must match BOTH the declared slug AND the resolved root, so a user mount that merely borrows a well-known slug (say, naming a mount app while pointing it somewhere harmless) is never mistaken for the package-owned source. Rootless sandbox kinds (vector memory, machine sandboxes) match by kind alone, since their slugs are per-user.

origin_package appears on the sources listing and on each source's runtime-stack row, where the agent sees it as a [pkg: …] tag — so the Sources panel can group sources by the package that owns them, and an agent can tell a first-party source from one a teammate attached. A source's description follows the declaration too: an explicit per-source description wins, otherwise the agent falls back to the declared description, then the connector kind's default — so a runtime-stack row is never blank.

Enabled toggle versus URI policy

Two independent controls answer two different questions:

  • URI policy answers which URIs may this caller touch. Per-source path rules are evaluated against the caller's role and agent at four enforcement layers — route, UI, tool (including the sandboxed shell), and the vector index — see URI policies. Connectors are deliberately not one of them: a connector is the raw transport for a URI scheme, and every gate sits one layer above it. A read: false path is hidden on both axes: its content is refused (read/diff/show), and its name is dropped from the file tree and search — the listing and search agree, so a denied path never leaks its existence. One authoring note about globs: a rule on foo/** hides everything inside foo, but not the folder node foo itself — to hide the folder's name too, add a rule on foo (or **/foo).
  • The enabled flag answers is this source live at all. SourceConfig.enabled is true, false, or 'auto' (default true). A disabled source returns a structured error (code: 'source_disabled', with enable_source as the suggested next step) from every fs_* tool and route, regardless of what the policy would allow.

'auto' exists for machine sources only: the effective state derives from a live probe of whether the machine sandbox is currently running, so a stopped machine reads as disabled without anyone flipping a switch. The toggle is exposed as POST /sources/enabled with { source, scope, enabled } and as a switch in the Files UI sources panel.

  • Connector liveness answers can the connector reach its backing store right now. This is a display-only signal, separate from both the enabled flag and URI policy: when a source's connector is disconnected — a local mount that has gone missing, a sandbox container that isn't running — its row dims (icon and name) in the file tree and the sources panel, and the agent's runtime view marks it offline with a reason. Already-synced content stays listed and readable; the full colour returns once the connector reconnects.

Manifest baselines and the persisted config

When a source is first attached, the filesystem layer unions every loaded package's matching manifest uriPolicies baselines into the source's persisted configuration JSON. From then on the persisted config is authoritative: owners and admins edit rules through the sources panel or the policy routes, and the manifest baseline is only re-seeded when a source is added or replaced.

Immutable floors are the exception, on purpose. A restrict-only floor rule from a trusted manifest — or from a source connector's own declaration — is re-asserted on every boot, so a floor added by a newer package version reaches sources that already exist, and one that went missing comes back. Ordinary rules never do; if you edited a path rule away, it stays away — and a no-op policy edit no longer disturbs the repair bookkeeping, so a rule you deliberately removed is not resurrected by a later restart.

A floor also decides indexing, not only reading: what cannot be read cannot be embedded, so a path a floor denies read on is never newly embedded into the vector store. That is one declaration covering both axes — nothing needs to be added to a sync-exclude list, and nothing editable can undo it. The merge semantics live on URI policies; the conversational editing workflow is the manage-uri-policy skill on brain-core skills.

Each source also carries an optional one-line description, editable from the sources panel and the source routes. Agents see it in their runtime-stack source table, so a good description directly improves how an agent picks the right source — see runtime stack.

Discoverable sources

The "add a source" picker offers a union of two suggestion families. The first is declared discoverable sources: a package surfaces an on-demand source — for example the privileged platform app zone — and the declared label and description prefill the add form. These suggestions carry their owning origin_package. The second is host-mount suggestions: directories the operator has bound into the runtime (the Neuralis tree, the projects zone, the whole host, or any named mount). Those depend on a runtime mount rather than a package, so they have no owning package. Either way, the access controls that gate seeing a high-blast-radius suggestion and attaching a source rooted in a protected zone remain authoritative — a declaration can tighten visibility, never loosen it.

A suggestion may also carry defaults for what happens after the attach: a sync include allowlist, a sync trigger, and uri-policy path rules. Path rules offered this way are ordinary editable rules, never immutable floors — only a trusted package manifest can author one of those, and a create request carrying a floor marker is rejected. The distinction that matters when reading a suggestion: an include narrows what is indexed, while what an agent can read is decided by the source root and the path rules. A wide root with a narrow include is still a wide root. Where a zone offers both a narrow and a wide proposal, the narrow one is listed first.

Deleting a source

DELETE /sources/:slug first cancels any running or queued sync job for the source (so an in-flight sync can never write entries for a configuration that no longer exists), then removes the source-config JSON and the in-memory connector binding. The underlying folder is never touched — brain-core owns the wiring, not your content. Vector index points that belonged to the source are left in place by default; pass ?purge=1 (the Files UI exposes this as a Purge vector index entries checkbox, off by default) to remove them. The purge runs as a background job with a single bounded index filter — the delete call returns immediately with the job reference instead of blocking on index size. Leftover orphan points can always be cleaned up later through the vector health repair tools — see memory and sync.

Sources are config, files are content

Attaching, disabling, or deleting a source never modifies the files behind it. Source operations edit configuration and index state only — content mutations go exclusively through the write tools and routes, where policies and approvals apply.

On this page