@agent-core

Providers and models

The multi-provider model layer: the declarative catalog, thinking controls, custom endpoints, and credential scoping.

agent-core speaks to every LLM through one provider abstraction. A StreamProvider wraps a single model and yields a normalized event stream (text, thinking, tool calls, usage); everything provider-specific — request framing, thinking parameters, caching markers, usage field names — lives in the per-provider adapter and in the model's declarative catalog entry, never in the runtime loop.

Provider families

The catalog covers these provider families, each with its own stream adapter:

FamilyNotes
Anthropic (Claude)Native Messages API, adaptive/budget thinking, prompt-cache markers, and a fast-mode variant on the top Opus models (same model id, premium rates).
OpenAIAPI-key models on the Responses surface (the adapter streams every catalogued model through /v1/responses).
OpenAI CodexChatGPT subscription sign-in via OAuth — separate backend, subscription-billed.
Google GeminiThinking budgets and levels, explicit cache lifecycle.
xAI GrokReasoning and non-reasoning variants are separate model ids, paired in the catalog.
DeepSeekOpenAI-compatible endpoint.
Moonshot KimiOpenAI-compatible endpoint, including the always-on multimodal K2.7 Code coding flagship with a highspeed fast-mode variant.
Alibaba QwenOpenAI-compatible endpoint with implicit caching.
MiniMaxAnthropic-compatible endpoint, with a high-speed variant per model (the highspeed variant is priced separately).
Custom endpointsAny Ollama-native or OpenAI-compatible server you configure — your own hardware (Ollama, LM Studio, vLLM, llama.cpp, LocalAI) or a remote gateway with an API key.

Cloud models are static catalog entries. Custom-endpoint models are discovered dynamically from each configured endpoint and get ids of the form local:<endpoint>:<modelName>. Per-model overrides let an admin toggle tool support, thinking style, visibility, pricing, context window or output ceiling per discovered model, and Ollama-native endpoints carry a keep_alive setting so the chat model stays resident in VRAM across an agentic loop.

Adding a custom endpoint

Endpoints are managed in the admin Credentials tab, under Custom endpoints (it needs the platform-config feature — the list is platform-global, so it affects every project on the instance). Each entry carries:

  • a name, a base URL and a framing (ollama-native or openai-compat);
  • authentication — None, or a bearer API key. The key lives in the credential store as llm.endpoint.<id>, never in platform config and never in an environment variable. It resolves like every other provider key, most specific first: agent → project → user → global. A member can therefore bring their own key for an endpoint, and only their own runs use it;
  • an address class — private (your own machine or network) or public. A public endpoint's address is checked on every connection, and the connection is pinned to the address that was checked, so a hostname that changes where it points cannot redirect the request to something internal. Saving one through the admin form checks it there too, so a mistyped address is refused immediately instead of failing unexplained on the first run. A private endpoint is an explicit operator exemption from that check — which is the whole point of pointing at localhost;
  • optional extra headers — a routing tag, an org id. Keep them non-secret: the platform refuses the names it owns itself, but anything else you put there is stored as plain text in the platform configuration, so a key belongs in the key slot instead;
  • an optional model list and optional pricing.

Discovery uses the platform-global key, because the model list is fetched once per endpoint and shared by everyone. Member- and project-scope keys are used for streaming only. If you cannot set a global key, list the models explicitly on the entry — a declared list skips discovery entirely.

Pricing decides what a $0 usage row means. Give the endpoint a price (or a per-model one) and its rows bill as metered like any cloud model. Leave it empty and an unauthenticated endpoint on a private address is local — your own hardware, where $0 is the truth — while a keyed or public endpoint is unpriced: the usage is recorded and flagged rather than passed off as free. Spend limits are expressed in dollars, so they cannot bound an endpoint you have not priced.

The declarative model catalog

Every model is a ModelInfo entry. The stream adapters read these fields instead of pattern-matching model ids:

  • thinkingStyle — how reasoning is requested: none, adaptive, budget (token budget), effort-param (Responses-style effort), level, separate-id (paired reasoning/non-reasoning models), or always-on.
  • supportedEfforts / defaultEffort — the reasoning-effort values the model accepts (from minimal up to xhigh and max). The UI only offers what the entry lists.
  • temperaturePolicy — allowed range, fixed value, forbidden-with-thinking, or rejected-by-API; the adapter never sends what the policy forbids.
  • contextWindow / extendedContextWindow — standard window plus an opt-in larger window (with the beta header when one is required). A local model whose endpoint does not report its context length leaves the field unset rather than inventing a number.
  • pricing — real provider rates in $/1M tokens: uncached input, output, cache-read, cache-write, optional long-context tiers where a large request re-prices the whole call, and — where the provider prices audio input at a premium (Gemini's flash family) — separate audio-input and cached-audio rates. A model whose provider publishes no cache discount carries its input rate as its cache-read rate, so a cached token on a priced, selectable model is never billed at $0.
  • billing — how the model is paid for, and therefore what a $0 usage row means: metered (the catalog carries the provider's published list price), subscription (a plan pays for it elsewhere, so usage records no per-token cost), local (your own hardware — there is no vendor rate to record) or unpriced (a catalog entry whose provider publishes no per-token price, or a keyed or public custom endpoint the operator has not priced — an unpriced keyless endpoint on a private address is local). Every entry declares one, and a 0 / 0 price is never metered. Usage that cannot be priced is logged rather than passed off as free.
  • tier — one of flagship | value | mini | nano | code, used for grouping and sorting in the model selector, plus per-model value/capability comparison scores (1–10) and the list price for the selector's quick-compare view.
  • capability — chat (default) or embedding. Embedding models are filtered out of the chat model selector and surface only where embeddings are configured.
  • capabilities — per-modality media-input descriptor (image / audio / video / pdf), with verified per-model limits where the provider documents them. A modality is listed only when the platform can actually deliver it to that model's API — and a single delivery boundary in the stream runtime decides, per request, whether each image / audio / video / PDF block is passed natively, adapted, or replaced with an honest text marker. Tools and connectors always produce full media; models that cannot ingest a modality receive a clear marker instead of silently losing content.
  • fastVariantApiId / fastVariantRequest / fastVariantPricing / pairedVariantId — fast-mode switches. A high-speed variant is reached either by swapping the wire model id (fastVariantApiId) or by keeping the id and adding a request parameter plus a beta header (fastVariantRequest); exactly one applies per model, and either way the request is billed from fastVariantPricing when fast mode is on. pairedVariantId is the paired non-reasoning model id for families that split reasoning into a separate id.

Individual model ids are deliberately not listed here — the catalog tracks provider releases, and the live list is always available from the models route and the chat model selector.

No model catalog is rendered into an agent's context either, which is a deliberate shape rather than an omission: an id an agent has not looked up is a guess, and a guess that misses does not run a weaker model — the run ends without completing a step. Agents read the same route the UI does, through the read-only inspection skill's model script, before setting a model on a subagent, an agent's default, or a workflow.

Availability and selection

Model policy can be set at three layers. An explicit conversation selection wins per field, then the agent default, then the provider/catalog default. Agent Config can set the default model, reasoning effort, an exact catalog-declared base or extended context, and fast mode; clearing a field restores inheritance, while an explicit Fast Off overrides an inherited On. Controls are shown from declarative catalog capabilities, and incompatible combinations are rejected rather than silently adjusted.

An agent default is inherited into every conversation, including ones running a different model. When the model that actually runs cannot honour an inherited or stale field, the stream drops that one field and says so in the conversation — a model_policy_skipped note naming the field and whether it came from the agent or the request — instead of failing the turn. Everything else in the request is unaffected, and the value stays stored for the models that do support it. The same resolved policy is what the usage meter prices, so a fast-mode run is always billed at the fast-variant rate.

A model is available when its provider's API key (or OAuth connection) resolves from the credential store for the caller asking. The registry checks availability per provider — one credential lookup for the whole family — and marks every catalog entry accordingly; the models route and the chat UI only offer available models.

Availability is therefore personal, not a property of the deployment. It is resolved through the same most-specific-wins scope cascade described below, so a member who stores their own provider key (see credentials) sees that provider's models offered to them and to nobody else, and the model the stream actually runs on is gated by the very same lookup that fetches the key — the picker cannot offer a model the request would then reject.

Conversation selection is the highest-precedence layer: the chat composer sets model, reasoning effort, fast mode, and an exact catalog-declared base or extended context window for the next request. Any absent field falls back independently to the agent's configured default policy, then to the platform/provider default — the Default Model platform-configuration setting, editable by an administrator without a restart. The DEFAULT_MODEL environment variable is its boot fallback, used only where platform configuration is not yet loaded.

Credential scoping

Provider keys are encrypted credentials, never environment variables — see credentials. Every key lookup carries the caller's scope (userId, projectId, agentId) and resolves most-specific-wins: agent → project → user → global. Two projects on one deployment can use different OpenAI keys, and a single agent can be pinned to its own key without affecting anyone else.

Resolved keys sit in a short-lived in-process cache (five minutes), keyed by credential and scope. Credential writes, rotations, and OAuth reconnects invalidate exactly the affected scope entries through a credential change bus, so a rotated key takes effect immediately without flushing other tenants' entries. The cache and the OAuth refresh machinery (rotation persistence, automatic purge of permanently-broken credentials) are anchored as process-wide singletons, so every server module instance shares one coherent state — a mid-stream credential purge banner also names the exact project whose login expired, and reconnecting targets that scope directly.

Usage metering

Each provider's parser normalizes usage into the same five billable categories: uncached input, cache reads, cache writes (5-minute and 1-hour TTL where the provider distinguishes them), and billable output. Reasoning-token counts are display-only and never billed, and provider-reported totals are recomputed rather than trusted. The per-stream usage meter prices each step against the catalog entry's pricing block, including long-context tiers. Where a provider bills audio input at a premium, the parser additionally partitions the input into text and audio buckets (a partition of the same totals, never an addition) so each is priced at its own rate.

Codex subscription usage limits

When you sign in to Codex with a ChatGPT subscription, the provider reports your subscription usage limits as one or more windows, each a percentage used. The windows are dynamic — their real length comes from the data (a 15-minute window, a 5-hour window, and a 7-day window are all possible), and every label is derived from that length rather than a fixed assumption.

Neuralis reads this state three ways: from the live stream, and — when you open the chat config Usage Signal panel — with a single on-demand refresh from Codex. That refresh persists the latest reading (keyed to the workspace the credential resolves for — project, then user, then the instance default — so members sharing a project's Codex account see one shared figure), and if the on-demand read is unavailable it falls back to the last stored reading, marked cached. There is no background polling; while a stream is running, its usage updates the open panel directly. Each window's label sits inside a small ring that fills as the window nears its refresh — hover it for the exact time until reset. Crossing 80%, 90%, 95% or 99% during a stream raises a brief, auto-dismissing notice above the composer so you are not surprised by a limit mid-task.

Codex prompt cache

Codex models keep their prompt cache across the whole agentic turn and across turns of the same conversation: the adapter replays the model's own reasoning and tool-call items exactly as the backend produced them, and runs each conversation on one persistent connection with server-side continuation, the same wire the official Codex CLI uses. A background delegate gets its own connection; a lost connection is reopened on the next step. Three platform settings bound it — codexWebsocketEnabled, codexWebsocketMaxSockets and codexWebsocketFrameTimeoutMs (the stall bound, not the model's thinking time) — and the stream timing table names the transport of every step.

Codex subscription savings

Codex models are billed through your ChatGPT subscription, not per token, so their usage records at $0 — and, for the same reason, a Codex login is outside the USD spend caps and the credential-use counter, wherever it is stored. A Codex account configured at the instance-wide (global) scope is therefore available to every project unmetered; setting one requires the cross-scope permission, which no role holds by default. To show the value the subscription is producing, Neuralis computes the savings — what the same tokens would have cost at OpenAI's API pricing for the equivalent model — and shows it (in an accent colour) wherever a Codex model would otherwise read $0. The admin dashboard adds a per-subscription view: total savings and a day-precise "has it paid off" bar that compares the savings against the plan's estimated monthly fee, pro-rated onto whatever time range you select. Plan fees are estimates from the reported plan tier; the savings figure itself is derived from the real token counts and the catalog's list API pricing, never promotional rates.

The savings count Codex tokens only. When a Codex-model agent delegates work to a subagent running a paid model, that subagent's tokens are priced on its own model — a real, billed cost recorded under that model, and one that counts against your spend limits like any other. It is not folded into the subscription's savings, and a spend cap set on a project using Codex will see it. This applies to usage recorded from this release onward: earlier records keep the attribution they were written with, so a long-running installation's historical savings figure still includes some delegated work.

When a turn fails or is cut short

Provider failures are surfaced uniformly across every model rather than ending a turn silently. Codex connection errors retain their safe provider detail, and a connection that closes before a terminal response is treated as a recoverable failure rather than a completed turn. If a request exceeds the model's context window, the chat shows a Context window full banner with Summarize now and Continue — summarizing the earlier turns frees the window, while Continue remains available. Errors and turns cut short for another reason (output-length cap, a content filter) also show a Continue button. The context-window indicator next to the model picker updates live while a turn streams, so you can see the window fill in real time — and as it approaches the summarize threshold, capable models are nudged to compact the conversation on their own before it overflows.

Stopping a turn stops the delivery, not the generation. Pressing Stop — or a subagent hitting its time budget — closes Neuralis's side of the provider connection, so the turn ends promptly and nothing further is written. It does not send a cancellation to the provider, so the model may finish generating on their side and the request may still be billed by them. Treat Stop as "this run is over", not as "that request never happened".

And a stopped step is recorded, not written off. The provider had already read the request by the time you pressed Stop, so the step books what the provider reported it used — and where a provider reports nothing until the answer finishes, it books the size of the last measured request as the input instead. Nothing is guessed from the text on screen: on those providers the stopped step's output is simply not billed, and the usage record says so rather than presenting the gap as free. The same holds for a browser disconnect, a server restart taken mid-answer, and a provider error partway through a step. A turn cut short by a server restart says so — it reads as stopped by the restart rather than by you, and it can be continued.

A provider that asks you to wait is waited on. When a provider says "retry in N seconds", Neuralis honours that exact figure instead of guessing, and each step gets its own small retry allowance so one blip early in a long piece of work does not spend the budget for the rest of it. Every one of those waits is interruptible: Stop ends the turn immediately rather than after the wait.

Titles, rewrites and summaries run on your conversation's own model. The three one-shot helper calls — the conversation title, the composer's prompt polish, and the summary — go through the same provider connections a turn does, so anything you can chat with can also title and summarize: a ChatGPT subscription, a Responses-only model, a model running on your own hardware. By default each one inherits whatever model the conversation is set to; the context-window popover can pin the summary to a specific model instead, and the polish control has its own picker. Titles and rewrites deliberately run at the model's lowest reasoning setting — they are short answers on a short clock — and the summary keeps the conversation's full setup, because it is the memory the next turn works from.

A conversation is named the moment you send. The name comes from your own first message immediately, with no model call in the way, and the model's shorter title replaces it a few seconds later. Stopping the turn before it answers, or a turn that is refused, keeps the first name — a conversation never stays "New conversation". If you rename it by hand in the meantime, your name stays.

Summarizing never loses history. A summary is a non-destructive projection: the full transcript is always kept on disk, shown whole in the timeline, and readable by the agent. While a summary runs, the composer shows a "Summarizing older turns…" banner with a Stop button, and the input is briefly locked until it finishes. When a summary is active the model reads it in place of the older turns; a Show full control switches the model back to the complete history at any time (and Resume summary switches back), without ever truncating the conversation.

A summary is checked against the record, not only remembered. Before the model writes anything, the platform reads the conversation for what it verifiably did — files changed with their line counts, every command run (commits marked), background runs and their logs, subagent runs and their transcripts, messages received, tool failures — and hands that to it as established fact. The same record is available on its own, so an agent can check what happened rather than recall it. You can also give a conversation a standing summary focus in the context-window popover: a sentence about what its summaries should always keep, which every later summary follows, including the automatic ones that run with nobody watching.

On this page