Providers and models
The multi-provider model layer: the declarative catalog, thinking controls, custom endpoints, and credential scoping.
agent-core speaks to every LLM through one provider abstraction. A
StreamProvider wraps a single model and yields a normalized event stream
(text, thinking, tool calls, usage); everything provider-specific — request
framing, thinking parameters, caching markers, usage field names — lives in
the per-provider adapter and in the model's declarative catalog entry, never
in the runtime loop.
Provider families
The catalog covers these provider families, each with its own stream adapter:
| Family | Notes |
|---|---|
| Anthropic (Claude) | Native Messages API, adaptive/budget thinking, prompt-cache markers, and a fast-mode variant on the top Opus models (same model id, premium rates). |
| OpenAI | API-key models on the Responses surface (the adapter streams every catalogued model through /v1/responses). |
| OpenAI Codex | ChatGPT subscription sign-in via OAuth — separate backend, subscription-billed. |
| Google Gemini | Thinking budgets and levels, explicit cache lifecycle. |
| xAI Grok | Reasoning and non-reasoning variants are separate model ids, paired in the catalog. |
| DeepSeek | OpenAI-compatible endpoint. |
| Moonshot Kimi | OpenAI-compatible endpoint, including the always-on multimodal K2.7 Code coding flagship with a highspeed fast-mode variant. |
| Alibaba Qwen | OpenAI-compatible endpoint with implicit caching. |
| MiniMax | Anthropic-compatible endpoint, with a high-speed variant per model (the highspeed variant is priced separately). |
| Custom endpoints | Any Ollama-native or OpenAI-compatible server you configure — your own hardware (Ollama, LM Studio, vLLM, llama.cpp, LocalAI) or a remote gateway with an API key. |
Cloud models are static catalog entries. Custom-endpoint models are
discovered dynamically from each configured endpoint and get ids of the
form local:<endpoint>:<modelName>. Per-model overrides let an admin toggle
tool support, thinking style, visibility, pricing, context window or output
ceiling per discovered model, and Ollama-native endpoints carry a keep_alive
setting so the chat model stays resident in VRAM across an agentic loop.
Adding a custom endpoint
Endpoints are managed in the admin Credentials tab, under Custom endpoints (it needs the platform-config feature — the list is platform-global, so it affects every project on the instance). Each entry carries:
- a name, a base URL and a framing (
ollama-nativeoropenai-compat); - authentication —
None, or a bearer API key. The key lives in the credential store asllm.endpoint.<id>, never in platform config and never in an environment variable. It resolves like every other provider key, most specific first: agent → project → user → global. A member can therefore bring their own key for an endpoint, and only their own runs use it; - an address class — private (your own machine or network) or public.
A public endpoint's address is checked on every connection, and the
connection is pinned to the address that was checked, so a hostname that
changes where it points cannot redirect the request to something internal.
Saving one through the admin form checks it there too, so a mistyped address
is refused immediately instead of failing unexplained on the first run. A
private endpoint is an explicit operator exemption from that check — which is
the whole point of pointing at
localhost; - optional extra headers — a routing tag, an org id. Keep them non-secret: the platform refuses the names it owns itself, but anything else you put there is stored as plain text in the platform configuration, so a key belongs in the key slot instead;
- an optional model list and optional pricing.
Discovery uses the platform-global key, because the model list is fetched once per endpoint and shared by everyone. Member- and project-scope keys are used for streaming only. If you cannot set a global key, list the models explicitly on the entry — a declared list skips discovery entirely.
Pricing decides what a $0 usage row means. Give the endpoint a price (or
a per-model one) and its rows bill as metered like any cloud model. Leave it
empty and an unauthenticated endpoint on a private address is local — your
own hardware, where $0 is the truth — while a keyed or public endpoint is
unpriced: the usage is recorded and flagged rather than passed off as free.
Spend limits are expressed in dollars, so they cannot bound an endpoint you
have not priced.
The declarative model catalog
Every model is a ModelInfo entry. The stream adapters read these fields
instead of pattern-matching model ids:
thinkingStyle— how reasoning is requested:none,adaptive,budget(token budget),effort-param(Responses-style effort),level,separate-id(paired reasoning/non-reasoning models), oralways-on.supportedEfforts/defaultEffort— the reasoning-effort values the model accepts (fromminimalup toxhighandmax). The UI only offers what the entry lists.temperaturePolicy— allowed range, fixed value, forbidden-with-thinking, or rejected-by-API; the adapter never sends what the policy forbids.contextWindow/extendedContextWindow— standard window plus an opt-in larger window (with the beta header when one is required). A local model whose endpoint does not report its context length leaves the field unset rather than inventing a number.pricing— real provider rates in $/1M tokens: uncached input, output, cache-read, cache-write, optional long-context tiers where a large request re-prices the whole call, and — where the provider prices audio input at a premium (Gemini's flash family) — separate audio-input and cached-audio rates. A model whose provider publishes no cache discount carries its input rate as its cache-read rate, so a cached token on a priced, selectable model is never billed at $0.billing— how the model is paid for, and therefore what a$0usage row means:metered(the catalog carries the provider's published list price),subscription(a plan pays for it elsewhere, so usage records no per-token cost),local(your own hardware — there is no vendor rate to record) orunpriced(a catalog entry whose provider publishes no per-token price, or a keyed or public custom endpoint the operator has not priced — an unpriced keyless endpoint on a private address islocal). Every entry declares one, and a0 / 0price is nevermetered. Usage that cannot be priced is logged rather than passed off as free.tier— one offlagship | value | mini | nano | code, used for grouping and sorting in the model selector, plus per-model value/capability comparison scores (1–10) and the list price for the selector's quick-compare view.capability—chat(default) orembedding. Embedding models are filtered out of the chat model selector and surface only where embeddings are configured.capabilities— per-modality media-input descriptor (image / audio / video / pdf), with verified per-model limits where the provider documents them. A modality is listed only when the platform can actually deliver it to that model's API — and a single delivery boundary in the stream runtime decides, per request, whether each image / audio / video / PDF block is passed natively, adapted, or replaced with an honest text marker. Tools and connectors always produce full media; models that cannot ingest a modality receive a clear marker instead of silently losing content.fastVariantApiId/fastVariantRequest/fastVariantPricing/pairedVariantId— fast-mode switches. A high-speed variant is reached either by swapping the wire model id (fastVariantApiId) or by keeping the id and adding a request parameter plus a beta header (fastVariantRequest); exactly one applies per model, and either way the request is billed fromfastVariantPricingwhen fast mode is on.pairedVariantIdis the paired non-reasoning model id for families that split reasoning into a separate id.
Individual model ids are deliberately not listed here — the catalog tracks provider releases, and the live list is always available from the models route and the chat model selector.
No model catalog is rendered into an agent's context either, which is a deliberate shape rather than an omission: an id an agent has not looked up is a guess, and a guess that misses does not run a weaker model — the run ends without completing a step. Agents read the same route the UI does, through the read-only inspection skill's model script, before setting a model on a subagent, an agent's default, or a workflow.
Availability and selection
Model policy can be set at three layers. An explicit conversation selection wins per field, then the agent default, then the provider/catalog default. Agent Config can set the default model, reasoning effort, an exact catalog-declared base or extended context, and fast mode; clearing a field restores inheritance, while an explicit Fast Off overrides an inherited On. Controls are shown from declarative catalog capabilities, and incompatible combinations are rejected rather than silently adjusted.
An agent default is inherited into every conversation, including ones running a different model. When the model that actually runs cannot honour an inherited or stale field, the stream drops that one field and says so in the conversation — a model_policy_skipped note naming the field and whether it came from the agent or the request — instead of failing the turn. Everything else in the request is unaffected, and the value stays stored for the models that do support it. The same resolved policy is what the usage meter prices, so a fast-mode run is always billed at the fast-variant rate.
A model is available when its provider's API key (or OAuth connection) resolves from the credential store for the caller asking. The registry checks availability per provider — one credential lookup for the whole family — and marks every catalog entry accordingly; the models route and the chat UI only offer available models.
Availability is therefore personal, not a property of the deployment. It is resolved through the same most-specific-wins scope cascade described below, so a member who stores their own provider key (see credentials) sees that provider's models offered to them and to nobody else, and the model the stream actually runs on is gated by the very same lookup that fetches the key — the picker cannot offer a model the request would then reject.
Conversation selection is the highest-precedence layer: the chat composer
sets model, reasoning effort, fast mode, and an exact catalog-declared base or extended context window for the next
request. Any absent field falls back independently to the agent's configured
default policy, then to the platform/provider default — the Default Model platform-configuration setting, editable by an
administrator without a restart. The DEFAULT_MODEL environment variable is its
boot fallback, used only where platform configuration is not yet loaded.
Credential scoping
Provider keys are encrypted credentials, never environment variables — see
credentials. Every key lookup carries the
caller's scope (userId, projectId, agentId) and resolves
most-specific-wins: agent → project → user → global. Two projects on
one deployment can use different OpenAI keys, and a single agent can be
pinned to its own key without affecting anyone else.
Resolved keys sit in a short-lived in-process cache (five minutes), keyed by credential and scope. Credential writes, rotations, and OAuth reconnects invalidate exactly the affected scope entries through a credential change bus, so a rotated key takes effect immediately without flushing other tenants' entries. The cache and the OAuth refresh machinery (rotation persistence, automatic purge of permanently-broken credentials) are anchored as process-wide singletons, so every server module instance shares one coherent state — a mid-stream credential purge banner also names the exact project whose login expired, and reconnecting targets that scope directly.
Usage metering
Each provider's parser normalizes usage into the same five billable
categories: uncached input, cache reads, cache writes (5-minute and 1-hour
TTL where the provider distinguishes them), and billable output.
Reasoning-token counts are display-only and never billed, and
provider-reported totals are recomputed rather than trusted. The per-stream
usage meter prices each step against the catalog entry's pricing block,
including long-context tiers. Where a provider bills audio input at a
premium, the parser additionally partitions the input into text and audio
buckets (a partition of the same totals, never an addition) so each is
priced at its own rate.
Codex subscription usage limits
When you sign in to Codex with a ChatGPT subscription, the provider reports your subscription usage limits as one or more windows, each a percentage used. The windows are dynamic — their real length comes from the data (a 15-minute window, a 5-hour window, and a 7-day window are all possible), and every label is derived from that length rather than a fixed assumption.
Neuralis reads this state three ways: from the live stream, and — when you open the chat config Usage Signal panel — with a single on-demand refresh from Codex. That refresh persists the latest reading (keyed to the workspace the credential resolves for — project, then user, then the instance default — so members sharing a project's Codex account see one shared figure), and if the on-demand read is unavailable it falls back to the last stored reading, marked cached. There is no background polling; while a stream is running, its usage updates the open panel directly. Each window's label sits inside a small ring that fills as the window nears its refresh — hover it for the exact time until reset. Crossing 80%, 90%, 95% or 99% during a stream raises a brief, auto-dismissing notice above the composer so you are not surprised by a limit mid-task.
Codex prompt cache
Codex models keep their prompt cache across the whole agentic turn and across
turns of the same conversation: the adapter replays the model's own reasoning
and tool-call items exactly as the backend produced them, and runs each
conversation on one persistent connection with server-side continuation, the
same wire the official Codex CLI uses. A background delegate gets its own
connection; a lost connection is reopened on the next step. Three platform
settings bound it — codexWebsocketEnabled, codexWebsocketMaxSockets and
codexWebsocketFrameTimeoutMs (the stall bound, not the model's thinking
time) — and the stream timing table names the transport of every step.
Codex subscription savings
Codex models are billed through your ChatGPT subscription, not per token, so
their usage records at $0 — and, for the same reason, a Codex login is
outside the USD spend caps and the credential-use counter, wherever it is
stored. A Codex account configured at the instance-wide (global) scope is
therefore available to every project unmetered; setting one requires the
cross-scope permission, which no role holds by default. To show the value the subscription is producing,
Neuralis computes the savings — what the same tokens would have cost at
OpenAI's API pricing for the equivalent model — and shows it (in an accent
colour) wherever a Codex model would otherwise read $0. The admin dashboard
adds a per-subscription view: total savings and a day-precise "has it paid off"
bar that compares the savings against the plan's estimated monthly fee,
pro-rated onto whatever time range you select. Plan fees are estimates from the
reported plan tier; the savings figure itself is derived from the real token
counts and the catalog's list API pricing, never promotional rates.
The savings count Codex tokens only. When a Codex-model agent delegates work to a subagent running a paid model, that subagent's tokens are priced on its own model — a real, billed cost recorded under that model, and one that counts against your spend limits like any other. It is not folded into the subscription's savings, and a spend cap set on a project using Codex will see it. This applies to usage recorded from this release onward: earlier records keep the attribution they were written with, so a long-running installation's historical savings figure still includes some delegated work.
When a turn fails or is cut short
Provider failures are surfaced uniformly across every model rather than ending a turn silently. Codex connection errors retain their safe provider detail, and a connection that closes before a terminal response is treated as a recoverable failure rather than a completed turn. If a request exceeds the model's context window, the chat shows a Context window full banner with Summarize now and Continue — summarizing the earlier turns frees the window, while Continue remains available. Errors and turns cut short for another reason (output-length cap, a content filter) also show a Continue button. The context-window indicator next to the model picker updates live while a turn streams, so you can see the window fill in real time — and as it approaches the summarize threshold, capable models are nudged to compact the conversation on their own before it overflows.
Stopping a turn stops the delivery, not the generation. Pressing Stop — or a subagent hitting its time budget — closes Neuralis's side of the provider connection, so the turn ends promptly and nothing further is written. It does not send a cancellation to the provider, so the model may finish generating on their side and the request may still be billed by them. Treat Stop as "this run is over", not as "that request never happened".
And a stopped step is recorded, not written off. The provider had already read the request by the time you pressed Stop, so the step books what the provider reported it used — and where a provider reports nothing until the answer finishes, it books the size of the last measured request as the input instead. Nothing is guessed from the text on screen: on those providers the stopped step's output is simply not billed, and the usage record says so rather than presenting the gap as free. The same holds for a browser disconnect, a server restart taken mid-answer, and a provider error partway through a step. A turn cut short by a server restart says so — it reads as stopped by the restart rather than by you, and it can be continued.
A provider that asks you to wait is waited on. When a provider says "retry in N seconds", Neuralis honours that exact figure instead of guessing, and each step gets its own small retry allowance so one blip early in a long piece of work does not spend the budget for the rest of it. Every one of those waits is interruptible: Stop ends the turn immediately rather than after the wait.
Titles, rewrites and summaries run on your conversation's own model. The three one-shot helper calls — the conversation title, the composer's prompt polish, and the summary — go through the same provider connections a turn does, so anything you can chat with can also title and summarize: a ChatGPT subscription, a Responses-only model, a model running on your own hardware. By default each one inherits whatever model the conversation is set to; the context-window popover can pin the summary to a specific model instead, and the polish control has its own picker. Titles and rewrites deliberately run at the model's lowest reasoning setting — they are short answers on a short clock — and the summary keeps the conversation's full setup, because it is the memory the next turn works from.
A conversation is named the moment you send. The name comes from your own first message immediately, with no model call in the way, and the model's shorter title replaces it a few seconds later. Stopping the turn before it answers, or a turn that is refused, keeps the first name — a conversation never stays "New conversation". If you rename it by hand in the meantime, your name stays.
Summarizing never loses history. A summary is a non-destructive projection: the full transcript is always kept on disk, shown whole in the timeline, and readable by the agent. While a summary runs, the composer shows a "Summarizing older turns…" banner with a Stop button, and the input is briefly locked until it finishes. When a summary is active the model reads it in place of the older turns; a Show full control switches the model back to the complete history at any time (and Resume summary switches back), without ever truncating the conversation.
A summary is checked against the record, not only remembered. Before the model writes anything, the platform reads the conversation for what it verifiably did — files changed with their line counts, every command run (commits marked), background runs and their logs, subagent runs and their transcripts, messages received, tool failures — and hands that to it as established fact. The same record is available on its own, so an agent can check what happened rather than recall it. You can also give a conversation a standing summary focus in the context-window popover: a sentence about what its summaries should always keep, which every later summary follows, including the automatic ones that run with nobody watching.