Per-provider daily budget and provider health
On 2026-10-01 the platform's OpenRouter key reached OpenRouter's own daily key limit, and every
model call on it failed with HTTP 403 … Key limit exceeded (daily limit): chat, the GLM-5.3 PR
reviewer, and the memex search embeddings. Nobody noticed. The failure showed up only inside
individual thread cells. MeshWeaver had no page that showed a provider's limit, its spend today or
its health.
This page describes the three things that fix that:
- a daily limit per provider in MeshWeaver, enforced before a round calls the provider, from spend that MeshWeaver itself records;
- a health record per provider key, written when the provider refuses the key and when the key works again;
- a section on the Providers page and on each provider page that shows limit, today's spend, remaining amount, health and last success.
Code: src/MeshWeaver.AI/ProviderBudget/ (ProviderBudgetRule, ProviderHealthRule,
ProviderBudgetLedger, ProviderBudgetGuard, ProviderBudgetView) and the call sites in
ThreadExecution.
The limit
ModelProviderConfiguration.DailyLimit, a decimal? on the provider node:
| Value | Meaning |
|---|---|
null / blank |
no limit — nothing is enforced; spend is still recorded and shown |
300 |
$300.00 per UTC day on this key — a guard rail, not a hard cap (see Overshoot) |
- Currency: USD, the platform providers' own pricing currency (OpenRouter publishes USD). The limit is compared with spend priced in USD. A spend row in another currency is not converted and not dropped. It makes the day's total undetermined (see below).
- The day starts at 00:00 UTC. This is OpenRouter's own daily reset, so MeshWeaver's counter and the provider's counter roll over together. The day bucket is a machine fact and is never converted to the viewer's time zone. The page shows the next reset as an instant in the viewer's zone.
- Where to set it: on the provider's page (
/Provider/OpenRouter), field Daily spending limit. That field is the node-bound editor (MeshNodeContentEditorControl) and only a platform admin can writeProvider/{id}. The editor writes a text value, so the property reads leniently (LenientDecimalJsonConverter): a number or a numeric string is the amount, and blank or unparseable text means no limit. A typo can therefore never break the provider node, and with it the key. The page then shows no limit, so the typo is visible. - Recommended setting:
DailyLimit: 300onProvider/OpenRouteron each instance that carries the platform OpenRouter key.
Which key a round spends: the key holder
A limit belongs to the node that holds the key. Provider/OpenRouterEU has no key of its own.
Its CredentialFrom names Provider/OpenRouter, so it is the same account reached through another
endpoint. Every round through either provider spends the one budget on Provider/OpenRouter.
ProviderBudgetRule.KeyHolderOf follows exactly the hop that ChatClientCredentialResolver
follows:
- only between two root-catalog providers (
Provider/{id}); - one hop only;
- the borrower's own key, when one is set, wins (it is then its own holder).
A limit set on a borrower is not read. Only root-catalog providers are budgeted. A user's own BYOK
key ({user}/_Memex/{provider}) is not the platform's key and is never metered here.
What counts as today's spend
One record per round, at Admin/ProviderBudget/{holder}/{yyyy-MM-dd}/{roundId}
(ProviderSpendCharge). Each record is priced once, when the round finishes, with the same rate
the credit meter uses (ModelCreditGuard.ObserveRate → the model node's authored price, else the
built-in table). This is one cost model with three readers. Today's spend is the fold of the day's
records (ProviderBudgetRule.Fold).
Why a third record, next to the two that already exist:
| Existing record | Why it cannot answer "spent on this key today" |
|---|---|
TokenUsage ({thread}/_Usage/{model}) |
a cumulative per-(thread, model) total with no time dimension; a thread that ran yesterday and continues today carries both days in one number |
ModelCreditCharge (Admin/ModelCredit/…) |
per round and timestamped, but written only for rounds billed to a subscriber's plan; a platform-internal round (the PR reviewer, a scheduled automation) spends the same key and appears nowhere in it |
What is counted: every thread round (chat, agents, delegations, the PR reviewer) on the
MeshWeaver harness whose model sits under a root-catalog provider, whatever subscriber or plan the
round belongs to. Errored and cancelled rounds count what they reported, or a character-based
estimate when the provider reported nothing (marked IsEstimated; the estimate errs high, which is
the safe direction for a budget).
What is NOT counted, and is not faked:
- Embeddings (search indexing,
search_chunksqueries). The embedding provider is configured separately (Embedding:Provider/Embedding:ApiKey,MeshWeaver.Hosting.Embeddings), knows nothing ofModelProvidernodes and records no usage at all. If an instance points its embedding config at the same OpenRouter key, that spend is invisible to this budget, and OpenRouter's own limit stays the only bound on it. - Non-thread utility calls that use a bare chat client: description and icon generation, PR
draft text, the indexing summarizer and image describer. They record no
TokenUsagetoday, so there is nothing to price. They are also not gated by the budget (follow-up: a gate at the bare client). - CLI harnesses (Claude Code, Copilot, Codex, …) run their own credential and are out of scope.
- Unpriced models: a round on a model with no price on the node and none in the built-in table records nothing and logs a warning naming the model.
Enforcement
ProviderBudgetGuard.Check runs in ThreadExecution next to the credit gate, at the last point
before the round calls a provider (the model that will answer is known, nothing has been sent).
It is reactive end to end and never faults. The verdict (ProviderBudgetRule.Decide):
| Outcome | When | Round |
|---|---|---|
NoLimit |
the holder has no DailyLimit |
runs |
Within |
spent today < limit | runs |
Reached |
spent today ≥ limit | refused |
Undetermined |
limit set, but the records could not be read within the bound, or a row is in a foreign currency | runs, and a warning is logged |
A refused round ends as an Error cell whose text names the provider, the spend, the limit and the
reset (key chat.providerDailyBudgetReached):
OpenRouter daily budget reached: $300.00 of $300.00 spent today; resets 00:00 UTC. This round did not run — nothing was sent to the provider.
🚨 An undetermined spend does not block. The budget is a guard rail in front of the provider's
own limit. It is not the only bound on a key. A broken counter (an unreadable Admin partition, a
slow read, an operator who prices one model in EUR) must not take every model on the platform key
down, because that is exactly the outage this feature exists to prevent. The gap is logged as
[ProviderBudget] today's spend on … is UNDETERMINED … the round RUNS.
🚨 There is no fallback. A refused round is not silently re-routed to another provider, because that would spend a key nobody chose for it.
Mid-round crossings complete. As with the credit meter, a round that is already streaming has already spent its tokens. The round that crosses the limit finishes and records its spend. The next round is refused.
Overshoot — a guard rail, not a hard cap. The check reads today's recorded spend and reserves nothing, and a round records its spend when it finishes. Rounds that start together (parallel agent delegations, several users at once) can all read "within" and all run, so the day can end above the limit by what those in-flight rounds cost. That is accepted deliberately: reserving an estimate up front would refuse rounds on a guess and still be wrong in both directions. For a hard ceiling keep the provider's own key limit (e.g. OpenRouter's per-key credit limit) as the outer backstop, a little above this one.
Provider health
ProviderHealth at Admin/ProviderBudget/{holder}/_Health, one node per key holder:
- Failing when a round fails with a refusal that says the key is unusable
(
ProviderHealthRule.Classify): HTTP 401 (credential rejected), 402 (insufficient credits), 403 (Key limit exceeded→KeyLimitExceeded; aninsufficient creditsbody →InsufficientCredits; otherwiseForbidden), 429 (rate limited;Key limit exceeded→KeyLimitExceeded), or, with no status, a body naming one of those conditions. A 404, a 5xx or our own fault says nothing about the key and is ignored. Otherwise one wrong model id would paint the key red. - OK when a round on the key completes.
- The stored message is sanitised (
ProviderHealthRule.Sanitize): the provider's own"message"when the body is JSON, otherwise the line naming the condition. Bearer values,sk-/mw--style keys and any 32+ character opaque token are replaced with[redacted], and the message is capped at 240 characters. A secret is never stored or logged. - Written only on a state change (
OnFailure/OnSuccess), never once per round. The triggers are: OK → failing, a refusal of a different kind or status while failing (the streak's start is kept), failing → OK (the last refusal is kept as history), and, while healthy, a last-success stamp older than 10 minutes. The state is the kind and the status, not the message: a refusal body that carries a per-request value (a request id, a retry-after) differs on every round, so while a key keeps failing the same way its newest message is refreshed at most every 10 minutes. A busy key therefore writes at most every few minutes, failing or healthy.
Where it is shown
/Provider/AiProviders: a Daily budget & health section above the catalog tabs, embedded as its own area (AiProvidersBudget), with one row per root-catalog provider that does not borrow another's key. Routes are listed with their holder ("shares its key with OpenRouterEU"). A provider with no key on its node has a row too ("no limit", "no rounds recorded yet"): a vault-seeded key reads as absent on the node, and a key that went missing is a row an operator must see. The alarm line above the grid quotes a provider's refusal as plain text, never markdown.- Each provider page (
/Provider/{id}): the same section, narrowed to that provider's key holder, embedded as its own area (ProviderBudget). The page also carries the Daily spending limit field.
Columns: provider, daily limit (or no limit), today's spend, remaining, health (OK / failing
since … : last refusal), last success. A reached budget or a failing key is marked 🔴 in the cell
and in a line above the grid (a grid cell cannot be coloured). Timestamps are rendered in the
viewer's zone (DisplayTimeExtensions.ToDisplayTime with the zone captured on the render turn).
🚨 Platform admins only. The records live under Admin, and the provider configurations are
read as System, so the section sits behind the same positive-confirmation gate as the AI usage &
cost tab (AiUsageCostSettingsTab.ConfirmedAdmin). A non-admin, a slow evaluator or a faulted one
sees nothing, and no System read is issued.
After deploying
- Set
DailyLimit: 300onProvider/OpenRouteron every instance that carries the platform OpenRouter key (through the provider page; the root catalog is admin-writable). - The budget gate and the health writes run in the thread hubs'
ThreadExecution, which binds the AI module once per activation. Running thread hubs pick up the new engine only when they are recycled or when the pod rolls. The provider pages (Provider/OpenRouter,Provider/OpenRouterEU) andProvider(the Providers catalog area) render from the per-node hub configuration and need a recycle to show the new section.
Follow-ups
- Gate and meter the non-thread utility clients (description/icon generation, PR draft, the indexing summarizer) at the bare chat-client factory.
- Meter embeddings once the embedding provider can name a
ModelProvidernode as its key. - A per-deployment health line in the fleet app (
Hosting/Models). It needs a cross-instance read of each deployment'sAdmin/ProviderBudget, which no cheap read provides today. - Classify 403 Key limit exceeded as its own condition in the thread cell. Today the cell shows
the provider's raw sentence, because
ProviderFailureClassifierdeliberately leaves 403 unclassified.