A round refused by the daily limit is queued again, not spent

The defect (Plugins #3076)

On the control instance, OpenRouter answered rounds with HTTP 403 Key limit exceeded (daily limit) after 0 output tokens. The dispatch pool did the right thing for every round that came after the refusal: it closed the key until the next UTC midnight, and those rounds waited Queued, with Thread.QueuedUntil telling every replica's supervisor that the wait has an end (#3092).

The round that met the refusal got none of that. ThreadExecution wrote its response cell as Error, flipped the thread to Idle with nothing pending, and nothing ever came back to it. Its input was spent on a refusal. For a person's chat that is a visible failed message they can resend. For the control plane's own threads it is a silent stop:

#3060 named the remaining gap: "Key fail-over for the round that hits an exhausted key is unchanged… The round that met the 403 ends Error."

The fix: the round is not spent

RefusedRoundRequeue (src/MeshWeaver.AI/RefusedRoundRequeue.cs) is applied in the Error branch of ThreadExecution, right after the round's failure is reported to the dispatch pool. Because it sits in the engine, it covers every thread kind at once.

When a round is re-queued

RefusedRoundRequeue.Decide re-queues only when all of these hold:

Condition Why
The failure closed the pool (ThreadDispatchPool.FailureOf → ClosedUntil), and it is a key's daily limit (IsDailyLimit → ProviderKeyFailureKind.KeyLimitExceeded). This is the one refusal that lifts by itself at a known time. A credit refusal or a rejected key needs a person, so it still ends Error. A 429 closes nothing.
The round produced nothing: no streamed text, no tool call, no node change. Nothing it did can be repeated twice.
The round carries the pending messages it drained (RoundParams.Inputs). A resume carries none.
The pool's answer for the next round (ThreadDispatchPool.NextAdmission, asked without recording anything) is a wait, or a move to an alternative model. If the pool would admit the repeat at once on the same model, re-queuing would run the refusal again at hub speed. No pool means the same thing, so neither case is re-queued. A delegated sub-thread falls here too, because it is admitted on its parent's seat.

What the re-queue writes

RefusedRoundRequeue.Apply makes one terminal write to the thread:

The round that is re-queued pushes no "stopped" bell notification, because it has not stopped. REFUSED_ROUND_REQUEUED in the log names the thread, the cell, the end of the wait and the pool's reason.

One cell per round, however many probes

The response id is derived from the drained messages' timestamp and text (ThreadSubmission.DeriveDeterministicResponseId). Both are kept, so the re-dispatch derives the same id and finds that cell already terminal (Error). Normally a terminal cell means the round is over, and the input is settled as answered (the supervisor-settle race). Here the dispatch sees the RequeuedFromResponseId marker (RefusedRoundRequeue.RunsInto) and runs into that cell instead.

The pool lets one probe through a closed key every probe interval. However many probes it takes, the thread therefore carries one response cell for the round:

No Error cell is added per attempt, and the model's history never holds a refusal as an assistant turn, because the round's own cell is excluded from its history.

A follow-up submitted while the round waits joins it. The drained set changes, so the round derives a new cell. The commit then retires the refusal's cell from Messages (RefusedRoundRequeue.SupersededCells), because the new cell answers that input too. Left in Messages, the old cell would reach the model's history as an assistant turn reading *Error: …*.

The input handed back is exactly what the round's commit drained (committedInputs in ThreadSubmission). It is never the pre-commit plan: an id that another drain consumed in between is not this round's to replay.

What it does not do

Tests

src/MeshWeaver.AI.Test/RefusedRoundRequeueTest.cs: