A supervisor relaunch that was never counted
Measured on a production portal on 2026-10-09 and filed into bug triage. The page is the companion of A supervisor that re-settles the thread it gave up on.
What it looked like
The thread was
Hosting/Triage/_Thread/bug-systemorph-meshweaver-plugins-2958/ea882e98/read-only-probe-for-a-bug-thread-do-not-5adfa63d6423764f9d50.
It is a triage probe running on z-ai/glm-5.3 in the bug-fix lane.
- Loki held 105 lines of
[ThreadSupervisor] relaunching … (1/2)between 08:52Z and 14:51Z, exactly 240 s apart. They came from every replica, and every one of them said (1/2). - The live node never carried
supervisorRetries,supervisorLastActionAtorsupervisorNote. The cap (maxRetries: 2) therefore could never be reached, and the thread was never settled or filed. - No other
[ThreadSupervisor]line about that thread reached Loki, not even a warning. The follow-up lines that would have explained the stand-down are logged at Information, and Information from this category does not reach Loki: a six-hour query forsweep:,waking,delegationandepisode movedfound 0 lines.
The step that did nothing
Act's relaunch arm went through four steps:
- log
relaunching (n/max); DelegationInFlight, which stands down when the active cell has an opendelegate_to_agentcall;Recycle;- write the count.
The active cell held such a call. It pointed at
…/4a437305/read-only-log-forensics-probe-do-not-fil-fffdc767a239ff1dc35f, a child that had finished on
2026-10-05 at 23:52Z: it was Idle, its summary was the OpenRouter HTTP 403 Key limit exceeded, and it
had used up its own relaunches 2 of 2. The parent's hub delivers a delegated result into the parent, and
that hub had gone while the parent waited. So the tool call stayed open with no result. Step 2 read that
as "still delegating" and returned a null row. Nothing was recycled, nothing was counted, and step 1 had
already logged a relaunch. The heartbeat ticker that the stand-down defers to lives in the parent's hub,
which was the hub that had gone, so nothing would ever end the wait.
Not established: what started the GLM rounds seen on that thread (one began at 14:50:30Z, 33 s after a relaunch line). The relaunch arm itself cannot have started them, because it never got past step 2. A read of the parent's stream activates its hub, and the recovered hub resumes a Streaming cell, so a read by an investigating agent or a viewer could have started them. That was not measured.
What changed
- A delegation is in flight only while its CHILD is.
DelegationInFlightnow reads each open call's delegated thread (ChildStillRunning): if it is executing or has unanswered queued input, the supervisor stands down; if it is finished, unreadable or gone, the parent is relaunched. Its fresh activation is what picks the result up. - Count, confirm, then recycle. The count is written first, and
WriteConfirmedchecks that the node the owner returns carries it. The activation is recycled only after that. If the owner takes no write before the recycle (the wedged hub the recycle exists for), the recycle goes first and the count is written to the fresh activation, where it must be confirmed or the action fails as anActionFailedrow with a warning. A counted relaunch whose recycle does not confirm a fresh activation also fails asActionFailed. The count stays on the node, so the cap is still reached. - The line is logged after the fact.
relaunched {Path} (n/max), the count on the node before|after the recycleis written at Warning once the relaunch has taken effect. A stand-down is never logged as a relaunch.
The test
ThreadSupervisorMeshTest.AStaleParent_WhoseOpenDelegationAlreadyFinished_IsRelaunched_AndTheRelaunchIsCounted
seeds the production shape: a stale Executing parent whose cell has an open delegation to an Idle
child. It asserts supervisorRetries == 1 on the stored node. On the old code it fails, because the
count stays at 0 for ever, which makes the old code the negative control. Its control,
…_WhoseDelegationIsStillRunning_IsLeftToTheHeartbeat, keeps the child running and asserts that the
supervisor touches nothing.
Stopping it on the running portal
The supervisor only ever stood down, so the rounds being paid for belonged to the thread itself. The
stopgap on 2026-10-09 was the thread's own Stop request: requestedStatus: Cancelled, the same write as
the chat view's Stop button. It cascaded to the running child. A governed Logs action on the control
instance then read 0 relaunch lines for that thread between 14:53Z and 15:03Z.