A supervisor relaunch that was never counted

Measured on a production portal on 2026-10-09 and filed into bug triage. The page is the companion of A supervisor that re-settles the thread it gave up on.

What it looked like

The thread was Hosting/Triage/_Thread/bug-systemorph-meshweaver-plugins-2958/ea882e98/read-only-probe-for-a-bug-thread-do-not-5adfa63d6423764f9d50. It is a triage probe running on z-ai/glm-5.3 in the bug-fix lane.

The step that did nothing

Act's relaunch arm went through four steps:

  1. log relaunching (n/max);
  2. DelegationInFlight, which stands down when the active cell has an open delegate_to_agent call;
  3. Recycle;
  4. write the count.

The active cell held such a call. It pointed at …/4a437305/read-only-log-forensics-probe-do-not-fil-fffdc767a239ff1dc35f, a child that had finished on 2026-10-05 at 23:52Z: it was Idle, its summary was the OpenRouter HTTP 403 Key limit exceeded, and it had used up its own relaunches 2 of 2. The parent's hub delivers a delegated result into the parent, and that hub had gone while the parent waited. So the tool call stayed open with no result. Step 2 read that as "still delegating" and returned a null row. Nothing was recycled, nothing was counted, and step 1 had already logged a relaunch. The heartbeat ticker that the stand-down defers to lives in the parent's hub, which was the hub that had gone, so nothing would ever end the wait.

Not established: what started the GLM rounds seen on that thread (one began at 14:50:30Z, 33 s after a relaunch line). The relaunch arm itself cannot have started them, because it never got past step 2. A read of the parent's stream activates its hub, and the recovered hub resumes a Streaming cell, so a read by an investigating agent or a viewer could have started them. That was not measured.

What changed

The test

ThreadSupervisorMeshTest.AStaleParent_WhoseOpenDelegationAlreadyFinished_IsRelaunched_AndTheRelaunchIsCounted seeds the production shape: a stale Executing parent whose cell has an open delegation to an Idle child. It asserts supervisorRetries == 1 on the stored node. On the old code it fails, because the count stays at 0 for ever, which makes the old code the negative control. Its control, …_WhoseDelegationIsStillRunning_IsLeftToTheHeartbeat, keeps the child running and asserts that the supervisor touches nothing.

Stopping it on the running portal

The supervisor only ever stood down, so the rounds being paid for belonged to the thread itself. The stopgap on 2026-10-09 was the thread's own Stop request: requestedStatus: Cancelled, the same write as the chat view's Stop button. It cascaded to the running child. A governed Logs action on the control instance then read 0 relaunch lines for that thread between 14:53Z and 15:03Z.