A silent cell owner must not mislabel a round

Every cell of a thread ({thread}/{cellId}, a ThreadMessage) is its own per-node hub. On the control instance those owners were regularly slow to answer: a version-1 user cell nobody had read since it was written, the cells of a thread whose hub the supervisor had just recycled, every cell of a thread on a pod fresh out of a roll, and a response cell whose first writes were never acknowledged inside the write queue's 5 s bound ([UpdateQueue] ADVANCE_WITHOUT_HANDOFF → [MergeGuard] refused stale/reordered cross-hub write → [UpdateRemote] LATE_NACK_REENQUEUE). Why an owner is that slow is a platform question (the query fan-in stall family, core #5315 / #5344 / #5478) and is NOT answered here. What this page records is where the AI engine let one silent owner decide a whole round: two of them are fixed (sections 2 and 3), and the history load is recorded as still open (section 1).

1. History: an owner read is the only authoritative one — still open (MeshWeaver#5862)

ThreadExecution.LoadFullConversationHistoryFromMesh reads every prior cell from its OWNER, in parallel, each under a 5 s budget, and an Unseen cell fails the round (HISTORY_LOAD_FAILED, #982 — a history with an unknown hole is never handed to the agent). Measured on the control instance: 1 of 272 prior cells … did not answer within 5s on the fleet coordinator (the unseen cell had been created on the same pod 93 s earlier), 27 of 31 … on a pod fresh out of a roll, 1 of 1 … on a relaunch 30 s after a recycle.

Rejected remedy, recorded so it is not tried again: reading settled cells from ONE children listing instead of from their owners. A listing's Content comes from the read-side index, which lags every committed write (Doc/Architecture/CqrsAndContentAccess — "Query .Content is always stale — never read it"), and thread cells are NOT write-once: ResubmitMessage rewrites a user cell's text, and a daily-limit re-queue runs the next round INTO its Error cell. A stale row would put an old prompt or an old refusal into the model's history with nothing reporting it — #2226's silent wrong answer. The owner read stays the source.

What the round does today when an owner is silent: it ends in a NAMED state — the response cell carries the localized "history could not be loaded" error, the thread settles Idle with that Summary, and a delegating parent is told the round failed. During the thread hub's own teardown it is not failed at all (HISTORY_LOAD_INTERRUPTED, #2618) and the next activation resumes it. The open question is the owner latency itself (below).

2. Error rounds: the thread settles on its own write (MeshWeaver#6137)

Before. The Error branch of ThreadExecution chained the thread's terminal write onto the error CELL's write under a bare .Timeout(10 s). When the cell's owner did not acknowledge in time, Rx's message-less TimeoutException: The operation has timed out. faulted the round, and the outer catch handled it as an initialization stall:

Now. As in the Completed and Cancelled branches, the cell push and the thread's terminal write are independent writes to two nodes. The cell push reports its own verdict (the write pipeline bounds it; a failure is logged); the thread settles on its own write — Idle, the classified Summary, the inbox fold, the re-queue — whatever the cell's owner is doing. The consumed-inbox snapshot is now taken synchronously in the branch, before the round's finally resets the channel; it used to be taken in the cell push's callback, after that reset.

Pinned by RefusedRoundRequeueTest.ARefusedRound_SettlesTheThread_WhileTheErrorCellsOwnerIsSilent: the response cell's owner is silenced from the moment the provider is called; the thread must settle with the input pending again and a Summary naming the refusal, and the cell, once released, carries no "stalled" suffix. Negative control: before the change the thread settles Idle with nothing pending and an empty Summary — the production shape exactly.

3. A post to a node that is not a thread is refused by name (MeshWeaver#6003)

Before. ThreadInput.AppendUserInput's update lambda read the target with a logging ContentAs<Thread> and returned the node unchanged when it was something else. A post into the TRIAGE ITEM Hosting/Triage/issue/systemorph-meshweaver-6002 (whose threads are paths recorded in its content) therefore logged As<Thread> … value is TriageItemContent … not convertible at Error and vanished; the poster was told nothing — submit_message read an onError that a cross-hub write can only call later, so it answered "Message queued" for every post.

Now. ThreadInput.AppendUserInputObservable is the append as a cold observable that FAULTS, naming the node and saying to post to the thread path the node records, when the target declares another NodeType or holds content that is not a thread. hub.ObserveSubmitMessage(...) is SubmitMessage as an observable (SubmitMessage subscribes it and reports the reason through its onError), and the submit_message tool answers from the write's verdict through ToolTask.Bridge, so the agent hears the refusal and can follow the item's thread path. A signed-in user's post already went through ThreadContinuation, which refused a non-thread target.

Pinned by SubmitToANonThreadTest. Negative control: on the old lambda the test reproduces the production Error line verbatim and the refusal never arrives.

What is NOT established