A bounded fan-out must not recurse: MergeBounded, never a bare Merge(n)

Rx's Merge(maxConcurrent) is the platform's bounded fan-out: an export's per-node lookups, a recursive delete's pre-validation, a static-repo import's upserts, a GitHub sync's blob reads. It keeps the inners beyond the bound in a queue — and when an inner completes, it subscribes the next queued inner inline, inside that inner's OnCompleted, on the same stack.

That is harmless while inners complete later, on another turn. It is unbounded recursion the moment queued inners complete synchronously during Subscribe: each completion subscribes the next, whose completion subscribes the next, so the stack depth is proportional to the length of the queue. Measured in BoundedMergeStackDepthTest: about 58 frames per queued leg for the export's lookup shape — 1,363 frames behind a queue of 20, 8,903 behind a queue of 150. A thread-pool thread runs out of stack at a few thousand queued legs, and a StackOverflowException cannot be caught: the process dies.

Inners complete synchronously more often than the happy path suggests

The production crash (#5649)

memex-cloud, pod …-v4txp, 06:56:01Z on 2026-09-24, during a roll (the sibling pod …-5hflx died the same way at 06:49:48Z). createdump reported Unwind: exception type System.StackOverflowException. The runtime's stderr trace, read with a Logs action filtered to stream="stderr", shows one MessageHub.HandleCallbacks at the bottom (under MessageService.DrainLoop) and above it a block that repeats until the stack ran out:

Merge.ObservablesMaxConcurrency.Inner  →  Select (ValueTuple)  →  CombineLatest
  →  Catch  →  Select  →  Timeout  →  Take  →  Select  →  Defer / Finally / Do
  →  MessageHub.WrapWithCancelOnDispose  →  AsyncSubject (already terminated)
  →  Timeout  →  Catch  →  ReturnImmediate  →  CombineLatest.SecondObserver  →  Merge.Inner  → …

That is, frame for frame, MeshOperations.GetNodeCollectionConfigs — two ReadHub.Observe(new GetDataRequest(…)) legs, each .Take(1).Timeout(…).Select(…).Catch(→ Return), combined by CombineLatest and projected to (NodePath, Configs) — running under Merge(NodeCopyHelper.DefaultBatchSize) for an MCP export of Ops/Logs (~2,200 nodes) — the MCP call dropped seconds before the replica died. The first batch of lookups was answered normally (the HandleCallbacks); by then the replica was shutting down, every queued lookup's response subject was already faulted, and the dequeue chain recursed through the rest of the queue.

The …-5hflx trace (06:49:47Z) carries the same repeating MessageHub.WrapWithCancelOnDispose frame on stderr; on …-v4txp a read of the Merge.ObservablesMaxConcurrency frames alone was cut at its limit of 50, so the recursion was at least that many levels deep. What this does NOT establish: the exact depth, or that the export was the only bounded fan-out recursing at the time. The MergeBounded fix covers every bounded fan-out in src/ and memex/, so it does not depend on which one it was.

The fix: MergeBounded

MeshWeaver.Messaging.BoundedMergeExtensions.MergeBounded(n) is Merge(n) with the recursion replaced by a drain loop — the work-in-progress loop Rx's own Concat uses. Starting the next inner is one loop per subscription, entered by whichever signal made a slot or an inner available (an outer OnNext, an inner's completion, the outer's completion). A signal that arrives while that loop is already running — an inner completing synchronously inside the subscription the loop is making is exactly that case — only records that there is more to do and returns; the running loop takes it on its next iteration. The next inner is therefore always subscribed from the loop's frame, never from inside the previous inner's OnCompleted, so the depth no longer depends on the queue.

Everything else is Merge's: at most n inners are subscribed at once, inners start in arrival order, an inner that fits the bound is subscribed inline on the signal that delivered it (no scheduler, no thread hop — so an outer that emits an inner and then faults in the same turn still has that inner's synchronous values delivered before the fault), values are forwarded serialised, the first error terminates everything, and the sequence completes when the outer and every inner have.

A first version subscribed each inner through Scheduler.CurrentThread instead. Review caught what that costs: inside a running trampoline every inner subscription is DEFERRED, so an outer that emits an inner and faults in the same turn drops that inner's values, and a long-running outer on the trampoline starves the inners. AnInnerEmittedBeforeAnOuterFault_DeliversItsValuesFirst pins it.

// ❌ recurses once per queued inner when inners complete synchronously
nodes.Select(n => Lookup(n)).ToObservable().Merge(NodeCopyHelper.DefaultBatchSize)

// ✅ same bound, flat stack
nodes.Select(n => Lookup(n)).MergeBounded(NodeCopyHelper.DefaultBatchSize)

Rx's other fan-in operators are not affected: both Concat overloads already drain through a loop or a trampoline, and an unbounded Merge() / SelectMany keeps no queue.

Enforcement

Satellite repositories (MeshWeaver.Plugins and the node repos) carry their own Merge(n) sites and are not covered by the guard; they can adopt MergeBounded once a sealed platform set carries it.