Compiling Off the ThreadPool
A portal silo has ONE shared .NET ThreadPool. Orleans grain turns, the routing pool's subscribe legs
(IIoPool.SubscribeThroughPool, see Reading a Routing Saturation Report)
and every reactive continuation that delivers a cross-hub response all run on it. Its minimum worker
count is Environment.ProcessorCount — 6 on an AKS portal pod (CPU limit 6) — and past that it
grows by hill-climbing, roughly one or two threads a second. Anything that holds pool workers for
hundreds of milliseconds therefore stops the silo delivering messages while nothing is stuck.
The two ways a compile held the pool
1. The CPU-bound half of the compile ran on a pool worker. MeshNodeCompilationService starts a
compile through OnThreadPool(() => CompileAsync(...)). CompileAsync returns a Task, so overload
resolution binds the Func<Task<T>> overload — Task.Run — and not the Func<T> overload that uses
CompileThread's dedicated thread. CompileAsyncCore then awaits the debug source write
(EnableSourceDebugging is on by default) and, for #r "nuget:" sources, the NuGet restore; after
the first await the rest — parse, source generators, bind, emit — resumed on whatever pool worker the
await completed on. So every production emit ran on the pool, although CompileThread's own comment
said the compile leaf did not.
2. Roslyn fanned each emit back out onto the pool. CSharpCompilationOptions.ConcurrentBuild
defaults to true. With it, Roslyn compiles declarations and method bodies through Parallel.For and
Task.Run on TaskScheduler.Default, with no degree-of-parallelism cap — from whatever thread
started the emit. A dedicated thread therefore kept only the CALLER off the pool, never the work.
The measurement
A standalone probe (Environment.ProcessorCount = 6): every 5 ms queue one ThreadPool work item and
record how long it waits to START, while N concurrent emits of a 20-class × 20-method source run.
| N | launch | ConcurrentBuild | probe p50 | probe p99 | pool queue max |
|---|---|---|---|---|---|
| 0 | — | — | 0.01 ms | 0.04 ms | 1 |
| 1 | Task.Run |
on (was prod) | 7.7 ms | 41 ms | 19 |
| 1 | dedicated | off (now) | 0.01 ms | 0.40 ms | 1 |
| 6 | Task.Run |
on (was prod) | 46 ms | 417 ms | 115 |
| 6 | Task.Run |
off | 326 ms | 687 ms | 114 |
| 6 | dedicated | on | 33 ms | 354 ms | 114 |
| 6 | dedicated | off (now) | 0.01 ms | 0.18 ms | 1 |
| 12 | Task.Run |
on (was prod) | 124 ms | 812 ms | 147 |
| 12 | dedicated | off (now) | 0.02 ms | 0.68 ms | 1 |
Both halves are needed: switching ConcurrentBuild off alone pins one pool worker per compile (row
"6 / Task.Run / off" is the WORST row), and a dedicated thread alone still fans out. Wall time for the
batch is unchanged at N ≥ 6 (729 ms vs 799 ms at N = 6; 1404 vs 1216 ms at N = 12); a lone compile
takes longer on one core (about 150 → 500 ms for this source), which is the price of not borrowing the
silo's grain-turn threads to finish it.
The fix
EmitPipeline.CreateRunCompilationOptions()= the canonicalCreateCompilationOptions()withConcurrentBuildoff.CreateEmitCompilationand the compilation inputs the diagnostics and language-service compilations are built from use it.- The canonical factory is unchanged, and the content key (
EmitPipeline.OptionsFingerprint) still renders it —ConcurrentBuildis a scheduling choice that does not change the emitted bytes, and changing the fingerprinted factory would re-key every NodeType in the fleet. CompileAsyncCorekeeps its IO prefix and hands the synchronous half (EmitCompiled) to the CPU lane (below). The first cut of this fix handed it toCompileThread.Run— off the pool, but one new thread per distinct NodeType compiling at once, which a roll turns into hundreds.
The CPU lane (IoPoolNames.CompileCpu)
Every purely CPU-bound Roslyn leaf goes through ONE bounded lane: an IIoPool whose
InvokeBlocking leaves run on at most IoPoolOptions.CompileCpu (default: processor count) threads
the lane starts itself — named mw-cpu-lane, never ThreadPool workers. The bound is the pool's
ordinary blocking scheduler (LimitedConcurrencyLevelTaskScheduler) with dedicatedThreads: true:
the same queue and the same cap, only the drain loops run on their own threads. A burst of hundreds of
distinct-type compiles now QUEUES behind the cap instead of becoming hundreds of threads, and the OS
time-slices the lane's threads fairly against the ThreadPool the grain turns run on.
Its leaves:
- a NodeType's
EmitCompiled(generators, bind, emit, artifact write) and the failure-pathCompileDiagnostics.DiagnoseInputs; - every language-service diagnostics bind — the node workspace (
GetDiagnostics), a speculative proposal (SpeculativeCompilation.Diagnose) and a script cell. The IO half of each (NuGet refs,Project.GetCompilationAsync) stays on theCompilepool; only the bind moves. The script workspace now also binds withConcurrentBuildoff.
🚨 The lane is separate from the Compile pool on purpose. Compile also runs kernel SCRIPTS — user
code that may activate a NodeType and wait for its compile — and cell-surface assembly loads. An emit
queued behind the same cap is the nested-gate deadlock the compile leaf already hit once (every slot
held by a script waiting for a compile that needs a slot). A lane leaf does nothing but compute: no
pool call, no mesh read, no wait on another leaf, so it cannot nest by construction. Keep it that way —
a leaf that needs IO splits it off, as the three diagnostics arms do.
Tests, each with a negative control that fails on the unfixed code:
CompileCpuLaneIsBoundedAndOffThePoolTest(Hosting.Test) — six parked leaves on a cap-2 lane: at most 2 run at once, all 6 run, none on a pool worker; the control, the ordinaryCompilepool with the same cap, runs all 6 on pool workers. With the registry building the lane without dedicated threads the first test fails.ServiceCompileStaysOffTheThreadPoolTest— a real compile through the service reaches Roslyn on anmw-cpu-lanethread withConcurrentBuildoff. Fails when the emit is routed through theCompilepool.LanguageServiceDiagnosticsRunOnTheCpuLaneTest— all three diagnostics arms bind onmw-cpu-lanethreads. Fails when the binds are routed through theCompilepool.RoslynEmitStaysOffTheThreadPoolTestpins the emit itself deterministically: anAsyncLocalchange handler counts every ThreadPool worker that starts running under the emit'sExecutionContext(Roslyn's fan-out tasks flow it). With the run options the count is 0; the negative control — the same emit withConcurrentBuildon — counts more than 0. A third case pins that the content key did not move.
What the production evidence does and does not say
The routing back-pressure flood of 2026-09-22 (core issues #5226–#5307) was read against the pods named
in each issue's evidence table. Of the 65 distinct crossings quoted in those bodies that printed
waiting for a pool slot ≥ 45 with subscribing ≤ 1 — the thread-shortage reading — 63 were
logged while two or more ReplicaSets of the same deployment were live within ±10 minutes, i.e. during
a roll (memex-cloud went through five ReplicaSets between 10:50Z and 13:05Z; memex rolled at 17:33Z).
The two exceptions are both memex at 18:20Z, where no second ReplicaSet appears in the sample. The
starved pod was as often a SURVIVOR of the old ReplicaSet (5c89c7f5f9-sddb5, 12:34–13:02Z, while
7cc85f47c rolled in) as a freshly booted one (7bc86fb6f-qnznf, 11:32–11:43Z), so this is "during a
roll", not "on the new pod". The [ReleasePostCondition] compile lines on …-sddb5 at 12:38Z fall
inside its starvation window. A roll is exactly when compiles happen on serving silos: a framework
identity change makes NodeTypes framework-stale, and release requests and first activations compile.
Read the 63/65 against its base rate: the sample is the log lines the incident filer quoted (at most five per issue), and on memex-cloud a roll was in progress for most of the sampled day, so "during a roll" is the expected state of most samples, not a rare one.
What was NOT established: the portal's own ThreadPool metrics for those windows (the control instance
was unavailable, so no Logs/Prometheus read was taken), and therefore how much of each window's
starvation the compiles account for. Other CPU on a rolling survivor — re-activating the grains of the
pods being replaced — competes for the same cores and is not addressed here. Still on the ThreadPool
by design: the language service's hover and completion requests (Roslyn's async QuickInfoService /
CompletionService, which schedule their own work with ConfigureAwait(false) and cannot be pinned
to a thread), the IO halves of the diagnostics requests, and CompileResultFromAssembly (assembly load
and user type initializers — bounded by its own timeout, on CompileThread, not a CPU compile).
Related
- Reading a Routing Saturation Report — which gauge sees the wait.
- Node Type Compilation — the compile pipeline this changes.
- Controlled IO Pooling — the pools and why the compile leaf is not on one.