Runner Pools and Dispatch Queues

Policy ci-two-pools-priority-queues. This page is the platform's statement of the rule. The mechanism lives in MeshWeaver.Plugins (Hosting/CiDispatchQueue.md, get Hosting/CiDispatchQueue on the control instance), and the runner configuration lives in Systemorph/Memex (deployments/aks/ci-runners/README.md).

The rule

  1. Two pools, split by technical need only. Plain (vars.MW_RUNNER → aks-silos) and Docker (vars.MW_RUNNER_DOCKER → aks-silos-dind). They share ALL capacity. No pool, label or PriorityClass is reserved for a kind of work.

  2. Priority comes from the queue. Every CI run takes a place in the control instance's ci queue (Hosting/Queues/operator/ci) in one tier:

    Tier For
    express blocking work: a run on main while main is red; a settle / stamp / publish change on main (or its bot PR); an explicit unblock or hotfix
    trunk every other main / schedule / dispatch run
    gate core's dependent suites (MeshWeaver.Plugins core-candidate.yml)
    pr pull requests

    The queue dispatches: a run does not start its heavy legs until the queue starts it, in tier order. A measured share of the pool (20 % by default, relative, never a count) is kept free so express never waits behind a lower tier.

  3. The order is mesh state. It lives on the queue node and its job nodes. A platform admin lists and moves it through MCP (queue_list, queue_reorder), and the next dispatch follows.

  4. CI never depends on the control instance to run. An enqueue that fails, a trunk run not admitted within 15 minutes, and a held pull request the dispatcher never re-runs all start anyway, loudly. The last of these is handled by a GitHub-side fallback lane that watches the dispatcher's activity, never the control instance.

Why reserved lanes were retired

Until this policy the fleet had five scale sets. Three of them, trunk (8), Docker trunk (12) and gate (12), bought precedence by reserving runners under their own label and PriorityClass. GitHub hands a label's queued jobs out first come, first served, and has no job priority. A reservation is always either too big, idling while other work queues, or too small, so its own work queues. Measured over the 40 most recent MeshWeaver.Plugins Plugin Catalog CI runs (2026-10-04 20:21Z → 10-05 06:57Z), jobs on the reserved trunk sets waited p90 6.6 / 3.8 min. Pull-request jobs on the shared sets waited p90 0.6 / 1.0 min. Ordering one shared pool fails in neither direction.

Where the queue lives

GitHub cannot express the order. A concurrency group holds one pending run and replaces it with the newest, and nothing in GitHub can be re-ordered by an admin. The queue is therefore an ordinary Hosting/Queue execution queue on the control instance (The activity execution model, MeshWeaver.Plugins Hosting/ActivityExecutionModel.md). Its own hub is the only writer of the order, and the tiers are priority bands of that ONE queue. One hub ordering one capacity has no cross-queue arbitration that could deadlock. Its dependency on the control instance is cut by rule 4.

How a run waits, and how it is dispatched

Agent work: the same queue model, no limit

The same queue model, tiers and MCP re-ordering apply to every control-plane agent thread:

Each one is a job on Hosting/Queues/operator/agents, which the queue's hub dispatches in tier order. Two rules differ from the CI queue:

Fail-safe as for CI: if the queue never stamps a job, the thread is started directly, with a warning. The mechanism, the bounds this removed and the tests are in MeshWeaver.Plugins Hosting/CiDispatchQueue.md §7a–7c.

What the gates refuse

Migration

The order leaves no unguarded moment:

  1. Memex raises the two pools to absorb the retiring sets. The namespace quotas are unchanged, and the old sets stay installed until nothing names them.
  2. MeshWeaver.Plugins stops naming the lanes and ships the queue. The queue stays bypassed (MW_BUILD_QUEUE=off, loud) until the control instance runs it.
  3. Setting MW_BUILD_QUEUE=dispatch turns it on. This policy is in force from that moment, and rollback is off.
  4. Delete the trunk variables. A follow-up Memex change then removes the retired sets, applied through the governed Hosting path.

The Staged Pull Request Pipeline decides WHEN a pull request may start its heavy legs. This policy decides in what ORDER the started work gets runners.