The CD Ledger Records a Failure, Not a Cadence

main-cd.yml's hourly reconciler files an issue when it finds main's HEAD without its deployment image set:

CD: main <sha> has an incomplete image set

main's HEAD <sha> is missing part of its deployment image set in ACR, so every self-updating install stays on the previous image.

It had filed that issue 109 times by 2026-09-18, at up to 15 a day. Every one of them was read, by anybody who looked, as a delivery outage — that is what the words say, and the label is ci-failure.

Not one of the 109 recorded a delivery that had failed.

This page is what the alarm was actually firing on, how that was measured, what changed, and the one question it deliberately leaves open.

The predicate, and why it is almost always false

gate asks check-image-set.sh whether main's HEAD has a complete set. That script answers about four artifacts in ACR:

Artifact What it means
memex-portal-ai:<core-short> the portal image an install pulls, as a linux amd64+arm64 index
memex-migration:<core-short> the schema migration Job's image
mw-plugin-test:<core-short> the tester every satellite resolves
memex-portal-ai:<core-short>-p<plugins-short> the pair tag — provenance, not a thing anything pulls

The first three are deliverability: an install that cannot find them stays on its old image. The fourth is provenance: promote phase A stamps it to record which MeshWeaver.Plugins commit the portal hosts were built from (#2622 — the hosts live in that repo, so a merge there changes what the image should contain while core's HEAD does not move).

All four shared one exit code. So did the reconciler's verdict, and so did the alarm.

Window 1 — the pair tag is unsatisfiable in steady state

gate resolves MeshWeaver.Plugins main afresh on every tick and demands the pair tag for that commit. Measured per day over 2026-09-10…17, MeshWeaver.Plugins main takes 25, 42, 39, 61, 48, 48, 18 and 43 pull-request merges — a mean of 40 a day, so a mean interval near 36 minutes and under 25 on the busiest day. A CD publish takes about 43 minutes (run 35282036886: started 22:26:33Z, promote tags written 23:02:57Z, closed 23:06:50Z).

So on a normal day the ref moves at least once while a publish is running. The pair tag a run writes names a plugins commit that has typically been superseded before verify-images reads it back, and the next tick finds the pairing stale again. It is the shape the memory note "a LIVE CENSUS cannot measure PROGRESS — an arrival cancels a completion one-for-one" describes; the quietest day measured (18 merges, one per 80 minutes) is the one where a tick could occasionally find it current.

That is the rate. The observation is stronger than the rate and does not depend on it: three consecutive hourly ticks on one unchanged core commit each resolved a different plugins HEAD.

The registry records it plainly. Core commit 0dadacc sat as main's HEAD for four hours on 2026-09-17 and was built three times:

3.0.0-ci.N written pair tag digest
3.0.0-ci.8883 22:03:12Z 0dadacc-pf23ce32 sha256:9bfa9888…
3.0.0-ci.8885 23:04:24Z 0dadacc-pa89a016 sha256:597a7651…
3.0.0-ci.8886 00:05:29Z 0dadacc-p27029a5 sha256:3251daeb…

Three issues (#4667, #4669, #4670), three "✅ Healed" closures, one unchanged source tree, and the 0dadacc tag moved twice — orphaning the two earlier images. That is the very shape main-cd.yml's own comment calls a defect under ONE COMMIT, ONE IMAGE SET (#3376).

Across the retained staging tags: 418 publishing builds over 340 distinct core commits — 78 rebuilds of trees that already had a complete, deliverable set. That is a floor, not a total; a staging tag can age out.

The same cycle, measured live

The historical count says what happened. One full cycle was also watched as it ran, on 2026-09-18, which is the same mechanism with no archaeology in it:

time (Z) event
02:38 CD run #8892 publishes 3e7dc7f. Every publishing job green, Verify every image shipped included. Pair tag 3e7dc7f-p752c379.
03:55 MeshWeaver.Plugins main advances to 2648215.
04:33 The tick demands 3e7dc7f-p2648215, does not find it, files #4678.
05:01 The heal writes 3e7dc7f-p2648215 and 3.0.0-ci.8895.
05:06 #4678 self-closes — "main's HEAD now has the complete image set … verified against ACR".
05:30 The next tick demands 3e7dc7f-pefb2449, files #4686. 24 minutes after the heal was confirmed.

That run's gate annotations name one artifact and one only:

memex-portal-ai:3e7dc7f-pefb2449 could not be verified in ACR (az exit 3)

3e7dc7f, memex-migration:3e7dc7f and mw-plugin-test:3e7dc7f all resolved, both architectures, in the same job — and 3e7dc7f-p752c379 and 3e7dc7f-p2648215 were both still in the registry at 05:30. Nothing was deleted; the heal wrote exactly the tag it was asked for; the question simply changed underneath it.

Window 2 — a publisher that has not been created yet

The other window has nothing to do with plugins. gate's only wait condition was PENDING || (!GREEN && AGE_MIN < 120) — it waits while CI is running and, the instant the required check concludes green, treats a missing image set as unhealable. But at that instant nothing has published, and nothing can have:

The INFLIGHT probe that #3376 added cannot cover this: it looks for runs that already exist, and only ones with a lower run id. On 0dadacc the scheduled run (35276745073, created 21:27:14Z) was created 34 seconds before the genuine push-path delivery run (35276798708, 21:27:48Z) — newer id, later creation, invisible to the probe in both directions.

What was measured

Every ci-failure issue titled has an incomplete image set — 109 of them, the complete population — was resolved to its heal comment, to that run's gate job, and to that job's failure annotations, which name the exact repo:tag the probe could not verify.

What the annotations named Issues
Only the pair tag — all three images present and multi-arch 28
The core sha tags too (i.e. HEAD had no set yet), pair tag included 79
The core sha tags only (pre-dating the pair check, 2026-08-30) 2
Unclassified / annotations unreadable 0

For the 25 most recent of the 79, the CD runs on that head sha were then listed, excluding the run that filed the issue and judging each by its state at the moment of filing:

State when the alarm was filed Count
No CD run for that commit had ever been created 19
The only prior run was a zero-job supersede cancellation 3
A run was mid-publish, newer than the tick, so the probe could not see it 3
A prior run had completed and failed to produce the set 0

A zero-job cancellation is the one-pending-slot supersede rule; AGENTS.md says in as many words that it "cannot have torn a set or a seal".

So: 28 alarms over a complete, deliverable set; 81 over a set whose publisher had not run yet or was still running. Zero over a failed delivery.

What it was NOT

Two hypotheses were tested and refuted, because "something is deleting images" and "the heal does not heal" are very different bugs and neither is this one.

What changed

check-image-set.sh answers three things instead of two.

exit 0  complete, and paired with the plugins commit asked about
exit 1  an image of the set is missing or malformed, OR any read failed   → NOT DELIVERABLE
exit 2  every image present and good, pair tag CONFIRMED ABSENT (az 3)    → DELIVERABLE, HOSTS STALE

Two things make exit 2 safe to act on, and both were added after review caught them missing.

Ordering: a run that lost a leg and whose plugins HEAD moved exits 1, never 2 — exit 2 asserts the set is intact, and saying that over a torn set is the one failure the file exists to prevent.

Only a CONFIRMED absence counts. Exit 2 promises that every deliverability image was verified, and gate acts on that promise: complete, no ledger entry, refresh. A 503, a refused pull or an expired credential establishes nothing about the tag, so treating it as "merely behind" would be this same defect one layer down — an unreadable answer read as a benign one. Azure CLI exits 3 for ResourceNotFoundError and 1 or 2 for everything else, so the discriminator is the code, never the absence of an answer; anything but 3 stays RED, naming the failed read.

Both are non-zero, so verify-images and release.yml, which simply run the script, keep today's behaviour exactly: a pair tag a run just wrote and cannot read back is still red there. Only gate inspects the code.

A stale host pairing publishes, silently — and only on a green HEAD. gate still rebuilds (#2622 is unchanged) but says what it is doing and writes nothing to the ledger, the discipline the batching path already follows. The branch sits above the required-check guard, which is right for its neighbours — bake_only ships no image, rebuild is an operator's explicit dispatch — and was a hole here: a HEAD whose check settled red still has its old complete set, so COMPLETE is true and an unguarded refresh would have built and shipped an untested tree. Not green ⇒ fall through to bake_only, which builds nothing.

The ledger requires evidence of an attempt, and evidence that nobody is mid-publish. The fingerprint is in the registry, not in the API: every publishing leg pushes staging-<sha>-<run_id> before promote applies a single consumer-visible tag, and the probe reads all three publishing repositories — the .NET legs run in parallel, so a run whose migration leg staged and whose portal leg died before its push leaves no portal marker at all, which is exactly the torn set the ledger exists to record. That is #1026's shape. Without a marker, the tick is not repairing anything: it is the publisher, and it says so.

A marker is not evidence of failure while its run is still alive, though. The in-flight tie-break at the top of decide filters to runs with a lower id on purpose — two runs deciding at once would otherwise both defer and neither would publish — so a newer run can be mid-publish right here (on 0dadacc the scheduled run was created 34 s before the genuine delivery run; 3 of the 25 issues measured were that). The ledger therefore asks a second question the tie-break deliberately does not: is any run of this workflow live on this sha, other than me? It decides only whether to WRITE, never whether to publish, so it cannot re-create the mutual deferral the id filter prevents.

It is a positive, clock-free condition throughout. No grace period, no retry, no widened bound — a delivery that has not started is a different fact from a delivery that failed, and the probes ask which. Both fail towards the alarm: an unreadable registry answers "attempted", an unanswered run probe answers "nobody else is publishing", because one extra comment costs less than a missed delivery hole — and each says which, so a reader is never guessing.

What this costs

The attempt budget now starts one tick later. A first, ledger-free publish that dies leaves the staging tags that make the next tick open the issue and count attempt 1 — four builds instead of three before cd-unhealed, for a commit that genuinely cannot build. The run that dies is still a CD failed on main alert in its own right, from the verdict job, which this does not touch.

The controls

Both live in the existing harnesses, which extract the real step and the real script rather than restating them.

test-check-image-set.py — a confirmed absent pair tag alone is exit 2 and a ::notice::; an unreadable pair read (503, refused) is exit 1 and an ::error::; a missing image outranks a stale pair; and a complete correctly-paired set is still exit 0 (the inert control, so "exit 2" cannot pass by being answered unconditionally).

test-cd-steps.py — the decide step's cases assert what it asked GitHub to do, read from the recorded gh calls, because a decision that prints nothing alarming and still calls gh issue create is precisely the defect and is invisible in stdout. A host-stale set publishes and files nothing; the same set on a red required check publishes nothing at all; a green commit with no staged layers publishes and files nothing; a newer live run on the sha suppresses the ledger entry without touching the publish decision; and three mutation controls pin the other direction — ATTEMPTED=true with no live run still opens the issue and still records attempt 1/3, an empty probe still fails towards the ledger, and a settled-red check is still reported.

test-cd-steps.py also executes the registry probe itself, against a stub az that models acr repository show-tags per repository. Without that the probe's --query, its repository list and its three branches would be covered by nothing — every decide case hands ATTEMPTED in as an environment value, so a malformed query would pass all of them while the production probe answered wrongly every hour. The cases: nothing staged anywhere ⇒ false; a marker in each of the three repositories alonetrue; a marker for a different commit ⇒ false (the prefix filter is the step's own, run through real jmespath); one unreadable repository ⇒ true, with a warning naming it.

Run against the pre-fix tree: two of the script cases and three of the decide cases fail there, including the one that shows the old step calling gh issue create with the body "every self-updating install stays on the previous image" for a commit no publisher had touched.

The decision that is still open

This does not stop the rebuild loop, and stopping it needs a decision nobody has made. It is tracked on #4688, which carries the same three options and the measurements below.

The reconciler still rebuilds the portal image whenever MeshWeaver.Plugins main has moved — so, at a mean 40 merges a day against an hourly tick, up to 24 full multi-arch builds a day, each publishing a 3.0.0-ci.N and each a publication event the fleet rolls on. The measured sample of plugins merges driving those rebuilds is dominated by commits that cannot change the portal image at all: lock regeneration, i18n mirror syncs, doc and CI changes, Merge main into <branch> — only generated manifest.lock files conflicted.

Core's own side of the same decision already has a relevance filter — gate's relevance step skips a build when every changed path since the newest published set is image-irrelevant. The plugins side has none, and the core filter cannot supply it: when only plugins moved, core HEAD is the newest published set, the compare is empty, and relevant stays true.

Three ways out, none of them free:

  1. A plugins-side relevance filter in core — compare the published image's plugins provenance (recoverable from its own pair tag) against plugins HEAD, and rebuild only when a changed path can enter the image. Same shape as the core filter, same fail-open discipline. It encodes another repository's layout in this one, which nothing here compiles or checks.
  2. The producer moves to the repo that knowsMeshWeaver.Plugins' own lane dispatches core's main-cd with rebuild: true (the operator door that already exists, documented for exactly this case) when a merge touches host paths, and the pair tag leaves the reconciler's publish trigger entirely. Architecturally the right owner; it is a cross-repo change, and per AGENTS.md the half that removes the automatic producer must land last.
  3. Move the host refresh to the daily cadence. The maintainer's 2026-09-12 directive for how this fleet reacts to an upstream repo moving is "do not recompile and run all tests on every platform build; the full run happens once a day". Applying it here is an analogy, not a quote.

Whichever is chosen, the alarm is no longer the thing that has to be silenced to choose it.

The issue body claims every self-updating install stays on the previous image. Measured on the control instance (memex.systemorph.com) on 2026-09-18, the fleet is behind for reasons this defect neither causes nor would fix:

If anything, the loop has been over-supplying images. What holds the portals is update policy and pins, and those are operational decisions, not this lane.

What is NOT established here, said plainly: only memex-cloud's instance-side Admin/UpdatePolicy was read directly. pearl and build have no MCP endpoint in reach, so their records' updatePolicy: Continuous is intent — and both records warn in as many words that the field alone changes nothing the instance does, because the operative setting is the instance's own node. partnerre reads notScraped: true with an empty replica list, so its live state is unknown rather than empty. None of that changes the finding — a fresher image cannot move an install that is not rolling — but it is three instances taken from their records rather than from themselves.

See also

Reconnecting…
The connection to the server was interrupted. Trying to restore it…
Trying again…
The connection could not be restored. Reloading the page…
The server was updated. Reloading the page to pick up the latest version.