A Fault Does Not State the Verdict

A line that reports a fault may state the fault. It may not state the decision that another branch makes after it runs.

This sounds like a style rule about wording. It is not. When the decision moves — when someone adds a second door, a fallback, a recovery path — the fault's sentence does not move with it, and it becomes a confident assertion about something that did not happen. Nothing fails to compile, nothing goes red, and the claim is now the only thing the observability pipeline keeps.

The shape

  ┌─ the fault is detected ────────────────────────────────┐
  │  logs at Error: "X failed, SO THE OUTCOME IS Y"        │  ← states a verdict it cannot know
  └────────────────────────────────────────────────────────┘
                            │  (the exception propagates)
                            ▼
  ┌─ a later branch DECIDES the outcome ───────────────────┐
  │  recovered → logs at Warning: "actually, not Y"        │  ← never captured
  │  refused   → logs at Error:   "Y"                      │
  └────────────────────────────────────────────────────────┘

The fault line is written first, so at the moment it runs the verdict does not exist yet. On the recovered path it is simply wrong.

Why it is not cosmetic: the capture rule turns the claim into the record

The load-bearing property is the capture rule, not the fingerprint (see Log Watch Triage): only fail: and crit: are captured — Error and Critical. A Warning is not collected, not fingerprinted, and not ticketed.

So on the recovered path:

The false claim becomes the permanent, ticketed record of an event that went fine, and the pipeline holds no trace of the recovery that contradicts it. That asymmetry is the defect, and it does not depend on how incidents are folded.

What the wording does and does not do

🚨 It is tempting — and wrong — to say "the fingerprint is the message". The identity (StructuralLogIncidentIdentity.Compute) is a hash over three parts: WHERE (the top application stack frame, or (category, eventId) when the burst names no frame), WHAT (the exception type by simple name), and WHICH (the masked exception message — the logged message only when there is no exception). The contract is explicit that the discriminating text is the exception's message, never the reporter's prose, and a LogIncident node's stored normalizedMessage field is not the identity. That distinction is worth holding precisely — a half-remembered version of it is inherited by every later reader.

Two consequences worth getting right, because the obvious reading of each is wrong:

So the benefit is narrower than "the harmful case gets its own ticket", and it is still worth having: the ticket a benign event opens now describes a transport fault rather than a held rollout. An issue titled after the bad outcome stops being reopened by events that are not instances of it — not because the identity now discriminates, but because the wording no longer describes that outcome at all.

The measured instance

The pre-warmer's build handshake is a SubscribeRequest to the build coordination root (see Build Coordination). When it goes unanswered, BuildProtocolDriver.RetryUnreachableCoordination exhausts its attempts and logs, at Error:

BuildProtocol: could not reach the build coordination node 'Admin/Build' in 3 attempt(s) — the pre-warm sweep never started, so this process has verified NOTHING about its NodeTypes on this image. This is a refusal, not a pass: readiness stays refused and the rollout holds the previous image. A restart re-attempts.

That last sentence was true while the subscription was the only door. Then WhenTheSubscriptionDoorIsShut was added: it catches this exception and asks the durable witness, and when the witness already carries the GO for this framework it grants readiness — reporting that at Warning.

From then on, every unanswered subscription published one Error asserting that readiness was refused and a rollout was held, whether or not either happened. Measured consequence: the incident for held rollouts was correctly closed twice with evidence, and reopened by an occurrence on an image three framework builds newer than the door that fixed it — datable from the log line itself, whose Queue(…) diagnostic carried handledWhileWaiting and a Trail: block that the original samples do not have. The reopen was legitimate by the pipeline's own rules; the line it rested on was not.

The cost is paid twice. An operator reading it goes to look at a rollout that is fine, and the issue queue carries an open item whose condition is not occurring.

The rule, and what it does not say

State what you know. Name what you could not do. Do not name the consequence unless you are the code that decides it.

Concretely, for the fault line:

And for the deciding branch: it states the verdict, at the severity the verdict deserves — the refusing branches at Error, in as many words, so the operator-facing claim about held rollouts is still made exactly when it is true.

🚨 This is not a severity change and not a visibility reduction. The fault keeps its Error. The temptation is to "fix" the false ticket by demoting the fault to Warning so the watcher stops seeing it — that deletes the only evidence that the transport is broken, which is the opposite of the goal. The fix is what the line claims, never whether it is reported.

🚨 It is also not a reason to stop logging early. The fault must be logged where it happens, with its inner exception and its diagnostic detail, because the deciding branch has less context. Both lines are wanted; only one of them may state the outcome.

How to tell whether you have this defect

Ask one question of any line that names an outcome:

Is this line written before or after the code that decides the outcome?

Before ⇒ it must not name it. The test is mechanical and does not require judgement about wording.

Two supporting smells:

Testing it, with a control on each side

A single assertion here is worthless, because "no line claims a refusal" passes trivially against a test fixture that supplies its own message. Two things make it real:

  1. Drive the production chain, not a fixture. Compose the real retry with the real door, so the message under assertion is the message minted in production. A fixture that hard-codes the wording tests the fixture.
  2. Assert from both sides with ONE predicate. The recovering case asserts the predicate does not fire; the refusing case asserts it does, on a line at Error that also names the witness it read. A predicate that could never match — a typo, a reworded production line — would make the recovering case pass having checked nothing, which is this defect's own failure shape one level up.

That pair establishes the property that matters: the claim moved to the code that decides it, rather than having been deleted.

Measured against the fix in place, both cases pass; with the pre-fix message restored, the recovering case fails on exactly that assertion. See PreWarmerReadsTheDurableGoTest.TheTransportFaultStatesNoVerdict_WhenTheDurableGoGrants and …TheDoorThatRefuses_DoesStateTheRefusal.

The family this belongs to

The same root, in both directions:

All three are one question: does this line speak for something it actually established?

Applying it beyond logs

The rule generalises to anything a later stage can overturn:

In each case the same repair applies: say what happened here, name where the outcome is decided, and let the deciding code say it.