One Log Site, Many Terminals

A bug-filing fingerprint is computed from the message a log site reported, not from the code path that reported it. That single fact produces two opposite errors at once, and a cluster of timeout tickets will contain both:

what it looks like in the tracker what it costs
the SPLIT two or more tickets, different fingerprints, the same log site, differing only in the INNER exception the same line is fixed twice, or fixed once and the other ticket outlives the fix
the FOLD one ticket whose samples carry several different inner exceptions one occurrence count over several roots; a fix is judged against a counter it cannot move

Neither is visible from a ticket's title. Both are visible from its samples.

Group first, and group on the TERMINAL

The grouping key that works is not the logger category and not the exception type the outer frame carries — those are what the fingerprint already used. It is the innermost terminal: the sentence produced by the code that actually gave up.

The mesh's write-wait — GetMeshNodeStream(path).Update(…) waiting for the owning per-node hub to deliver its first frame — produces four of them, and they have four different causes and four different owners:

terminal what it means whose defect
BaseStateTimeoutException"NO change item at all reached the mirror" nothing arrived inside the bound unresolved: the owner's activation, its routing, or its initial read
InvalidOperationException"this hub's mirror ended without ever carrying the node's state" the mirror COMPLETED without state mirror lifecycle — acquired and not held for the write
MeshNodeStreamException"Host is shutting down, cannot route to …" the write was issued during shutdown the caller: a reconcile that keeps working while the host stops
MeshNodeStreamException"the owner returned no verdict for this update within …" the delivery reached a hub and no handler ran delivery/routing for that owner

The third is the one that matters most for grouping, because it is not a mesh defect at all. A write refused at shutdown is the platform answering correctly; what is wrong is that the caller issued it and then reported the refusal as a failure. A ticket whose samples are mostly that terminal is a classification ticket, and reading its occurrence count as a mesh-reliability number is reading a shutdown census.

How to tell a SPLIT from two real tickets

Two signals, both on the ticket itself, neither of which requires reading code:

  1. The ancestor fingerprint. A re-addressed incident records the identity it inherited. Two tickets naming the same ancestor are one log site the identity function later separated — that is a split, and one of them is a duplicate of the other.
  2. The log line prefix. Two tickets whose samples share the same message template up to the ---> are the same site whatever their fingerprints say.

When both hold, close the newer as a duplicate naming the survivor, and say which terminal the closed one contributed — otherwise the fold loses the one piece of information the split had.

When a "duplicate" must NOT be closed

A diagnosed root and an undiagnosed symptom family that share a log site are not interchangeable. Folding the family into the root buries whichever of the two is wrong: if the root is fixed and the family keeps firing, the fix reads as ineffective; if the family is closed on the root's fix, its remaining causes vanish from the queue.

So the rule has a direction: a split is closed toward the older, better-diagnosed ticket only when the terminal is the same. Same site, different terminal ⇒ keep both, and cross-reference them by terminal rather than by title.

How to tell a FOLD

Read every retained sample's inner exception, not the first one. A ticket folding several terminals is a ticket whose count is a sum over different defects — so:

A frozen lastSeen therefore has three causes and only one of them is good news: the fault stopped, the fingerprint moved, or the ingestion that folds sightings is itself down. Say which one you established.

And before any of that: ask where the log statement SITS

A count is only a count of failures if the line that produced it was written after the outcome was decided. Often it is not. A handler that logs at Error and then runs the branch which retries, recovers or reclassifies emits one line per attempt — so the counter measures the log statement's position in the code, not the defect.

That makes the count unsound in both directions, which is worse than merely useless:

The tell is in the message itself. A line that says "it still carries its request and is picked up again" — or anything else naming its own recovery — is reporting an attempt, and its count is a census of retries. Read it that way before treating it as damage, and prefer a per-terminal reading of the retained samples: which code path was reached is a fact the line CAN establish, because the inner exception is produced at the point it names.

Full account of the general shape: A Fault Does Not State the Verdict.