A teardown verdict arrives on a CAUSE, not on a clock

Some of this system's most important answers are minted during a hub's teardown, not during its working life. The owner-disposing patch NACK is the canonical one: when a per-node hub goes down owing an answer for an in-flight PatchDataRequest, a RegisterForDisposal registrant mints MeshNodeErrorCode.OwnerDisposing and hands it to the waiting caller (DataExtensions.RegisterOwnerDisposingNack, #2778).

That registrant runs in the ShutDown phase — the LAST phase of the hub's teardown. Everything in front of it must finish first: Quiescing, then the whole DisposeHostedHubs join over the hub's own subtree.

The rule this page exists to state is:

Nothing bounds how long that takes, on purpose. So nothing downstream may budget it in seconds — not production code, and above all not a test.

Why there is no bound, and why that is right

HostedHubsCollection.DisposeHubsReactive used to cap the hosted-hub join at a flat Timeout(5s). #1317 removed it, and the reasoning is worth restating because it is exactly the reasoning that applies one layer up:

What replaced it is a stall detector, not a deadline (MessageHub's disposal watchdog, #1701). It is re-armed by SubtreeDisposalProgress on every RunLevel transition anywhere in the subtree. So a large teardown that keeps making progress never trips it, however long it takes — which is the correct behaviour and is precisely why "the teardown is taking a while" produces no report.

Read those two facts together and the consequence is unavoidable: a healthy teardown has no upper bound in time, and no signal that says how long it took.

The failure this produced (measured)

NackReachesTheWaiterDuringTeardownTest is the regression test for #2778. It parks the owner's merge turn, posts a patch, fences on the caller being armed in LatePatchResponseRegistry, disposes the mesh, and then asserted that the armed watch was consumed within 6 seconds.

That bound was chosen for two good reasons and one bad one. The good reasons: it is far below LateResponseWatchBound (30 s), so an entry that merely expired could not be mistaken for one that was answered; and it is below the 8 s disposal stall budget, so a stall verdict could not be what produced the answer. The bad one, unstated: it also raced the owner's entire teardown, which the paragraphs above say is unbounded.

The rate, and why the asymmetry is the finding

repo / suite failures executions examined window rate
MeshWeaver.Plugins, src/MeshWeaver.Hosting.Monolith.Test slice #2 5 138 2026-09-07T15:19Z → 2026-09-08T07:12Z 3.6% — 1 in 28
MeshWeaver (core), test/MeshWeaver.Graph.Test twin 0 796 runs 2026-09-05T15:24Z → 2026-09-08T07:23Z 0%

Same test body — they are twins, held in step by Plugins' TeardownTwinParityTest — and the same framework. What differs is the size of the mesh being torn down, and therefore how long the owner's teardown takes. A flat 6 s is comfortable in core's fixture and marginal in the portal's, which is exactly what "asserting a bound the framework does not provide" looks like from the outside: it tracks mesh size, not correctness.

Two measurement notes worth keeping, because both would have understated it:

The failing run in detail

Run 34195323935, Portal hosts (shard 2), 2026-09-08. The evidence is entirely negative, and that is the tell:

signal present? what its absence rules out
LATE_NACK_REENQUEUE (Warning) no the caller's watch never received any verdict, early or late
[PatchAck] … reached no route (Warning) no the owner never minted a verdict and lost it on both routes
[LateWatch] VERDICT_EXPIRED (Warning) no no verdict arrived after the 30 s window either
Message delivery failed … cannot route (Warning) no routing never refused the patch or its answer
[QUIESCE-WAIT] / [QUIESCE-CUT] no no hub re-armed its quiesce budget on an owed reply
DISPOSAL DEADLOCK DETECTED / [DISPOSE-WEDGE] no the stall detector never fired: the teardown kept moving
MEM_WATCHDOG on its normal 5 s cadence yes the process itself was healthy and not starved

The one warning in the window was [QUIESCE-TIMEOUT] cache/…: 0 callback(s) still pending after 2s — Pending: <none>, which OnQuiesceComplete documents as the benign shape: the poll said not-drained and the dictionary was empty by the time the verdict was rendered.

Nothing had gone wrong. The owner had simply not reached its ShutDown phase yet. The test asserted a timing guarantee that does not exist, and the failure it produced was a bare System.TimeoutException naming a line number — no state, no armed ids, no run level. Six CI-artifact greps were needed to establish "nothing happened", which is why this had recurred without ever being diagnosed.

The shape that is correct

Assert the cause, then read the state. The ordering inside MessageHub.HandleShutdownCore is fixed and is what makes this exact rather than lucky:

ShutDown phase:
    CancelCallbacks()
    DisposeImpl()          ← disposables.Dispose() runs the RegisterForDisposal registrants
                             — this IS where the NACK is minted and dispatched
    messageService.Dispose()
    RunLevel = Dead        ← set BEFORE the signal, deliberately (see its comment)
    SignalDisposalCompleted()

So RunLevel == Dead (equivalently, DisposalCompleted having fired) is the exact instant at which "the owner has had its chance to answer" becomes true. Waiting for that and then reading the registry is a decision, not a race:

// Wait for the CAUSE …
await Observable.Interval(50.Milliseconds()).StartWith(0L)
    .Where(_ => owner.RunLevel == MessageHubRunLevel.Dead)
    .FirstAsync().Timeout(ownerDownBound).Await(ct);

// … then the assertion is instantaneous and unconditional.
registry.ArmedRequestIds.Should().NotContain(armedId, "…the #2778 defect…");
registry.ExpiredVerdicts.Should().Be(0, "CONSUMED, not lapsed");

Three things get strictly better:

  1. It still catches the defect — sooner. Falsified by restoring the pre-#2778 shape (dropping the ILatePatchVerdictSink dispatch so only the run-level-guarded parent post remains): the test fails in 2 s instead of 6, and the message reads "Did not expect collection {"gKQ_…"} to contain "gKQ_…" because the owner minted an OwnerDisposing NACK …" instead of TimeoutException. The owner reaches Dead just as fast when it is broken; it is the watch that is still armed.
  2. "Consumed" stops being inferred from a stopwatch. The 6 s bound was standing in for "it did not merely expire". LatePatchResponseRegistry.ExpiredVerdicts counts that case directly — #3197 made an expired verdict a reported fact rather than silence — so the distinction is now observed, and it stays true however long the teardown took.
  3. The remaining bound guards a PRECONDITION and says so. If it trips, the failure is "the owner never finished disposing", which is a different defect with a different investigation — and the test's stall-verdict assertion (verdicts.Entries.Should().BeEmpty()) names it. It must stay below LateResponseWatchBound, because a question asked after the entry has lapsed cannot be answered either way; deriving it (LateResponseWatchBound - 10.Seconds()) keeps that true without a second literal, per Bounds must be ordered.

🚨 Correction (#4023, 2026-09-11): Dead is the instant the owner has had its chance to ANSWER — not the instant the answer has ARRIVED. Of the routes an answer takes to the armed watch, two complete on the owner's own turns: the ShutDown-phase registrant's direct Dispatch, and the undeliverable-reply sink that routing offers a reply to when the parent is past DisposeHostedHubs. For those, the reading above is exact. The third does not: since #3291 the parked merge usually COMMITS during teardown, and a success ack posted while the parent is still below DisposeHostedHubs travels the ordinary route — owner → mesh → the caller's hub — whose last two queues are crossed after the owner reaches Dead. Measured once in 3,600 stressed iterations (the reply sat in the caller's hub queue at Dead and was delivered 12 ms after the assertion). The other failures on that route were a real defect — the mesh's router dropped the reply — and are fixed; see Reading a Disposal Stall Verdict → "Resolved 2026-09-11". A twin that judges at the owner's Dead can still report the residual race.

What makes Dead safe to read: the run level is monotone, and that is enforced

The shape above rests on one property that is easy to assume and was not, until #3647, actually guaranteed: RunLevel never goes backwards. Read the ordering again and notice that every phase of HandleShutdownCore is entered by handling a ShutdownRequest, and that two of the three idempotency guards test RunLevel == <their own phase> rather than >=:

case MessageHubRunLevel.Quiescing:          if (RunLevel >= Quiescing)          return Ignored;   // ✅
case MessageHubRunLevel.DisposeHostedHubs:  if (RunLevel == DisposeHostedHubs)  return Ignored;   // ⚠️
case MessageHubRunLevel.ShutDown:           if (RunLevel == ShutDown)           return Ignored;   // ⚠️

A ShutdownRequest handled after RunLevel = Dead therefore does not hit its guard: it re-enters the phase and assigns the run level back down to DisposeHostedHubs or ShutDown. Every caller and every test that this page tells to wait for Dead and then read state would be reading a value that can be un-set behind it.

Nothing can produce that today, and the reason is worth stating because it is what the fix makes structural rather than incidental. ShutdownRequest is internal and is deliberately absent from the default TypeRegistry, so no wire sender can mint one; in-process there are exactly three posters (MessageHub.Dispose, OnQuiesceComplete, PostShutDownPhase) and each fires once per hub. The intake gate now says so as well: MessageService.RefusesIntake exempts teardown's own traffic only while a phase is left for it to advance — ShutdownRequest up to ShutDown, DisposeRequest up to Quiescing — instead of exempting both at every level forever. So the monotonicity Dead depends on is a property of the gate, not of who happens to post.

The same unbounded exemption was what produced #3647's report, and Teardown layers carries that half: a DisposeRequest admitted into a window where its handler is a proven no-op became a queued turn, and messageService.Dispose() filed it at Error as discarded work — while the pump, one stack frame below, drained and answered it 3 ms later.

A second defect the same change removes

The old assertion polled registry.ArmedCount == 0 — a whole-registry count. But dispatching the NACK runs the caller's LATE_NACK_REENQUEUE arm, which posts a fresh PatchDataRequest and arms a second watch microseconds later. ArmedCount therefore goes 1 → 0 → 1 → 0, and the zero the test was looking for exists for less than a millisecond. It passed only because the re-attempt is refused quickly during a teardown — i.e. the assertion was reading a transient it had no right to observe. ArmedRequestIds exists for exactly this (its own remarks say so): name the correlation the verdict has to land on, and a second, unrelated arm cannot falsify it.

The rule, generalized

If an answer is produced by a RegisterForDisposal registrant, a ShuttingDown handler, or anything else in a hub's teardown, then its latency is the teardown's latency — which is deliberately unbounded and deliberately unreported.

Reconnecting…
The server was updated. Reloading the page to pick up the latest version.