Source-Set Establishment

A NodeType's compile is only a verdict about its code if the compiler was shown that code. When the source set is short, Roslyn is still perfectly correct — it reports what it was given — and the diagnostic it produces is indistinguishable from a genuine break:

CS0246  The type or namespace name 'CessionData' could not be found
CS0103  The name 'CessionSampleData' does not exist in the current context
CS1061  'LayoutDefinition' does not contain a definition for 'AddSocialMediaPostLayoutAreas'

Every symbol in those three lines is declared in the same NodeType's own Source/ folder. The code was never broken; the pass that went looking for it did not find it.

The platform already knows this failure mode: SourceSnapshot in MeshWeaver.Compiler exists to keep "the set is empty" apart from "the set could not be established", and it says so at length. The mistake this page is about is not the absence of that idea — it is that the batched bake's implementation of it had a condition on it that switched it off for almost every NodeType in a real mesh.

The three shapes of zero

There is exactly one honest question to ask of a resolved set of size zero: does the mesh agree?

Evidence What zero means Verdict
CurrentSourceVersions is {} (explicitly empty) the sources were deleted, or the type is configuration-only Established. Classify NoSources; must never gate a rollout — no image can change what a mesh query matches
CurrentSourceVersions has n > 0 entries the mesh says this type HAS n source files Unestablished. The pass contradicts the record, and a contradiction is not a verdict
CurrentSourceVersions is absent (SQL NULL) no witness either way fall back to whether the type declares its own queries

CurrentSourceVersions is written by each NodeType's own sources watcher and persists on the NodeType's MeshNode, so it survives a failed compile and it is readable from a cold boot before anything has been compiled. It is the only independent witness available at discovery time, and it is used in both directions — the same field, read for both polarities, which is what makes the two classifications agree by construction rather than by coincidence.

NodeTypeBatchBake.DiscoveryUnestablished is the one predicate. When it answers true the whole batch is abandoned with SourceDiscoveryFailedException and the pod falls back to the activation-driven sweep, which re-resolves each type individually — slower, correct, and impossible to mistake for a content verdict.

What went wrong (issue #3663)

The predicate used to require that the type declare its own source queries:

// before
if (matched.Count == 0 && pending.DeclaresSources && !pending.SourcesKnownDeleted)

Very nearly no NodeType declares them. The default queries — namespace:{path}/Source scope:subtree nodeType:Code and the matching Test one — are how the population is authored, which DynamicTypePreWarmer.ClassifyCompileFailure had already been corrected to say in #1391: "an empty Sources does not mean configuration-only — it means uses the DEFAULT queries." The same conjunct sat in Assemble, in the opposite direction, and it meant the invariant guarded a tiny minority of types and waved the rest through.

So a short discovery pass produced compile errors, the errors were classified CompileError (the snapshot was populated, so NoSources did not apply either), and NodeTypeBakeGateState recorded four regressions on a healthy image.

The measurement

BatchBake logs the size of every pass. On memex.systemorph.com, every boot in the window 2026-09-07 22:40Z → 2026-09-08 07:37Z:

Boot (UTC) Code nodes resolved compileErrors= Readiness
09-07 22:40:02 1236 2 granted
09-07 22:43:42 1236 2 granted
09-08 00:31:22 1145 6 REFUSED
09-08 03:32:37 1237 2 granted
09-08 03:35:30 1237 2 granted
09-08 07:34:44 1241 2 granted
09-08 07:36:57 1241 2 granted

Two of those errors are the portal's standing baseline — types already at Error before the deploy, which MarkOutcome correctly refuses to let gate. The four extra are the entire Doc partition's NodeTypes (…/BusinessRules/Cession, …/PythonPandasNode/PandasExplorer, …/SocialMedia/Post, …/SocialMedia/Profile) — four of four, not four of forty — every one of which uses the default queries and carries a populated CurrentSourceVersions.

The denominator is 29 boots across both portals in fifteen hours, and exactly one of them refused. The image was identical on three of the six boots that granted readiness. Nothing about the image explains the difference; the size of one query pass does.

With a positive control, because a zero needs a denominator. The same instrument, the same 15-hour window, counted over the whole window on the instant endpoint:

count_over_time({namespace="memex"}       |= "NodeType bake regressed" [15h])   →  1065, ONE stream
count_over_time({namespace="memex-cloud"} |= "NodeType bake regressed" [15h])   →     0

The 1065 all come from 7d5d458cc4-cbztk's boot 0 and nowhere else — one line per readiness probe at the 10-second cadence for 2 h 58 m, which is the startup budget below, and the reason the memex count is non-zero is what makes the memex-cloud zero mean something. In fifteen hours, across two portals, exactly one container boot ever put the bake gate into Regressed. The simultaneous memex-cloud stall did not involve this gate at all.

The prebuilt bytes were already there. The same boot logged ShippedPrebuiltBundles: bundle Doc.zip: adopted 4/4 prebuilt assembly(ies) at 00:31:08 — ten seconds before the sweep enumerated 204 of 209 … need building — 5 already on the share and set about recompiling them. Adoption seeding and the sweep's own store probe disagreeing on a cold boot is what put those four types on the compile path at all; it is a separate seam, it is not fixed here, and it is tracked as issue #3703 — which carries both log lines, the control that makes it a disagreement rather than a cold-store fact (a boot with 77 already current sees 189 baked; this boot, with 78 adopted now, saw 5), and the measurement that would name the mechanism.

Why it lasted three hours, and why that is the dangerous part

A recorded regression is not sticky by design — NodeTypeBakeGateState.RetractRegression is level-triggered, and a type observed reaching a usable build on the same image has its regression withdrawn (issue #1214). But retraction is driven by the type recompiling, and nothing recompiled: the content never changed, so the park registry's source-change retry never fired. The verdict was therefore correct-and-permanent for as long as the process lived.

What ended it was the kubelet. The portal's startupProbe is failureThreshold: 1080 at periodSeconds: 10three hours — and the container was killed at 03:29:52Z, 2 h 59 m after it started at 00:30:23Z. The replacement boot's discovery pass came back complete and the pod went Ready at once.

🚨 A rollout that stalls for hours and then succeeds is a worse failure than one that never succeeds. It presents as a slow deploy, its recovery looks like the system healing itself, and the next occurrence reads as flakiness. The window is bounded by a probe budget, not by anything that understands the defect.

What did NOT cause it

What is NOT established: why the pass was short

This page fixes what a short pass is allowed to CONCLUDE. It does not explain why that one pass was short, and nothing here should be read as if it did.

The leading candidate is the completion rule. RunQuery accumulates a query's chunked Initial and treats one second of silence as "the answer is complete" (QueryQuietWindow, then .Throttle(…).Take(1)). QueryResultChange<T> carries no terminal marker — Initial, Added, Updated, Removed, Reset and nothing that says done — so a quiet window is the only completion signal available to the reader, and a chunk gap wider than it silently truncates the fold. The suspect boot ran its discovery 14 seconds after a 35-second burst of bundle seeding onto the shared volume, which is exactly the kind of contention that widens a gap.

That is a hypothesis, and it has not been measured. It is tracked as issue #3704, which carries this section's content plus the ruling-out below, so the next reader starts from the instrument rather than rebuilding it. What would settle it: instrument RunQuery to record, per query, the number of change events folded and the largest inter-chunk gap, then compare a short pass against a complete one on the same portal. A pass whose largest gap approaches QueryQuietWindow names the completion rule; one whose gaps are all small says the shortfall is upstream, in what the providers returned, and the search moves to the static catalog (which is the only route by which a Doc/** Code node reaches discovery — its nodes are served from the image, not from the Postgres code satellite, so they arrive through the unpinned nodeType:Code fetch alone).

🚨 Do not "fix" this by widening QueryQuietWindow. A longer window makes a short read rarer without making it impossible, and the invariant above is what makes rarity irrelevant: a short read now produces "I don't know" instead of a verdict. Widening the bound would trade a correctness property for a probability.

Re-measuring this

One Loki query answers whether a pass was short, and it needs no pod to still exist:

{namespace=~"memex|memex-cloud"} |= "source discovery resolved"

Read the Code-node count per boot and compare it against its neighbours on the same portal — the two portals hold different content and their absolute numbers are not comparable. A boot whose count sits below its neighbours' resolved a short set, and every compile verdict from that boot is suspect. Pair it with:

{namespace=~"memex|memex-cloud"} |= "warm-up complete"

whose compileErrors= is the portal's standing baseline plus whatever that boot invented.

Always write explicit start/end in nanoseconds — see Measuring a Live Portal Read-Only, whose first trap (since= is silently ignored) applies to both queries above.

Reconnecting…
The server was updated. Reloading the page to pick up the latest version.