A held gate and a spent code are recorded, not ticketed
Systemorph/MeshWeaver#5722 and #5254. Both follow the rule roll churn established: a line the platform cannot re-level and that reports a mechanism doing its job is recorded on its own incident and not triaged, and the classification is narrow enough that the fault next to it still files.
The held readiness gate (#5722)
nodetype_bake is the startup gate that keeps a rolling replica out of rotation until its dynamic
NodeTypes are built against the image it runs. While the bake runs its verdict is Unhealthy, and
it has to be: the health endpoint maps Degraded to HTTP 200, so a softer status would declare the
pod started mid-bake. ASP.NET's DefaultHealthCheckService logs every Unhealthy result as event
103 at fail:, on every probe. So every replica prints one red line per probe for the length of its
bake, on every boot.
What that did to one ticket, counted over the lines folded into #5722 between 2026-09-26 and 2026-10-11:
| line | count |
|---|---|
nodetype_bake — bake in progress (five phase wordings) |
422 |
nodetype_bake — regressed on this image |
29 |
db_version — check threw |
11 |
PostgreSql — name resolution failed |
3 |
The issue is titled for the last row. The in-progress line words its phase differently as the bake advances, so the site outran the per-site variant budget in almost every window and was folded to one site-level incident. The other three checks log through the same logger with the same event id and no application frame, so they share that site and went into the same ticket, under 160 recurrence comments about a gate that was working.
The rule (ReadinessGateClassifier, in MeshWeaver.Observability.Contract). A burst is a held
gate only when all hold: the category is the health-check service, the event id is 103, and the
line reports the check named nodetype_bake as Unhealthy with a message that starts with
NodeType bake in progress. Both strings are constants the host itself uses to register the check
and to build that description, and NodeTypeBakeProbeStatusTest builds the line from the check's
real results, so the two cannot drift apart.
What it does not classify. The same gate saying regressed or FAULTED: that is the gate refusing an image, and it must reach a person. Every other health check. A bake that never finishes is not hidden by this either — the startup probe's bound fails the pod and the rollout stalls, which the instruments that watch rolls report.
Taking the held-gate lines out of the site's variant count is what lets the remaining checks keep
their own identities: a PostgreSql name-resolution failure and a db_version fault are now two
incidents, each with one variant.
The spent authorization code (#5254)
When the provider's token endpoint answers an error, ASP.NET's OpenIdConnectHandler logs it at
fail: twice (events 52 and 17) before any application code runs. A filter on that category would
also hide a wrong client secret or a revoked app registration.
Two of those answers are the protocol working. An authorization code is single-use and short-lived
by specification: AADSTS54005 says it was already redeemed, which takes two submissions of one
callback in flight at once (a double POST from the browser), and AADSTS70008 says it expired
before it was posted (a tab left open). The portal redeems a code once — no retry handler reaches
the handler's back channel, and a plain replay fails earlier, on the correlation cookie — and the
next sign-in starts a fresh flow.
The rule (SpentAuthorizationCodeClassifier). The category is the sign-in handler, a line
carries the token response's error: 'invalid_grant', and the same line's error_description
starts with one of those two reasons, by number and with its colon.
What it does not classify. Every other invalid_grant reason (a code issued for another
redirect URI is a misconfiguration), every other OAuth error (invalid_client is a wrong or
expired secret), and the same text under any other category.
The portal's own line. ReportRemoteFailure logged the same event a third time, at Error,
under the provider's auth category. It now reads the token response's error and
error_description from the exception's protocol fields and logs a spent code at Information.
The cost of that level is nothing a reader loses: Information ships to the log store, so the
event stays countable and a change in its rate stays visible. What it buys is that a
double-submitted callback no longer opens an incident about a sign-in whose first submission had
already redeemed the code. Every other token-endpoint refusal keeps Error.
Recorded, never dropped
Each class goes to its own incident — Admin/_LogIncident/readiness-gate-{siteFold} and
Admin/_LogIncident/spent-authorization-code-{siteFold} — which the portal records as Suppressed
from its first sighting, with the classification and the evidence that produced it. They are
counted and sampled like any incident and ask for nothing.
LogIncidentClassification.IsRecordedNotTriaged is the one list of values the portal treats this
way. An unknown value, a module's own included, is an ordinary fault.
What has to roll
The classification runs in the watcher, which is its own image (memex-log-watcher). Until
that image is rebuilt from this change and rolled, the lines above keep folding as before. The
portal half (IsRecordedNotTriaged, the Information level) arrives with the portal image. The
contract is additive in both directions: an older watcher sends no classification, and an older
portal triages a readiness-gate-… report as an ordinary incident.
Related
- Roll churn is recorded, not ticketed — the first classification, and the mechanism these two reuse.
- One log site, one ticket — the site fold.
- The category is the logger, not the subject.