A held gate and a spent code are recorded, not ticketed

Systemorph/MeshWeaver#5722 and #5254. Both follow the rule roll churn established: a line the platform cannot re-level and that reports a mechanism doing its job is recorded on its own incident and not triaged, and the classification is narrow enough that the fault next to it still files.

The held readiness gate (#5722)

nodetype_bake is the startup gate that keeps a rolling replica out of rotation until its dynamic NodeTypes are built against the image it runs. While the bake runs its verdict is Unhealthy, and it has to be: the health endpoint maps Degraded to HTTP 200, so a softer status would declare the pod started mid-bake. ASP.NET's DefaultHealthCheckService logs every Unhealthy result as event 103 at fail:, on every probe. So every replica prints one red line per probe for the length of its bake, on every boot.

What that did to one ticket, counted over the lines folded into #5722 between 2026-09-26 and 2026-10-11:

line count
nodetype_bake — bake in progress (five phase wordings) 422
nodetype_bake — regressed on this image 29
db_version — check threw 11
PostgreSql — name resolution failed 3

The issue is titled for the last row. The in-progress line words its phase differently as the bake advances, so the site outran the per-site variant budget in almost every window and was folded to one site-level incident. The other three checks log through the same logger with the same event id and no application frame, so they share that site and went into the same ticket, under 160 recurrence comments about a gate that was working.

The rule (ReadinessGateClassifier, in MeshWeaver.Observability.Contract). A burst is a held gate only when all hold: the category is the health-check service, the event id is 103, and the line reports the check named nodetype_bake as Unhealthy with a message that starts with NodeType bake in progress. Both strings are constants the host itself uses to register the check and to build that description, and NodeTypeBakeProbeStatusTest builds the line from the check's real results, so the two cannot drift apart.

What it does not classify. The same gate saying regressed or FAULTED: that is the gate refusing an image, and it must reach a person. Every other health check. A bake that never finishes is not hidden by this either — the startup probe's bound fails the pod and the rollout stalls, which the instruments that watch rolls report.

Taking the held-gate lines out of the site's variant count is what lets the remaining checks keep their own identities: a PostgreSql name-resolution failure and a db_version fault are now two incidents, each with one variant.

The spent authorization code (#5254)

When the provider's token endpoint answers an error, ASP.NET's OpenIdConnectHandler logs it at fail: twice (events 52 and 17) before any application code runs. A filter on that category would also hide a wrong client secret or a revoked app registration.

Two of those answers are the protocol working. An authorization code is single-use and short-lived by specification: AADSTS54005 says it was already redeemed, which takes two submissions of one callback in flight at once (a double POST from the browser), and AADSTS70008 says it expired before it was posted (a tab left open). The portal redeems a code once — no retry handler reaches the handler's back channel, and a plain replay fails earlier, on the correlation cookie — and the next sign-in starts a fresh flow.

The rule (SpentAuthorizationCodeClassifier). The category is the sign-in handler, a line carries the token response's error: 'invalid_grant', and the same line's error_description starts with one of those two reasons, by number and with its colon.

What it does not classify. Every other invalid_grant reason (a code issued for another redirect URI is a misconfiguration), every other OAuth error (invalid_client is a wrong or expired secret), and the same text under any other category.

The portal's own line. ReportRemoteFailure logged the same event a third time, at Error, under the provider's auth category. It now reads the token response's error and error_description from the exception's protocol fields and logs a spent code at Information. The cost of that level is nothing a reader loses: Information ships to the log store, so the event stays countable and a change in its rate stays visible. What it buys is that a double-submitted callback no longer opens an incident about a sign-in whose first submission had already redeemed the code. Every other token-endpoint refusal keeps Error.

Recorded, never dropped

Each class goes to its own incident — Admin/_LogIncident/readiness-gate-{siteFold} and Admin/_LogIncident/spent-authorization-code-{siteFold} — which the portal records as Suppressed from its first sighting, with the classification and the evidence that produced it. They are counted and sampled like any incident and ask for nothing.

LogIncidentClassification.IsRecordedNotTriaged is the one list of values the portal treats this way. An unknown value, a module's own included, is an ordinary fault.

What has to roll

The classification runs in the watcher, which is its own image (memex-log-watcher). Until that image is rebuilt from this change and rolled, the lines above keep folding as before. The portal half (IsRecordedNotTriaged, the Information level) arrives with the portal image. The contract is additive in both directions: an older watcher sends no classification, and an older portal triages a readiness-gate-… report as an ordinary incident.