Instance Reboot

The rule (policy instance-reboot, register). An instance can be brought to a known-good, NEWEST state in ONE step. One durable request runs five steps in order, and each step reports its outcome on the request with a named reason when it is skipped or fails. A person's call is the signature, with no second approver, and the request records who made it. The instance's own watchdog may file the same request, without approval, but only on an explicit, conservative wedge predicate, at most once per interval, and every firing raises an alarm.

Why it exists

The control instance once spent a day unable to recover itself. Hosting needed MeshWeaver.AI 1.21, 1.20.4 stayed loaded, and every thread start threw MissingMethodException. Every remedy existed: a sync, a module landing, a restart, a newer image. But each was a separate act, some needed a second person, and the one instance able to run them was the one that was broken. A reboot is all of those acts as one request, executed by platform code that is shipped in the image, so it cannot be broken by the module it repairs.

The request

An InstanceReboot node lives at Admin/_Reboot/{id} and holds an InstanceRebootRequest (MeshWeaver.Graph.Configuration). It is written only by InstanceReboot.Request, as System, after the caller has authorised the requester. The executor acts only on a request whose createdBy is System. This is the same trust rule as Module Reload.

field meaning
reason, requestedBy, requestedAt why, and who: the person's user id or reboot-watchdog
trigger Person or Watchdog
wedged the instance was judged wedged. A watchdog reboot always is; a person may say so. The image step then does not wait for a dependent-suite verdict
wedgeEvidence, wedgeFingerprint for a watchdog reboot: the evidence that fired it, as one line, and its fingerprint
status Requested, then Preparing, then AwaitingRestart, ending in Done or Failed
steps[] Sync, Modules, Image, Restart, Verify. Each step is Pending, Running, Ok, Skipped (with the reason why) or Failed (with the reason why)
modules[] the module rows the landing produced: found, landed and target version, or a decline by name
runningImage, targetImage what ran before, and what the image step chose. A null target means a restart on the running image
restartRequestedAt stamped BEFORE the roll or restart is requested, so a resumed executor never asks twice
replicas{process} what each process that booted after the stamp reported from its own checks
failure, completedAt, log[] the red steps by name, when it ended, and the audit trail

The five steps

The executor is InstanceRebootExecutor, running on the request node's own hub, so exactly one process drives a request.

  1. Sync. Every GitSynced MODULE source is imported at its branch HEAD (ModuleSourceUpdate.UpdateAll). The executor makes the same call as git_hub_sync op=update (UpdateToLatestFromGitHub), once per source, as System, and follows each import's activity to its end. The seal is consulted exactly as it is for a person's update: a held source lands nothing and the hold is named. Each module is judged alone by its manifest hash. A floor above the running platform, or a dependency the loaded build does not meet, is declined by name (ModuleSyncDecision). A source counts as a module source when its config has recorded a module-bearing tree. Content-only spaces are not re-imported, and the step lists them as skipped, with the reason.

  2. Modules. The executor lands the newest COMPATIBLE published version of every installed module through ModuleReloadExecutor.ResolveAndLand, which is the module reload's own resolve-and-land and is not forked. Compatibility means the declared floor is at or below the running platform, never a seal. A floor above it is a decline, named on the step. A decline is not red, because it is the right answer; any other failure to land is red. Nothing is activated at this step, because the restart activates it.

  3. Image. The executor picks the newest platform image that this instance's own update policy admits (IInstanceRebootActivation.SelectImage, the self-updater's own selection: the registry's tags, the policy's channel and pattern, and the release-availability walk). A held release is never forced. The step then reads Skipped, names the hold, and the reboot restarts on the running image. A wedged reboot leaves the combo (dependent-suite) verdict out of the decision and says so on the step.

  4. Restart. The executor stamps restartRequestedAt, then makes ONE roll to the target, with the database migration first, or one restart on the running image when there is no target. This goes through the self-updater's one path: a self-patch, or a hand-over to the control lane (self-update-available or self-update-restart-pending). The roll floor does not defer it. If no roll or restart could be scheduled, the reboot is Failed, with the step naming why. Two refusals come before anything is issued:

    • a disruptive rollout strategy on the portal Deployment;
    • a control instance that cannot self-patch, which would otherwise hand its restart to itself.

    Both are described under "Operating lessons" below.

  5. Verify. Every process that booted after the stamp runs the registered checks (InstanceRebootAgent, IInstanceRebootCheck) and reports under its own key:

    • health:nodetype_bake waits for this process's bake to settle and is red when any CRITICAL NodeType (Reboot:CriticalNamespaces, default Hosting) regressed, is content-broken, failed with no baseline, or was not evaluated. Failures elsewhere are named, not red.
    • health:pending_module_activation is red while a landed module still waits for activation, or the state is undetermined.
    • health:content-types is red when a NodeType's content degraded to untyped on this process.
    • health:singletons-resumed is red when a configured singleton (the PR babysitter, the PR review sweep) has not stamped a pass newer than the restart within its budget.
    • smoke:thread-start (registered by the AI module in MeshWeaver.Plugins) starts a thread and is red when the start throws.

    A counted process that measured NO smoke check makes the verdict red. Without that rule, the incident's own failure shape would pass unmeasured. A health reading that is not registered on the host is NotMeasured: it is named in the verdict and counted neither green nor red.

A red step before the restart does NOT stop the reboot, because bringing the instance back is the point. It does make the request Failed, naming every red step. Only the restart failing ends the reboot early, because nothing would boot to verify.

Who can start one

surface authorisation how
MCP reboot_instance (MeshWeaver.Plugins McpMeshPlugin) IsGlobalAdmin MeshOperations.RebootInstance(reason, wedged). Returns {status, path} at once; read the node for the outcome
the Reboot button on a deployment's page (MeshWeaver.Plugins Hosting/Deployment) IsGlobalAdmin; the click is the signature files a Hosting/InstanceAction of kind Reboot. For this instance, the action files the request in-process; for another instance, it goes through the control lane
the control lane (ControlLaneOperation.Reboot, RebootOperation) the control instance signs with the deployment's own key the plan is fixed, with ONE step (RebootOperation.PlanFor), so its digest is bound without a dry run, and the requester is the approver. The lane run ends when the request is FILED on the target, and the reboot then reports on its own node there
the instance's own watchdog none; rate-limited and alarmed see below

Self-trigger: the watchdog

RebootWatchdog, armed on every process's mesh hub, evaluates once per Reboot:WatchdogInterval (default 1 minute). The predicate is explicit and pure (RebootWatchdogRules.Evaluate), and it holds when EITHER leg holds:

The rate limit is RebootWatchdogRules.Admit. It refuses to fire in three cases:

Every firing and every refused firing logs a Critical line, [RebootWatchdog] …, naming the evidence. Refused firings are logged once for each distinct evidence and reason per process. Setting Reboot:WatchdogEnabled to false turns the self-trigger off.

Configuration (InstanceRebootOptions)

key default
CriticalNamespaces Hosting
BakeSettleBudget / CheckBudget 15 min / 5 min
WatchdogEnabled / WatchdogInterval true / 1 min
ThreadStartFaultThreshold / ThreadStartWindow 3 / 15 min
CompileErrorFor 20 min
WatchdogMinInterval / WatchdogRepeatWindow 6 h / 24 h

Operating lessons the reboot must respect

These lessons come from the control-instance recovery attempts on the day this was built:

What is NOT established