Two questions were raised during the 2026-08-31 CI overhaul and tracked as issues: "why is the chart gate on core?", and "can we mount a drive where git is checked out". Both are real, both are larger than they look, and neither should have stayed in an issue thread. This page is the state of record for both, measured against the code on 2026-09-01.

The chart lives here; the application it deploys does not

deploy/helm/ is in the platform repo, so Chart Gate and Chart Drift correctly run here β€” a gate lives beside its subject, and this one exists because of a measured shape-impossibility incident: the chart described KEDA-min-2, PDB-2 and replicas: 1 simultaneously for a month, and one pod meant multi-minute 503s.

The architectural drift is equally real. The portal hosts the chart deploys have moved to the plugins repository; deployments are node-type-driven through the hosting operator; environment folders live in the private deployments repository. So the platform ships and gates a chart for an application it no longer contains.

What that costs today, measured: Chart Gate triggers on pull_request with no paths: filter, so every core PR renders every values combination and asserts every config key is read by a template β€” including PRs that cannot touch the chart. Chart Drift is schedule-only and costs nothing per PR.

🚨 The relocation checklist β€” everything that must move in ONE change set

A chart is not a directory; it is a directory plus everyone who resolves a path into it. Measured references, each of which breaks silently if it is left behind:

Consumer What it holds
main-cd.yml renders WebhookInbox__Targets__N from deploy/helm/templates/memex-portal/config.yaml (the FrameworkBroadcast__Subscribers__N slots were retired 2026-09-03 β€” the subscriber set is the Hosting/Deployment records')
homebrew.yml path filters on deploy/helm/** β€” a chart change is what triggers a tap rebuild
deploy/aks/scripts/check-chart-invariants.sh, check-values-are-read.py the gate's own implementation
deploy/aks/envs/example/deploy.sh, secretproviderclass.yaml first-time environment setup
deploy/aks/infra/modules/portal-identity.bicep infrastructure that must agree with the chart's service account
the hosting operator CHART="${HOSTING_CHART:-/opt/hosting/chart}" β€” the operator carries its own baked copy at a pinned path inside its image
the self-updater documented against deploy/helm/templates/memex-portal/ for the SA, Role and RoleBinding

🚨 The bake gate is live and helm upgrade REVERTS. A relocation that moves the chart without moving the gate's arming leaves a window in which an upgrade silently undoes it.

⚠️ A drift already exists in this set, independent of any move, and is worth fixing whether or not the chart relocates: the self-updater and its RBAC target memex-migration-deployment, a workload the chart does not render β€” helm template deploy/helm emits a Job, not a Deployment. That is the #1788 shape (a command aimed at a Deployment that does not exist either errors or keeps a cluster-only orphan alive that re-runs the migration forever).

🚨 The cheap interim is NOT safe here β€” do not path-filter the chart jobs

The obvious economy is to filter the PR-triggered chart jobs to deploy/helm/**, on the reasoning that a chart which did not change cannot newly break.

Do not do this while Chart Gate is a required context. GitHub paints a skipped job the same colour as a passed one, and this repository has already established that a SKIPPED or absent required context counts as SATISFIED. So a path filter converts "the chart gate did not need to run" and "the chart gate did not run" into the same green tick β€” and the day someone changes a values file through a path the filter does not match, the gate is silently absent rather than red.

The reasoning that makes a filter seem safe β€” "it cannot newly break" β€” is also not quite true: the gate asserts a cross-file property (every config key a values file sets is read by a template). A change to a template under a different path can therefore break an unchanged values file, which is precisely the class the filter would stop watching.

If the per-PR cost must come down, the sound options are to make the gate cheaper, or to move it with its subject β€” not to make its absence indistinguishable from its success.

The self-hosted runner pool

GitHub-hosted runners cannot mount persistent drives, which is what the original request asked for. The approximation already in place is per-digest blob caching; what remains per job is a once-per-digest fill and a checkout of a few seconds.

🚨 The first half of this shipped on 2026-09-07, and it did NOT use a dedicated node pool. The measured headroom on the existing pool made the standing cost avoidable; isolation comes from a negative priority class, a ResourceQuota and an explicit reserve instead. What was deployed, the caps and the arithmetic behind them are the state of record in Self-hosted CI runners on AKS. The rest of the shape below β€” the git mirror, the warm layer store, the opt-in labels β€” is still outstanding.

The full shape, recorded so it can be picked up rather than re-derived:

Also agreed for the lane, and independent of the pool: after the global build, run tests β€” one job builds the whole graph in one workspace, and test jobs consume its artifacts.

Constraints that shape it: the cluster is private, so kubectl reaches it only through az aks command invoke; environment configuration lives in the private deployments repository; and cluster changes route through the hosting operator rather than direct access.

The honest trade. This converts a variable per-job cost into a fixed standing one: a node pool that exists whether or not CI is running, plus a controller to keep alive, plus the failure modes of a stateful mirror (a corrupted or diverged mirror fails every job at once, where a fresh checkout fails none). It is worth doing when CI volume makes the per-job seconds dominate β€” and it is worth saying out loud that until then, the caching already in place is most of the benefit for none of the standing cost.

See also

Reconnecting…
The server was updated. Reloading the page to pick up the latest version.