A mesh node's identity is the pair (namespace, id). Its path is derived, not stored independently — in Postgres it is literally a generated column:

CREATE TABLE IF NOT EXISTS mesh_nodes (
    namespace       TEXT        NOT NULL DEFAULT '',
    id              TEXT        NOT NULL,
    path            TEXT        GENERATED ALWAYS AS (
                        CASE WHEN namespace = '' THEN id ELSE namespace || '/' || id END
                    ) STORED,
    ...
    PRIMARY KEY (namespace, id)
);

Two consequences follow, and together they are the whole of this page.

An id MAY contain a slash

There is no constraint anywhere forbidding it. Every mesh_nodes DDL — the three Postgres variants, the satellite-table script, mesh_node_history, and the SQLite adapter — declares plain id TEXT NOT NULL. A repo-wide sweep for a SQL CHECK constraint across both src/ trees finds none at all.

Slash-bearing ids are not an accident to be tolerated; several node families depend on them. Every LanguageModel node's id is the provider's wire idz-ai/glm-5.3, anthropic/claude-opus-5, openai/gpt-5.2 — because that string is what the provider's API expects and what the model is known by. The Postgres adapter states the rule in its own words (issue #2212):

🚨 THERE IS NO POSITIONAL (namespace, id) SPLIT OF A PATH — an id may contain '/'.

Splitting a path positionally is path-invariant and key-destroying

Because path is namespace || '/' || id, moving a slash from the id into the namespace leaves the path byte-identical while changing the primary key. These two rows have the same path and different identities:

namespace id generated path
Provider/OpenRouter z-ai/glm-5.3 Provider/OpenRouter/z-ai/glm-5.3
Provider/OpenRouter/z-ai glm-5.3 Provider/OpenRouter/z-ai/glm-5.3

That invariance is what makes the corruption invisible: every path-addressed read, every log line, every URL keeps reading the same. Nothing looks wrong.

The read/write asymmetry that turns it into data loss

The two sides of the adapter address rows differently, and both are correct in isolation:

So a write carrying a re-keyed (namespace, id) finds no conflict and INSERTs a second row whose generated path collides with the first. From that moment:

  1. A read WHERE path = $1 matches two rows and resolves an arbitrary one — not reliably the same one twice.
  2. The versions table (PK (namespace, id, version)) holds two independent chains, so a node with a long history can read back as having none.
  3. A delete WHERE path = $1 removes BOTH rows. One delete, and the node is gone.

What went wrong (#3894)

MeshOperations.SanitizeNodeId split a slash-bearing id at its LAST slash and moved the prefix into the namespace, on the stated premise that "the DB has a CHECK constraint blocking slashes in id". That premise was false, and had presumably always been false. Create and Update both ran through it, so no MCP write could address a flat-keyed slash-id node — every write minted a duplicate.

On the production portal this split Provider/OpenRouter + z-ai/glm-5.3 across two rows on 2026-09-09; the following morning the node was gone from the model list entirely. The MCP create tool's own documentation stated the same false rule ("id — the node's own slug, NO slashes"), so an agent following its instructions produced the duplicate by hand.

Both writes reported "did not land within the confirmation window" — a true negative: the confirmation reads the flat-keyed path while the write had gone to a different primary key.

Patch was never affected. It reads the existing node and writes existing with { … }, so it inherits that node's keying by construction — which is why the split could not be reproduced through patch alone.

The rules

The adjacent failure: a cached miss that outlives its invalidation

Worth knowing when a node "does not exist" right after you created it. A negative read is cached by the storm breaker inside MeshNodeStreamCache (_negative, keyed by path) — not by the per-node hub, and not by the unrelated MessageStormBreaker, which is a per-hub rate breaker with no path negative. An open window fast-fails reads and writes alike, and MeshOperations.Get reports it as Not found, indistinguishable from genuine absence and produced without ever reaching the owner.

A create at that path does clear it: the post-commit MeshChangeEvent.Created publish reaches MeshNodeStreamCache.OnMeshChangeResetFailureState(path), which drops the negative entry and evicts a faulted read entry. That behaviour is pinned by a test.

The remaining race was fixed in #3954. Every read/write that can conclude NotFound now claims the path before opening its owner round-trip. A change event revokes the current claim and clears only the negative entry belonging to that exact generation. Publication is a pair-exact compare-and-swap: an older probe cannot overwrite a newer probe's genuine miss, and its retraction cannot remove that newer entry either. Readers and writers also refuse an entry whose claim is no longer current, so a stale verdict cannot fast-fail even during the small interval before its owner retracts it.

Claims are not a second permanent path cache: a successful or transient probe removes its claim, and teardown removes a pending probe that never reached a terminal, while a genuine miss keeps one only for the lifetime of the existing negative entry. Natural re-probes replace the claim and keep the established exponential backoff; no timer, retry, or sweeper was added. This mirrors PathResolutionService._pendingFills: invalidation is authoritative over work that began in the older failure era.

See also

Reconnecting…
The server was updated. Reloading the page to pick up the latest version.