The Portal Heap Is Hubs
A memex-cloud portal replica grows ~215 MB/h, monotonically, through hundreds of full
collections, and never returns to an earlier floor. This page is the heap-dump evidence for
what is actually in that memory, taken read-only from a live production pod on
2026-09-04. It exists because the answer contradicts every hypothesis the investigation was
carrying, and because the measurement was expensive to assemble.
One sentence: the heap is 9 386 MessageHub instances β 89.6 % of them sync/{clientId}
stream hubs, of which 1 496 are fully disposed corpses held by the very streams that created
them β and it is large because each hub carries its own copy of the framework's
dependency-injection and type-registry metadata; not because of assembly load contexts, not
because of Roslyn, and not because of GC fragmentation.
The specimen
memex-cloud/memex-portal-deployment-5b9cbf46dd-gs6j6, 26 h old, working set 6 859 MiB,
container limit 16 Gi. A dotnet-dump collect --type Heap through an ephemeral
--profile=sysadmin container (the recipe is in
Debugging Disposal, Storms and Leaks), then
dumpheap -stat. The dump file was deleted from the pod afterwards; no cluster object was
mutated.
π¨ That last sentence understates the cost, and a later measurement proved it. On a replica
this size --type Heap freezes the process for ~106 s, which is longer than the liveness budget,
so taking one restarts the container β see The dump is not free below. No cluster object
is mutated; the replica is.
Total 49,827,806 objects, 5,544,881,838 bytes
Free 411,422 2,031,989,944
β3.51 GB live, 2.03 GB free. The heap is not mostly garbage awaiting a collection.
What is in it
The exact hub count, from the dump:
MT Count TotalSize Type
7f9d080a7478 9,386 2,928,432 MeshWeaver.Messaging.MessageHub
7f9d0953b320 962 123,136 MeshWeaver.Hosting.Orleans.MessageHubGrain
9 386 live hubs against 962 grain activations. dumpobj segfaults SOS on this dump, so the
addresses were read with a 40-line ClrMD walk instead (MessageHub β Configuration β Address
β Segments[0]), run inside the same ephemeral container against a second dump taken 22 minutes
later (22:32:46Z β 22:55:16Z):
TOTAL MessageHub: 9393
--- by address type segment ---
8419 sync
184 MeshWeaver
144 Ifrs17
67 Crm
61 Underwriting
47 Planning
26 ReinsuranceDemo
β¦ (the tail is partition names β node hubs)
--- sample full addresses ---
sync/oeTOuFSg1kSZgDMysMaQ4Q
sync/QP8kKCTa2k-bgSYV8-5moQ
Plugins/OpenAI
π¨ 8 419 of 9 393 hubs β 89.6 % β are sync/{clientId}. The ~974 remaining are node hubs
spread across partitions, which matches the 962 MessageHubGrain activations almost exactly. So
the grain-hosted working set is not what grows: each node hub carries ~8.6 synchronization
sub-hubs, and those are the heap. At the β390 KB a hub retains (below), 8 419 of them is
β3.2 GB of the 3.51 GB live heap.
β οΈ One honest caveat from the pair of dumps: creation is bursty, so do not read this curve as purely uptime-driven. 9 386 β 9 393 is +7 hubs in 22 minutes (β19/h) β while the pod's lifetime average is 9 386 over 26 h β 361/h. Both dumps were taken at ~22:40 UTC, off-hours. So hubs arrive with load and then never leave; the memory curve looks uptime-driven only because the leaving half is missing. That distinction matters for the fix: nothing here is a background timer minting hubs on an idle pod.
The MessageHub objects themselves are 2.9 MB. Everything below is what they root:
| Type | Count | Bytes | Per hub |
|---|---|---|---|
Autofac.Core.Resolving.Pipeline.MiddlewareDeclaration |
10,001,898 | 480 MB | 1,066 |
Autofac.Core.Registration.ExternalComponentRegistration |
2,455,373 | 275 MB | 262 |
Autofac.Core.Resolving.Pipeline.ResolvePipelineBuilder |
2,534,453 | 101 MB | 270 |
Autofac.Core.Service[] |
2,534,378 | 81 MB | 270 |
ConcurrentQueueSegment<IComponentRegistration>+Slot[] |
56,330 | 83 MB | β |
ExternalComponentRegistration+NoOpActivator |
2,455,373 | 59 MB | 262 |
Autofac.Core.IComponentRegistration[] |
278,690 | 35 MB | 30 |
System.Action<ResolveRequestContext> |
508,761 | 33 MB | 54 |
Autofac.Core.Registration.ServiceRegistrationInfo |
255,878 | 24 MB | 27 |
List<Autofac.Core.IComponentRegistration> |
434,807 | 14 MB | 46 |
| Autofac subtotal | β1.17 GB | β128 KB | |
MeshWeaver.Messaging.Serialization.TypeDefinition |
848,844 | 75 MB | 90 |
ConcurrentDictionary<String,TypeDefinition>+Node |
1,261,018 | 60 MB | 134 |
System.Func<String> / System.Func<KeyFunction> |
1,695,786 | 108 MB | 181 |
System.LazyHelper |
1,696,944 | 54 MB | 181 |
Lazy<String> / Lazy<KeyFunction> |
1,695,767 | 68 MB | 181 |
TypeDefinition+<>c__DisplayClass2_0 |
846,918 | 27 MB | 90 |
ConcurrentDictionary<String,TypeDefinition>+VolatileNode[] |
18,860 | 20 MB | β |
| TypeRegistry subtotal | β0.41 GB | β45 KB |
β1.58 GB of a 3.51 GB live heap β 45% β is framework metadata duplicated once per hub. 262 component registrations and ~90 type definitions, identical in content, held 9 386 times. Nothing in that table is application data. Averaged over the whole live heap a hub retains β390 KB, of which β173 KB is this duplicated metadata.
The rest of the live heap is the same shape one level up: LayoutAreaDefinition 61,610 objects
(7.9 MB) and ObservableRenderer 65,279 (4.2 MB) against 2,473 LayoutDefinition β the layout
catalogue, re-materialised per hub.
Why the LIVE ones are never released
HostedHubsCollection.messageHubs β src/MeshWeaver.Messaging.Hub/HostedHubsCollection.cs:29:
private readonly ConcurrentDictionary<Address, IMessageHub> messageHubs = new(AddressComparer.Instance);
It has no idle sweep, no TTL, no size cap and no LRU. The one removal is on the hub's own
disposal (:203):
hub.RegisterForDisposal(h => messageHubs.TryRemove(h.Address, out _));
So the chain is: an entry leaves iff the hub disposes; a grain-hosted node hub disposes iff
MessageHubGrain.OnDeactivateAsync runs (src/MeshWeaver.Hosting.Orleans/MessageHubGrain.cs:890-892
β hub.CancelCurrentExecution(); hub.Dispose();); and that method logged its own first line
(:864, LogInformation, *"Grain deactivating: reason="*) **zero times across five replicas and 514 687 log lines in 28 hours**, in a window carrying 2 791 other info:` lines.
Deactivation never happens because the activation is re-pinned faster than it can age out.
MessageHubGrain.cs:560 installs
.Set(new GrainKeepAliveCallback(() => TryDelayDeactivation(TimeSpan.FromMinutes(10))))
and every live sync stream fires a HeartBeatEvent at it every 45 seconds
(src/MeshWeaver.Data/Serialization/SyncStreamOptions.cs:18 β
JsonSynchronizationStream.cs:1062 β MeshExtensions.cs:537 callback.KeepAlive()). A 45 s
heartbeat against a 10-minute window is a 13Γ over-provision, and because the heartbeat is
itself a routed grain call it also resets Orleans' own idle clock independently. Orleans'
defaults (CollectionAgeLimit 2 h) are never configured anywhere in src/ or deploy/ and are
never reached.
The dump shows what is doing the heart-beating: 7 396 live synchronization streams
(SynchronizationStream<MeshNode> 4,205 + SynchronizationStream<EntityStore> 3,191) against
976 MeshNodeStreamCache+Entry. The stream cache's idle sweep works and is bounded; the
streams that outlive it are not.
π¨ The heartbeat is evidence that a transport is alive, not that anyone wants the data. It is today the only input to the lifetime decision, and it can only ever vote "keep".
β¦but that is only three quarters of the population
A third dump (23:03Z, same replica) read MessageHub.runLevel alongside the address. The
population is not homogeneous:
TOTAL MessageHub: 9398
6925 sync RunLevel=1Started
1495 sync RunLevel=6Dead
974 <node-or-other> RunLevel=1Started
4 sync RunLevel=0Starting
Dead is the enum's terminal state β "the hub is fully disposed and inert"
(MessageHubRunLevel.cs:22). 1 495 sync/ hubs have completed their entire shutdown and are
still reachable: β580 MB of corpses at β390 KB each. Unlike the Started population, that
admits no benign reading at all.
Who holds a corpse
A fourth dump (23:06Z) and one full pass over all 51 562 911 objects, recording every referrer
of a Dead hub:
dead hubs = 1496
--- referrers of DEAD sync hubs ---
19448 MessageHub+<>c__DisplayClass188_0 β«
13464 MessageHub+<>c__DisplayClass59_0 βͺ the hub's OWN internals β
1496 MessageHubConfiguration β¬ exactly 1496 of each,
1496 MessageService βͺ they die with it
1496 HierarchicalRouting / RouteConfiguration β
1485 MeshWeaver.Data.Serialization.SynchronizationStream<MeshWeaver.Mesh.MeshNode>
11 MeshWeaver.Data.Serialization.SynchronizationStream<System.Text.Json.JsonElement>
1 ConcurrentDictionary<Address, IMessageHub>+Node
--- GC roots pointing directly at a dead hub ---
direct-root hits = 0
Two readings, both decisive, and the first one corrects this page's own earlier section:
π¨ HostedHubsCollection.messageHubs is not what holds the dead ones β it did its job.
Exactly one of the 1 496 is still in a ConcurrentDictionary<Address, IMessageHub>. The
RegisterForDisposal removal at :203 fired for the other 1 495. The registry is where the
live hubs sit (6 925 sync + 974 node, none of which anything ever asked to dispose); it is
not the corpse-holder.
π¨ The corpse-holder is SynchronizationStream<T>.Hub. 1 485 + 11 = 1 496 β one stream per
dead hub, exactly. No GC root points at a dead hub directly (direct-root hits = 0); every one
is alive purely transitively, through the stream that created it.
And the stream was never told
A fifth dump (23:11Z) read isDisposed off every stream:
SynchronizationStream total=8461 disposed=11
4219 <MeshWeaver.Mesh.MeshNode> live
3193 <MeshWeaver.Data.EntityStore> live
992 <MeshWeaver.Data.InstanceCollection> live
11 <System.Text.Json.JsonElement> DISPOSED
Only 11 of 8 461 streams are disposed. SynchronizationStream.Dispose() sets isDisposed
before it disposes the hub, so a stream reading false never ran Dispose(). Put the two
measurements together:
π¨ β1 485
SynchronizationStream<MeshNode>instances are NOT disposed, and theirHubisDead. The hub was destroyed out from under a stream that still believes it is alive.
That is a defect with a name. A sync/ hub is a hosted hub: its parent disposes it during
DisposeHubsReactive when the parent goes down β a Blazor circuit ending, a DisposeRequest, a
recycle β while the stream that created it is owned by a workspace somewhere else entirely and is
never notified. The stream keeps isDisposed == false, keeps a strong reference to the corpse and
everything the corpse roots (its Autofac ILifetimeScope, its TypeRegistry), and keeps being
handed out as usable.
The codebase already knows the creation-time half of this hazard β SynchronizationStream.cs
refuses to hand out a stream whose host is winding down, because "a consumer cannot even detect
the corpse: ISynchronizationStream exposes no liveness/disposal member", and that cost a
production NRE. The post-creation half β the hub dying under a stream already handed out β has
the same root and is not covered.
What this falsifies
Three explanations were live before the dump. All three are refused by measurement.
Collectible AssemblyLoadContexts accumulating because nothing deactivates. The reasoning was
sound β a lifetime lease is taken per instance hub
(src/MeshWeaver.Graph/MeshDataSource.cs:1016-1036) and NodeAssemblyLoadContext.Dispose defers
Unload() while any lease is held. But the count is not there:
dotnet.assembly.count ({assembly}) 451 (gs6j6, 6.9 GB)
dotnet.assembly.count ({assembly}) 475 (5nqbz, 14.0 GB)
450 distinct .dll mappings in /proc/1/maps. The fattest replica carries 24 more assemblies
than the thin one, not thousands more. Whatever the ALC lease costs, it is not this heap.
Lazy first-compile from prebuilt-bundle coverage failure. Same refutation, same numbers β 368
uncovered NodeTypes cannot be a 10 GB retainer at 451 loaded assemblies. It remains a real CPU
cost and Modules:RequirePrebuilt remains the right lane for it; it is not the memory.
GC fragmentation, or committed memory the runtime declines to return. Plausible on the thin
replica (gen2 fragmentation 537 MB, LOH 1.27 GB) β and decisively false on the fat one. From
5nqbz at 14.0 GB:
dotnet.gc.last_collection.heap.size (By)[gen2] 10,251,252,048
dotnet.gc.last_collection.heap.fragmentation.size[gen2] 1,075,128
dotnet.gc.last_collection.memory.committed_size 12,391,981,056
10.25 GB of gen2 with 1.07 MB of free space in it β 0.01 %. That is live, reachable data sitting at the 12 GiB container-derived hard limit. It answers the question the issue asked head-on: 267 full collections free nothing because there is nothing to free.
Why it presents as CPU starvation rather than a restart
Nothing sets DOTNET_GCHeapHardLimit[Percent], so .NET's container default gives the managed
heap 75 % of the 16 Gi limit β 12 GiB. The runtime defends that ceiling with back-to-back
blocking gen2 rather than letting the kubelet restart the container: 0 OOMKills across 7 replicas
over 28 h against a GC pause share of 0.66 sustained for 2 h 15, with /alive answering
Healthy throughout. A memory shortage is something Kubernetes fixes in one second; CPU
starvation inside a process that still answers TCP is something it has no opinion about.
Three defects, not one
They are independent, and conflating them is how this stayed open.
A hosted hub can die under a stream that is never told. 1 496
Deadhubs, each held by an undisposedSynchronizationStream<T>.Hubβ β580 MB of pure corpse, and the only one of the three that is unambiguous.ISynchronizationStreamexposes no liveness member, so neither the stream nor its consumers can even detect the state.sync/hubs that are stillStartedare never released β 6 925 of them, ~7 per node hub. The same objects supply the 45 s heartbeat that pins their owning grain, so this defect is both memory and the reason nothing deactivates: the lifetime decision has exactly one input, and it can only ever vote "keep". Whether all 6 925 are garbage is the open question below.A hub costs ~390 KB of retained heap, ~173 KB of it duplicated framework metadata. A constant factor, but a large one: even a legitimate 962-node working set costs 962 Γ 9.75 Γ 390 KB β 3.5 GB, which does not fit a 12 GiB ceiling with room to serve. Each child
ILifetimeScopewraps 262 parent registrations inExternalComponentRegistration, each with its own ~4-stage resolve pipeline. Fixing (1) and (2) alone leaves a portal whose working set is set by its DI cost per actor.
π¨ None of the three is fixed by raising the memory limit, lowering the GC hard limit, recycling pods on a timer, or deleting pods by hand. The pod-deletion stopgap has been applied at least four times and is recorded as a stopgap.
What is still open, and the next measurement
Five dumps settle what is retained, which hubs, and who holds the dead ones. Two things are still inference rather than measurement, and both are named here so the next reader does not have to re-derive the whole chain.
1. What kills the 1 496 sync hubs. Measured: they are Dead while their stream is not
disposed, so the kill did not come through SynchronizationStream.Dispose(). Inferred: it came
from the parent, via HostedHubsCollection.DisposeHubsReactive β a circuit ending, a
DisposeRequest, a recycle. The confirming measurement is a log correlation, not a dump: count
[DISPOSE-CONTAINER] {Address}: lifetime scope closed (HostedHubsCollection.cs:338, Debug) for
sync/ addresses over a window and compare it with the growth in the Dead population. If they
match, the parent-teardown path is confirmed and the fix belongs where a hosted hub's death is
announced to whoever created it.
2. Whether the 6 925 Started sync hubs are garbage at all. They are in the registry, nothing
asked them to dispose, and their streams are live. That is consistent with correct behaviour β
8 450 undisposed streams for subscribers that really are attached. It is also consistent with
streams abandoned without Dispose(). The discriminator is a referrer walk on the Started
population's streams (the same one-pass ClrMD technique used above, retargeted): if they are
held by Workspace._localStreamCache / _remoteStreamCache entries whose subscribers are gone,
they are garbage; if they are held by live circuits, the portal's working set is simply larger
than its ceiling and the per-hub cost below is the whole story.
π¨ Two things block that walk today, and both are about WHICH replica you point it at (measured 2026-09-06, Systemorph/MeshWeaver#3432):
- No deployment carries the
Dead-half fix yet. #3427 merged asf41f8bdaat 15:49Z; CD run 7905 was that commit and was cancelled, so the earliest published image carrying it is3.0.0-ci.7906(71c194ca9β confirmed an ancestor-of check plus the tag present in ACR). Bothmemexandmemex-cloudrun3.0.0-rc9.ci.7693, 213 CD runs older, because prod was rolled back by hand that evening. π¨ Mind the two tag lines when reading those side by side: the running tag carries anrc9.segment that current CD tags no longer use, so compare the CD run NUMBER, not the string. A dump taken now therefore measures the pre-step-3 population and cannot answer "did theStartedcount move", which is the first row of this section's outcome table. - Dumping a replica fat enough to carry the population restarts it β see The dump is not free below. The walk itself is cheap and correct; getting a dump to run it against is the expensive half.
π¨ Do not skip to a fix on the strength of the corpse count alone. 1 496 dead hubs is β580 MB
β real, unambiguous, and worth fixing β but it is one sixth of the 9 398. Fixing only that
leaves ~2.7 GB in the Started population untouched, and (2) is what decides whether that
remainder is a leak or a sizing problem.
A third reading worth having: the growth rate of MessageHub count against replica age, on three
replicas of different ages. #3321 observed that working set tracks uptime rather than load β the
near-idle fxbhd holds 11 443 MiB on 512 log lines per 30 min while the busiest replica holds
4 134 MiB. The +7-in-22-minutes reading above refines that rather than contradicting it: hubs
arrive with load and never leave, so an idle pod holds its history without adding to it.
The dump is not free β it costs a container restart at this heap size
Measured 2026-09-06 on memex-cloud/memex-portal-deployment-58c859d4cc-gqqg7 (8 h old, working
set 6 609 MiB, image 3.0.0-rc9.ci.7693), while attempting the referrer walk this page's What is
still open section asks for.
21:37:38Z dotnet-dump collect -p 1 --type Heap (process suspended)
21:39:17Z kubelet: Container memex-portal failed liveness probe, will be restarted
21:39:26Z Complete (process resumed β 106 s of freeze)
21:39:35Z ephemeral container SIGKILLed with the restart
21:42:18Z pod reads 1/1 Running, RESTARTS 1
The probes, read off the live deployment, are what turns a long freeze into a restart:
| probe | path | period | timeout | failureThreshold | a freeze longer than |
|---|---|---|---|---|---|
| readiness | /ready |
10 s | 5 s | 3 | β30 s pulls the replica out of the Service |
| liveness | /alive |
15 s | 10 s | 6 | β90 s restarts the container |
π¨ The freeze is not write-bound, so it does not shrink by finding a faster disk. dd oflag=direct on the same node measured 169 MB/s (3.1 GB in 18.6 s). This dump's own file size
was never read β the container restart wiped it before it could be β but at the 5.54 GB the
2026-09-04 specimen produced, the write is only β33 s of the 106. The rest is createdump walking
the object graph β β50 M objects β with every thread suspended, so the freeze scales with the
very thing that makes this heap interesting.
Hence the tension that governs every future measurement on this page's open questions: the population only exists on a replica fat enough that dumping it restarts the replica. A young replica dumps in seconds and has nothing to show β a freshly restarted portal here reached 2 977 MiB within five minutes, so ~3 GiB is the floor, not the accumulation.
So do not treat "read-only, dump deleted afterwards" as meaning "no impact". Either:
- Spend the restart on purpose. The other four replicas stayed
1/1 RunningwithRESTARTS 0throughout, so the deployment served from four of five for β2 min; the cost is one replica's accumulated state, which is also the thing being measured. That can be a fair trade β but it is a decision to take deliberately and in advance, not a side effect to discover afterwards. (Whether any in-flight request was actually dropped was not measured.) - Or use the counters below, which read EventCounters over the diagnostic socket and never suspend the process. They cannot answer a referrer question, but they discriminate retention from fragmentation, which is what most of this page's falsifications turned on.
π¨ And note WHY the cheap option is so weak: the platform publishes no metric of its own.
src/ contains zero System.Diagnostics.Metrics meters β grep -rn "new Meter(" src is
empty β so every counter this page relies on (dotnet.gc.*, dotnet.assembly.count,
dotnet.timer.count) is a runtime built-in. Nothing anywhere reports live MessageHub count,
the sync/-versus-node split, the RunLevel histogram, or SynchronizationStream
total-vs-disposed. That is exactly why dotnet.timer.count appears below as a proxy for live
hubs, and why answering any question on this page currently costs a replica restart. The three
numbers this whole investigation turns on are cheap to compute in-process and are simply not
emitted β HostedHubsCollection.Hubs.Count() already exists and is used only in two log lines
(MessageHub.cs:1879, :2487). An observable-gauge trio over hub count by address type and run
level would make the population trackable per replica, continuously, at no risk β and would turn
"take a dump and hope" into "watch the curve and dump only when it says something changed".
Reading the counters yourself
The whole discrimination above is four counters, and they need no dump β dotnet-counters collect -p 1 --format csv System.Runtime inside an ephemeral --profile=sysadmin container with
TMPDIR=/proc/1/root/tmp:
| Counter | What it settles |
|---|---|
dotnet.gc.last_collection.heap.size[gen2] |
how much survives a full collection |
dotnet.gc.last_collection.heap.fragmentation.size[gen2] |
retention vs. fragmentation β the whole question |
dotnet.assembly.count |
whether ALCs are accumulating (they are not) |
dotnet.timer.count |
a cheap proxy for live hubs: 11,387 at 6.9 GB, 37,796 at 14.0 GB |
π¨ dotnet-gcdump is the wrong tool here and will mislead you. Its walk stopped at exactly
10 000 000 objects on this process and reported a 568 MB heap β a sixth of the truth β with a type
table skewed to whatever it reached first. The Total β¦ objects line from dumpheap -stat is the
number to trust.
π¨ And SOS dumpobj / gcroot segfault on this process's dumps, so anything that needs to
follow a field has to be ClrMD. That is not a hardship: the SDK image the debug container runs
has full internet, so dotnet new console + dotnet add package Microsoft.Diagnostics.Runtime
dotnet run -- <dump>works inside the pod, start to finish, in about ten seconds. The walk that produced the address histogram above is forty lines (MessageHubβ<Configuration>k__BackingFieldβ<Address>k__BackingFieldβ<Segments>k__BackingField[0]), and reading fields by name-substring offClrType.Fieldskeeps it robust against the compiler's backing-field spelling.
Four practicalities, all verified 2026-09-06 against Microsoft.Diagnostics.Runtime 4.0.732401:
- Pass the DAC explicitly.
DataTarget.LoadDump(path)records the runtime path as seen by the target, which does not exist in the SDK sidecar. Find it withfind /proc/1/root/usr/share/dotnet -name libmscordaccore.so(it wasβ¦/Microsoft.NETCore.App/10.0.11/) and pass it toCreateRuntime(dacPath, ignoreMismatch: true). - Compile the walk locally first.
dotnet new console+ the same package pins the identical version, so a localdotnet build -c Releasecatches a typo before it costs a dump β and, given the section above, a dump can cost a replica. - Get the analyzer into the sidecar with
kubectl cp, not a base64 heredoc:az aks command invoke --file Program.cs --command "kubectl cp Program.cs <ns>/<pod>:/hw/Program.cs -c <container>". Create the sidecar with-- sleep 7200and drive it with individualkubectl execcalls, so a failed step costs one command rather than the whole run. - The sidecar dies with the target. It is an ephemeral container in the same pod: if the dump trips the liveness probe, the sidecar is SIGKILLed (exit 137) along with the restart, taking the dump file and any un-flushed output with it. Print results as you go; do not batch them to the end.
The fields the open questions above need, beyond address and runLevel: SynchronizationStream's
isDisposed and Hub; Address.Host (which names the parent hub, so a Started sub-hub under a
Dead parent is separable from one under a live parent); and the subscriber count behind
SynchronizationStream.Store, reachable as ReplaySubject β _implementation β _observers β _data[]
β the discriminator for "is anyone actually listening to this stream".
Related
- Debugging Disposal, Storms and Leaks β the profiling recipe and the ClrMD root-chasing method.
- Hub Disposal Model β what disposal means for a hub, and the RunLevel phases.
- Mesh Node Stream Cache β the idle sweep that is bounded, and its "never release an entry with a live subscriber" contract.
- Dead Circuit Fan-Out Storm β the sibling failure where a departed subscriber keeps costing.
- Actor Model β why there is one hub per addressable thing in the first place.