Agent admission — one rule for every lane

The rule (maintainer, 2026-10-03/04): "must generally use up the queue" · "if we have empty queue ⇒ utilize" · "we can configure relative load but not absolute number" · "it applies only when we have exhausted pool" · "we don't exhaust in one species and then block" · "slowly floods rather than sudden flooding … waits for seeds to grow and spread" · "if a lot is free ⇒ should add with high likelihood" · "we must be very tolerant and try to always find an alternative. however, must also provide possibility to hard pin model."

Every agent round in a process asks ONE admission before it starts: the dispatch pool (ThreadDispatchPool, src/MeshWeaver.AI/Supervision/). Its state is AdmissionState (immutable, pure transitions) and its rules are LaneScheduler (pure). The lanes that start work — the bug pool, the triage issue sweep, the review steward — pace themselves by the same admission's Allowance(lane), so none of them keeps an absolute cap of its own. There is no number of concurrent rounds anywhere in configuration: what is configured is relative (floors, weights) or a declaration (alternatives, pins).

The layers

1. Bulkheads — one per model pool

A bulkhead is the measured state of one model pool (the model a round runs on). Exhaustion is measured per bulkhead, so an exhausted Opus pool never holds back work on GLM or Kimi. A bulkhead carries its estimate K̂ (the concurrency it serves now), the carrying capacity K (the load at which it last said "too much"), when and why, Thompson's success/failure counts, and the mean round time and cost.

What says "too much" — reported by the engine where it happens:

Signal Where
A provider refusal for load, quota, credit or key (429, 402, insufficient credits, Key limit exceeded, a rejected key) ThreadExecution error path → ThreadDispatchPool.ExhaustionOf → ReportExhausted
A stream the stall guard abandoned; the round cap same
The provider's daily budget reached ThreadExecution budget gate → ReportExhausted
A review round slower than a review needs (33 tokens/s for 40,000 tokens in 20 minutes) ReviewAdmission.ApplyRelease → ReportExhausted

A clean round reports ReportClean; every round's price reports ReportCost.

2. Seed-and-grow — the estimate of one pool

Rule What it does
Seed A new pool starts at Seed = 2 rounds — a starting point, not a cap.
Slow start While no carrying capacity is remembered and demand exceeds a utilized estimate (at least half running), K̂ doubles every SlowStartEvery (15 s) — 2, 4, 8, 16, 32 within a minute. A 429 answers within seconds, so that is evidence enough, and an idle pool does not crawl. Each clean round that ended at or above the estimate adds one more.
Multiplicative decrease A refusal sets K = the load that ran and K̂ = Decrease × K (0.5) — once per burst (DecreaseHold, 30 s): thirteen 429s of one burst are one decrease.
Logistic growth Inside the exhaustion window, at most every GrowEvery (1 min), (ClearAfter, 35 min = a round's whole life) K̂ grows logistically toward K: K̂ += r·K̂·(1 − K̂/K), never past it.
Recovery Past the window K is forgotten and the pool probes upward again by doubling.
Spreading A pool first used because its neighbour was full starts from a quarter of the neighbour's estimate instead of the bare seed.
Memory The bulkheads are written to Admin/Threads and re-read on start: a restart does not forget what was measured, so a pool measured at 18 is used at 18 at once.

3. The probability of a marginal admission rises with the free fraction

A round beyond its lane's floor is admitted with probability p(φ) where φ = 1 − running/K̂: certain while at least 30 % is free (HighFree), falling smoothly to 0 as the pool fills, and 0 at or above the estimate. A light lane competes less eagerly for the last slots (the probability is raised to heaviest weight / its weight, between 1 and 4). Upward probing happens by growing the estimate — never by letting each waiting round gamble above it. Refused rounds re-ask on every release and on the pool's 15-second clock, so a refused draw is retried within seconds.

4. Lanes — guaranteed floors, borrowing, weighted fair share

Lane Who Floor (default) Weight
express Blocking work: a red main (a bug-fix thread whose id starts ci-) 15 % 0.15
babysitter The babysitter, the PR fixer, platform builds 15 % 0.15
reviews Pull-request review rounds 25 % 0.25
bugfix Bug threads and their fix rounds 25 % 0.25
triage Issues, feedback, incidents 10 % 0.10
interactive Everything else — a person's own chat unreserved (10 %) 0.10

A thread's lane is its Threads-app group (Lanes.Of, AI/ThreadGroups); LaneOfAgent on Admin/Threads overrides it per agent. Floors, weights, alternatives and pins are all live configuration on Admin/Threads (laneFloors, laneWeights, modelAlternatives, pins, laneOfAgent); the old maxConcurrentAgents is retired and only warned about.

In one pool, in this order (LaneScheduler.Decide):

  1. Within its floor — a lane with fewer running than its guaranteed slots (its floor of K̂, allocated by largest remainder among the lanes that HAVE work; a lane with waiting work and a non-zero floor gets at least one) is admitted while a slot is free. On a pool never measured exhausted it is admitted even when the pool is full: the estimate is then a conservative start, not a measured limit.
  2. Reserved — a borrower is refused while the free slots are owed to other lanes below their floors that have work waiting.
  3. Cost guard — a lane whose spend today ON THIS BULKHEAD'S PROVIDER reached its floor share of that provider's daily budget borrows nothing more there (it keeps its floor). Spend is kept per lane per provider, so a lane's spend on one provider never stops it borrowing on another; the budget is the limit the platform's budget gate last read for the provider today (AdmissionState.Budget) — with no reading today the guard does not apply.
  4. Fair share — the borrowed part goes to the waiting lane with the least borrowed per weight (deficit-round-robin in effect; ties to the higher-priority lane).
  5. Probability — §3.

Idle lanes lend everything: a single flooding lane fills the whole pool.

5. Reclaim at round boundaries

Every round asks again (a slot is released when the thread leaves its executing states), so every round is a boundary. A borrower is refused at its next boundary when a lane below its floor waits — nothing is killed, and the lane short of its floor starts within one round (LaneScheduler.MustYield names the borrowers that will yield).

6. Aging inside a lane

Inside a lane the oldest waiting round goes first: a newcomer does not take a free slot while older rounds of its lane wait for it. Where a lane chooses WHICH item to start (the bug pool), urgency is (1 + severity) × (1 + age / AgingUnit) (LaneScheduler.Urgency); the bug pool weights a red main 32, sev:B 16, H 8, M 4, L 2, unlabelled 1, with an aging unit of a day — so an eight-day-old low bug comes before a fresh blocking one, and nothing waits for ever.

7. Starvation signal

A lane below its floor with work waiting longer than StarvedAfter (15 min) — and every hard-pinned round waiting that long — is listed under starving on Admin/Threads. The stuck watchdog reads it as data.

8. Fallback to an alternative, and the hard pin

A round its own pool refuses moves to a declared equivalent model (modelAlternatives: same tier, route rules respected — declared, never guessed) that admits it now. Among those, Thompson sampling picks: each alternative's success rate is drawn from Beta(successes + 1, failures + 1), discounted by its mean cost; an unexplored alternative draws from Beta(1, 1), so the choice keeps exploring. The move is written onto the round's drained messages in the claim write, so the round runs on the alternative.

A hard pin (pins: a lane or an agent) disables the fallback: the round waits for its model, the wait says hard-pinned to … on the pool's census row, and past StarvedAfter it is a pinned starvation signal.

8b. Fallback is ON by default — default alternatives per tier

"Very tolerant — always find an alternative", a hard pin the exception (maintainer, 2026-10-04). With nothing configured, a model in one of the default tiers (LanePolicy.Equivalences) falls back to the other models of its tier that the deployment's catalog actually has (every LanguageModel under Provider/, read live by the supervisor — ThreadDispatchPool.SetCatalog):

Tier Members, in preference order
heavy claude-opus-5.5 ↔ claude-sonnet-5.5 ↔ gpt-6-sol
standard glm-5.3 (and its configured variants: -review, -coding, …) ↔ kimi-k2.7-code ↔ deepseek-v4-flash ↔ mistral-large
light gpt-6-luna ↔ mistral-small ↔ gemini-3.5-flash-lite

Order (LaneScheduler.DefaultAlternatives): the same model on a DIFFERENT provider first (a direct Anthropic or Azure Foundry model beside an OpenRouter route — a key's daily cap is per key, so another provider is what keeps the lane alive), then the tier's other models in token order, own provider before others. Never off the EU route: from an EU-routed provider only another EU route or a direct provider, never the non-EU router (LaneScheduler.MayMoveTo). A model in no tier has no defaults; an explicit modelAlternatives entry on Admin/Threads replaces its defaults. A hard pin still waits.

9. The provider's DAY — closed pools and the budget spent by priority

On 2026-10-04 at 14:47Z the control instance's provider key hit its daily limit (HTTP 403: Key limit exceeded (daily limit)), and every lane stopped at once — reviews, triage, bug fixes, embeddings. Two rules answer it:

Measured the same day: from 16:31Z to ~20:01Z every review round of every pull-request head on the control instance ended within a second on that 403 (this PR's own seven rounds included) — the reviewer is bound to one model on one key, with no declared alternative to fail over to. Default alternatives per tier, including another provider first, follow in the stacked change #2875.

Not wired yet: OpenRouter publishes a key's own limit and remaining credit (GET /api/v1/key: limit, limit_remaining, usage_daily); this reads the platform's own ledger against the provider node's DailyLimit instead. A key whose limit is set only at the provider is therefore learned from its first refusal (and closed at once), not ahead of it.

10. Delegated rounds

A sub-thread of a round that holds a slot ({parent}/…) is part of that round's work: it is admitted at once on its parent's lane — refusing it could only stall the parent (a parent holding a slot while waiting on a child that waits for a slot is a nested-resource deadlock). It does NOT ride the parent's slot: it is COUNTED as a running round of that lane, because the provider sees it, and it is marked (Seat.DelegatedBy; delegated per lane × bulkhead on Admin/Threads) so a fan-out is visible. Nothing in admission bounds the fan-out: a bound there would reintroduce the deadlock. A bound belongs where the delegation is made (the delegate call refused visibly to the parent), not here.

11. The read model and the dispatch entry point (for a coordinator)

Admin/Threads carries, per bulkhead: capacity (K̂), carryingCapacity, exhausted, exhaustedAt, why, slowStart, free, running, waiting, successes, failures, meanSeconds, meanCostUsd; per lane × bulkhead: running, waiting, guaranteed, borrowed, bound, oldestWaitingSince, spentTodayUsd (on that bulkhead's provider), delegated; per bulkhead also closedUntil, budgetSpentUsd, budgetLimitUsd, budgetAt; and starving (a closed provider first). A coordinator decides what to start and how much: ThreadDispatchPool.Plan(lane, model, wanted) answers how many of wanted rounds the admission would start now (floors, fair share, pins and alternatives honoured, nothing seated). Enforcement stays in the admission: each round still asks.

The audit — every cap the lanes had (2026-10-04)

Read on origin/main fd3dbd1e1. (a) = a throughput throttle, replaced by the shared admission; (b) = a safety bound, kept, with why.

Where Cap Value Class Now
src/MeshWeaver.AI/Supervision/ThreadDispatchPool.cs:34, ThreadSupervisorStatus.cs:21 (Admin/Threads) maxConcurrentAgents — agent rounds per process 50 (a) Retired. Admission is per measured pool (this page). A configured value is warned about.
Hosting/Deployment/Source/BugFixPool.cs:2083 MaxActionsPerPass — starts + re-drives per 10-min pass 20 (a) Removed. BugFixPool.BudgetOf(Allowance(bugfix)): every due action with headroom, the lane's allowance once exhausted.
Hosting/Deployment/Source/BugFixPool.cs:2088 MaxStartsPerPass 1 (a) — its reason was the per-variant bound reading a stale index Removed. A start now counts on its variant at once (PassBudget.Move), so several starts in one pass cannot read one slot as free twice.
Hosting/Deployment/Source/TriageIssueSweep.cs:73,91 Hosting:Triage:IssueSweep:MaxPerRun — items created/re-opened per hourly run 20 (a) Retired. A run acts on every actionable issue with headroom, on Allowance(triage) once exhausted; a configured value is warned about.
Hosting/Deployment/Source/ReviewAdmission.cs:77 review share of an exhausted pool 10 % relative already Now the reviews floor of the shared admission (25 %), tightened by the process's admission (ReviewAdmission.Combine); review rounds that measure exhaustion feed every lane's pool.
Hosting/Deployment/Source/BugFixPool.cs:69, BugFixProcess.cs:173 per-variant maxConcurrent 2 / 3 (a) Not changed here — removed by the bug-pool change in flight; it adopts BudgetOf/Allowance.
Hosting/Deployment/Source/PrFixer.cs:73 MaxFixesPerDay — fix COMMITS pushed fleet-wide per UTC day 20 (b) Kept: a blast-radius guard on automated pushes to other people's branches, not a model-pool throttle (no model call is saved by it).
Hosting/Deployment/Source/PrFixer.cs:81 MaxFixThreadsPerPass 1 (b) Kept: correctness — every fixer answers on one page list that a patch replaces; two answers before the claim would lose one.
Hosting/Deployment/Source/PrFixer.cs:70 MaxAttemptsPerHead 2 (b) Kept: once-per-head loop guard.
Hosting/Deployment/Source/PrBabysitter.cs:1512 MaxProposalsPerHandoff 20 (b) Kept: the size of ONE validator thread's task, not how many threads run.
Hosting/Deployment/Source/PrBabysitter.cs:2318,2324 MaxLogTails, MaxChangedFiles 2 / 60 (b) Kept: evidence size per proposal.
Hosting/Deployment/Source/BugFixPool.cs:1673,1676 MaxAdoptionsPerPass, MaxAdoptionReadsPerPass (pre-hand-over backlog) 1 / 5 (b) Kept: GitHub reads of a legacy backlog; a migration trickle, not agent throughput.
Hosting/Deployment/Source/PullRequestSweep.cs:59 MaxPerRun — kicks per run only while the slot ledger cannot be read 20 (b) Kept: the fail-safe of an unreadable ledger; with a readable ledger the kick budget is the admission's.
src/MeshWeaver.AI/Supervision/ThreadSupervisorStatus.cs MaxRetries 2, LookbackLimit 500 (b) Kept: relaunch loop guard; scan page size.
PullRequestIntake.MaxReviewAttempts, ReviewAdmission.MaxOutageRounds 3 / 6 (b) Kept: once-per-head loop guards.
Hosting/Build/Source/BuildQueueLogic.cs:24 Hosting:Builds:Concurrency — CI heavy legs admitted 3 different resource Not the model pool: GitHub runner capacity. Left as it is; a runner-pool admission is its own change.
Hosting/Queue/Source/QueueLogic.cs the job queue's measured Bound (AIMD, TargetUtilization 0.1) relative relative already Unchanged; its jobs' agent rounds go through this admission like every other round.

The Monte Carlo comparison — why seed-and-grow

src/MeshWeaver.AI.Test/AgentAdmissionSimulationTest.cs runs the admission that ran until 2026-10-04 (step: everything admitted until a refusal, then a tenth of the measured load for 35 minutes, one bound for everything) against seed-and-grow (the shipped AdmissionState) over the load measured on the control instance on 2026-10-02 (Hosting/ReviewRoundCapacity): 35 review heads in two minutes (40,000 tokens each), the bug backlog (100 rounds of 20,000), three red mains five minutes in, triage every four minutes, the babysitter every six. The provider model is fitted to the measured rates — 61 / (1 + (n − 1)/(0.53·K)) tokens/s per round with n in flight (K = 18: 61 alone, 39 at five, 20 at twenty); a round started above K is refused with a 429 within 15 s with probability min(0.9, (n − K)/10); a round unfinished at 30 min is cut; a round slower than its tokens / 1,200 s reports the pool exhausted. Ten seeded runs per row, means; the test asserts the relations below and prints the table, so a run reproduces it exactly.

Scenario Admission Drain (min) 429s Cut at 30 min Slow rounds Longest wait to start, by lane (min)
burst (2026-10-02) step-unlimited 600 (not drained in 10 h) 666 165 65 express 28, babysitter 33, reviews 62, bugfix 0, triage 63
burst (2026-10-02) seed-and-grow 272 13 17 20 express 6, babysitter 10, reviews 227, bugfix 260, triage 46
idle, new pool, 6 reviews step-unlimited 17 0 0 0 reviews 0
idle, new pool, 6 reviews seed-and-grow 17 0 0 0 reviews 1
idle, measured pool (K̂ 18), 12 reviews step-unlimited 24 0 0 12 reviews 0
idle, measured pool (K̂ 18), 12 reviews seed-and-grow 24 0 0 12 reviews 0
bug flood on a small pool (K 6), reviews on another (K 18) step-unlimited 232 112 10 37 reviews 58, bugfix 0
bug flood on a small pool (K 6), reviews on another (K 18) seed-and-grow 369 9 0 8 reviews 1, bugfix 364

The reviews of the last scenario drained in 28 min under seed-and-grow against 81 min under step (its one bound held them behind the flood).

What it shows:

What this does not establish

The tests — one per rule

Every test below fails when its rule is removed. Checked 2026-10-04 by mutating each rule in LaneScheduler / AdmissionState / ThreadDispatchPool.FailureOf (27 mutations: the floor, the overdraft, reclaim, fair share, the probability, slow start and its interval, logistic growth, the decrease and its once-per-burst hold, aging, the cost guard, starvation, spreading, Thompson, the pin, the fallback, oldest-first, bulkheads, the closure and its provider-wide reach, the budget reserve and its day, delegation, the midnight reset, rate-limit-is-load, closed-first) and running the three suites: every mutation turned at least one named test red.

Rule Test
Bulkheads LaneSchedulerTest.Bulkheads_AnExhaustedPoolNeverBlocksAnother, AdmissionStateTest.PerProviderExhaustion_Isolates
Fallback + Thompson LaneSchedulerTest.Fallback_ThompsonPrefersWhatWorksAndKeepsExploring, AdmissionStateTest.Fallback_ARefusedRoundMovesToAnAlternative
Hard pin LaneSchedulerTest.Pins_NameALaneOrAnAgent, AdmissionStateTest.HardPin_WaitsVisiblyAndSignalsTheWatchdog
Seed LaneSchedulerTest.SeedAndGrow_ANewPoolStartsFromTheSeed
Slow start LaneSchedulerTest.SeedAndGrow_SlowStartDoublesWhileDemandExceedsAndNothingRefuses
Logistic growth to K LaneSchedulerTest.SeedAndGrow_LogisticTowardTheCarryingCapacityInsideTheWindow
AIMD (once per burst) LaneSchedulerTest.SeedAndGrow_MultiplicativeDecreaseOncePerBurst_AdditiveIncreaseOnSuccess
Recovery LaneSchedulerTest.SeedAndGrow_RecoveryForgetsTheCarryingCapacityAndProbesUpward, AdmissionStateTest.Recovery_LiftsTheBound
Spreading LaneSchedulerTest.Spreading_ANeighbourSeedsFromAFullPool
Probability rising with free capacity LaneSchedulerTest.Probability_RisesWithTheFreeFraction
Floors LaneSchedulerTest.Floors_AreGuaranteedToLanesWithWork, LaneSchedulerTest.Floors_ALaneBelowItsFloorIsAdmitted, AdmissionStateTest.ExhaustedPool_EachLaneGetsItsRelativeShare
Floor overdraft on a never-exhausted pool AdmissionStateTest.Flood_ExpressAndReviewsStartAtOnce_OnAPoolNeverExhausted
Borrowing (idle lanes lend) LaneSchedulerTest.Borrowing_AFloodUsesEverythingWhileTheOtherLanesAreIdle, AdmissionStateTest.IdleLanes_AFloodUsesTheWholePool
Weighted fair share LaneSchedulerTest.FairShare_TheLaneBehindPerWeightBorrowsFirst
Reclaim at round boundaries LaneSchedulerTest.Reclaim_BorrowersYieldToALaneBelowItsFloor, AdmissionStateTest.Flood_ExpressAndReviewsStartWithinOneRound_OnAnExhaustedPool
Lane bound under exhaustion LaneSchedulerTest.LaneBound_NullUntilExhausted_ThenFloorPlusUnowed
Aging inside a lane LaneSchedulerTest.Aging_AnOldLowSeverityItemComesFirstEventually, AdmissionStateTest.OldestFirst_ANewcomerDoesNotPassAnOlderWaitingRound, BugFixPoolTests.TheQueueOrder_IsSeverityTimesAge
Daily cost guard LaneSchedulerTest.CostGuard_ALaneOverItsShareBorrowsNothingButKeepsItsFloor, AdmissionStateTest.Cost_CountsAgainstTheLanesDay
Starvation signal LaneSchedulerTest.Starvation_ALaneBelowItsFloorWaitingTooLongIsReported, AdmissionStateTest.HardPin_WaitsVisiblyAndSignalsTheWatchdog
A key's daily limit closes the provider; every lane fails over; pinned waits; top watchdog item AdmissionStateTest.KeyLimit_ClosesTheProvider_AndEveryLaneFailsOverInsteadOfStopping
Rate limit = load, key/credit = closed AdmissionStateTest.Failures_RateLimitIsLoad_KeyAndCreditClose_OurOwnFaultIsNothing
The day's budget spent by priority AdmissionStateTest.Budget_TheRemainderIsSpentByPriority
Delegated rounds AdmissionStateTest.Delegated_ASubThreadOfARunningRoundStartsOnAFullPool
Default alternatives per tier (catalog only, other provider first, EU route kept, explicit replaces) DefaultAlternativesTest.* (6 tests, control's catalog)
A closed key fails over by default; a pin still waits DefaultAlternativesTest.ByDefault_AClosedKeyFailsOverToAnotherProvider_AndAPinStillWaits
Lanes from thread groups LaneSchedulerTest.Lanes_FollowTheThreadGroup
Every ready round starts with headroom, per lane AdmissionStateTest.HealthyPool_EveryReadyRoundStarts (all six lanes), AdmissionStateTest.HealthyPool_EveryLaneAtOnce_EverythingStarts
A late signal finds its pool AdmissionStateTest.Signals_ALateSignalFindsItsPool
The process pool on a mesh ThreadSupervisorMeshTest.OnAPoolMeasuredExhaustedAtOne_TheSecondThreadWaitsQueued_ThenRunsWhenTheFirstSettles
Bug pool pass budget BugFixPoolTests.ThePassBudget_IsTheSharedAdmissionsAllowance
Triage sweep count retired TriageIssueSweepTests.TheCadence_IsConfiguration_AndClamped
Reviews: floor + the process bound ReviewAdmissionTests.WhenExhausted_TheBoundIsARelativeShareOfThePool, AHerdOfHeadsQueuesBehindTheReviewBoundTest
The Monte Carlo comparison AgentAdmissionSimulationTest.SeedAndGrow_BeatsStepUnlimited_OnTheMeasuredLoad

Where the code is

Related: Review round capacity · Queues · The activity execution model · Thread groups.