Model labels and the two-stage Auto router
Maintainer, 2026-10-01: "auto should choose which model is appropriate. let's also calibrate this process that we have labels like coding where we can filter available models" · "auto must triage first the kind of task, then which model to take."
The three pieces of data
| Where | Field | Example | Meaning |
|---|---|---|---|
Model node (nodeType:LanguageModel) |
labels |
["coding", "review"] |
What the model is good for. Free-form, case-insensitive. Its tier counts as an implicit label. |
Agent (.md front matter) |
modelLabels |
modelLabels: [coding] |
The KIND of work the agent does. Pins Auto's stage 1. |
| Model selection | auto:<labels> |
auto:coding |
Auto, with the kind pinned. /model auto coding in the composer writes it. |
The kind vocabulary is data, not an enum: Ai:Router:Kinds (comma-separated), default
coding, triage, review, reasoning, chat, utility. An unknown label never drops a model or an agent;
it simply matches nothing (ModelLabels, the same rule as ModelDefinition.Tier).
The flow (one bounded routing pass, AgentChatClient.ApplyAutoRouterSelectionAsync)
It runs when the round's model is Auto (auto, Provider/Auto/auto, or auto:<labels>), or when the
serving agent declares modelLabels and no model was picked.
- Floor. The synchronous dispatch has already parked the round on a usable model: the agent's tier, or the deployment default. Every failure below keeps it. A residency-bound floor is never routed away from.
- Stage 1: kind and difficulty, on a CHEAP fixed model. The classifier is the
utilitytier (elsechat), never the floor, which may be the most expensive model. It answers{"kind", "difficulty", "reason"}. Difficulty is one ofsimple,standard,hard,hardest. With pinned labels the kind is fixed and only the difficulty is assessed. - Filter. The candidates are the loaded, non-router models that pass the caller's admission gate (usable credentials AND the instance's and agent's data residency, so EU-only holds by construction) and carry ALL the required labels (the pinned ones, else the classified kind). They are capped at 12. If no candidate is left, the floor is kept and that decision is recorded. With exactly one, it is taken without a call.
- Stage 2: the model. The floor model is shown each candidate's labels, its price in and out per million tokens, its reasoning efforts, and a calibration line for the kind. It is told to pick the CHEAPEST candidate that does this difficulty well. An id outside the candidates keeps the floor.
- Recorded.
AutoModelRouter.RoutingDecisionis logged at Information and stamped on the response cell asThreadMessage.ModelRouting:kind=coding (one-line fix); difficulty=simple; model=kimi-code (cheapest that can do it); floor=….
The decision logic is pure and unit-tested (AutoModelRouter.AfterStage1 / AfterStage2,
ModelLabelsAutoRoutingTest). The client only makes the two bounded calls (10 s each) and applies the
result.
Calibration
IModelCalibration.For(kind, model) returns ModelCalibrationStats, which the prompt renders as
observed for coding: 12 tasks, 7 succeeded, $3.10 per success, median 4 rounds.
The platform registers NodeModelCalibration (TryAdd, so an in-process producer's own registration
wins in either order). It is empty, exactly as EmptyModelCalibration, until
Ai:Router:CalibrationNode names a node. From then on, every replica's NodeModelCalibrationFeed serves
that node's calibration: [{kind, model, tasks, succeeded, totalCost, medianRounds}].
A producer builds ModelOutcome records (kind, model, input/output tokens, prices at the time,
succeeded, rounds) from its own items and aggregates them with ModelCalibration.Aggregate. A
node-native producer cannot register a service, and its pass runs on one replica while rounds route
on all of them, so it writes the rows onto that node. The bug-fix pool is the first producer, writing
to Hosting/BugFix:
- triage: succeeded when a difficulty was recorded;
- coding: succeeded when the fix merged, failed when the rounds or the ladder ran out;
- a bug blocked only for want of PR tooling is not counted (Essentials/BugTriage → The outcome ledger).
Entry points that use this router
| Entry point | Uses it | How |
|---|---|---|
| Composer Auto (thread chat) | yes | Model selection auto / Provider/Auto/auto |
Composer /model auto coding |
yes | Persists auto:coding (ThreadChatView.TryPinnedAutoSelection) |
| Headless rounds (bug pool, steward, triage) | yes | The agent's modelLabels, or an auto[:labels] selection |
| CLI harnesses (Claude Code, Codex, Copilot, ACP agents) | no, deliberately | They run their vendor's own CLI and model list, not this catalog. /model under a harness is forwarded to the harness. |
mw CLI |
not applicable | It has no model option; a thread it starts takes the composer's or agent's selection, so it routes the same way. |
Not yet
- The bug-fix pool is the calibration producer: it writes its ledger onto
Hosting/BugFix→calibration, and an instance withAi:Router:CalibrationNode: Hosting/BugFixserves those rows to the router on EVERY replica (NodeModelCalibration+NodeModelCalibrationFeed, a live query of that one node). Unset, stage 2 decides on labels, price and difficulty alone. - Model nodes generated from a deployment record's
ai.openRouterEU.modelslist carry nolabelsuntil the record format grows them. Label the nodes, or add the field there.