Docs · Routing & failover
Routing & failover
How Frontier decides which agent takes a task, and the narrow set of failures that move work to a different one.
Eligibility comes first
Before anything is scored, a provider must be:
- enabled and detected;
- below its concurrency limit and any optional token budget;
- outside a cooldown from a recent quota or unavailability error;
- below any plan-window utilisation its own CLI has reported.
Anything that fails one of those tests is not a candidate at all.
Then the ranking
Remaining providers are scored, highest first, on:
Your provider override
An agent pinned on the task wins outright.
Routing policy
Balanced, Quality first, or Token saver.
Task and provider affinity
How well that agent suits this kind of work.
Your priority order
The order you arranged the providers in.
Today's usage and current load
Locally estimated tokens and how busy the agent is right now.
Task classification
Every task is typed as coding, debugging, review, planning, documentation, or general work. The type feeds the affinity score — which is why the plain Ollama provider, having no file-editing tools, is only ever offered planning, review, documentation, and general tasks.
Routing policies
| Policy | Optimises for |
|---|---|
| Balanced | A sensible mix of capability and spreading load across your subscriptions. |
| Quality first | The strongest agent for the task type, accepting faster quota burn. |
| Token saver | Prefers local and cheaper agents, keeping subscription windows in reserve. |
When work moves, and when it stops
This is the distinction that matters most, and Frontier is deliberately conservative about it.
| What happened | What Frontier does |
|---|---|
| Quota / rate limit | Cools that provider down and fails over to the next eligible agent. |
| Overloaded / unavailable | Same — treated as temporary unavailability. |
| CLI is logged out | Classified as unavailable on a non-zero exit, so it cools down and fails over instead of failing the whole task. The real fix is still to log that CLI back in. |
| Agent failure | Stops the task. Re-running a partially completed coding task through another agent could duplicate or conflict with edits already on disk. |
| You cancelled it | Terminal for that run. Cancellation never enters automatic failover. |
Failover applies everywhere a provider runs: first turns, follow-up messages, and every stage of an orchestrated run. Reported 100% plan utilisation and configured tracked-usage limits also remove a provider from routing before launch, so a doomed run never starts.
Concurrency and budgets
Frontier runs multiple independent tasks in parallel while respecting a global limit and a per-provider limit (the shipped default is one concurrent run per agent). Optional per-provider daily token budgets give you a deterministic ceiling on top of whatever the CLI itself enforces — useful because subscription tools expose no universal "tokens remaining" interface.
Picking a model, not just an agent
Routing does not stop at "which agent" — each candidate model is placed on a tier (local, fast, standard, frontier), and the Balanced/Quality-first/ Token-saver policy shifts which tier a task is aimed at. Providers also accumulate outcome-aware routing signal over time — completion, whether your own repo checks passed, and whether you merged or discarded the resulting branch in the Review inbox — folded into one bounded, labelled factor per provider (and, once enough runs exist, per model) that nudges the ranking without ever overruling policy or an explicit pick. Turn it off in Settings and the router scores exactly as it did without it.
Jev — the optional routing advisor
Jev (TypeSafe's System One model) is an auxiliary service, not a coding agent: it never generates text and never runs on your behalf, which is why it falls under the narrowed, opt-in exception to Frontier's "no API keys" rule rather than breaking it — see security & data. Jev advises; the router still decides. Its answers only ever become bounded, labelled routing factors next to every other factor in this page — they never override eligibility, an explicit pick, or a user-picked model.
What Jev adds
One request asks a single typed, calibrated question set about the task: its type (the same six categories used everywhere else), a complexity score across four levels (trivial → single-file → multi-file → architectural), whether it edits files, whether it needs a large context, whether it is worth splitting into subtasks, and a best-fit target — a probability distribution over every eligible provider/model pair, each described by what that model is actually good at (and honest about what it can't do, such as an Ollama provider having no file tools).
Off / Shadow / Active
| Mode | What happens |
|---|---|
| Off | The default. Routing scores exactly as it would with no advisor at all. |
| Shadow | Jev is asked and its pick is recorded next to the real route — visible in the inspector's Route section and in the calibration view below — but it changes nothing about what actually runs. |
| Active | Jev's answers become real, bounded routing factors and can influence which provider and model are chosen, still subject to every eligibility rule and an explicit pick. |
Confidence threshold and the 3-second fallback
A task_type or complexity answer below a configurable minimum confidence
(0.5 by default) is ignored rather than trusted. Jev is also never allowed to stall the queue: each request has a
hard timeout with a single retry on a rate-limit response, all bounded by an overall 3-second
deadline — a request that has not settled by then is treated as a miss, and the task simply falls back to
Frontier's own built-in heuristic classifier. The queue itself is never blocked waiting on a Jev response; the
engine computes the heuristic answer synchronously and only fires the Jev call alongside it.
Per-subtask advice
When an orchestrated task's planner returns two or more subtasks, and the advisor is not off, Frontier makes one additional Jev call covering every subtask (up to the first eight) at once, asking complexity and best-fit questions per subtask. In active mode, each subtask that got advice routes on it independently — so a trivial rename and an architectural refactor inside the same plan can land on different-tier models of the same agent, instead of every subtask inheriting one provider default regardless of its own complexity. A subtask that got no advice, or any failure of the extra call, simply keeps today's default behaviour.
The calibration view
Settings → Routing buckets every finished, Jev-advised task by the confidence of its best-fit answer
(< 0.5, 0.5 – 0.8, ≥ 0.8) and reports, per bucket, how many tasks ran,
how many completed, how many passed your own repo's verification checks, and whether Jev's pick agreed with
whoever actually ran the task — so you can see whether higher-confidence answers actually correlate with better
outcomes before switching from shadow to active.
Jev is built and hosted by TypeSafe; see TypeSafe's documentation for the model itself.