Why this matters now:The FinOps Foundation's 2026 State of FinOps report found that 98% of enterprises now actively manage AI spending, up from 31% just two years ago, and named AI cost management the number-one skill teams are adding. At the same time, the price spread between commodity and frontier models has reached roughly 100x, meaning an enterprise sending all queries to a frontier model is paying frontier prices for work a much cheaper model could handle. A routing layer that sends 60-70% of queries to smaller models has been shown to cut cost per query by 37-46%. Most teams do not have one because building it badly costs more than skipping it.

Why enterprises overspend on LLMs and why routing alone does not fix it

The default pattern in enterprise AI is to pick one frontier model and route everything through it. It is the lowest-friction way to get something working, and for a pilot or early deployment it is often the right call. The problem is that frontier models charge frontier prices on every query, including the simple ones. A question like "summarize this paragraph" or "classify this support ticket into one of five categories" costs the same as a question requiring multi-step reasoning over a 50-page contract. The work is not the same, so paying the same rate for both is waste at scale.

The case for routing is simple on paper: cheaper model for easy queries, expensive model for hard ones, savings without quality loss. Execution is harder. A misclassifying routing layer degrades quality on the queries it gets wrong. A classifier that takes 600ms to decide on every request produces a net latency penalty regardless of the model savings. And a routing setup with no quality monitoring drifts silently as the query distribution shifts. The teams that get real savings from routing spend as much time on the classifier and the monitoring as they spend on model selection itself.

The RouteLLM research demonstrated 85% cost savings while retaining 95% of GPT-4 quality on MT Bench benchmarks, which is the number that typically gets cited in routing discussions. The more useful number for production planning is the range that well-tuned enterprise routing layers have actually delivered on real workloads: 37-46% cost-per-query reduction, which is meaningful at volume but assumes the routing classifier is accurate, the quality thresholds are calibrated, and the fallback logic works. Five architecture decisions determine whether you reach that range or build an expensive layer that underperforms a single well-chosen model.

"A routing layer that misclassifies 20% of queries does not save money. It degrades quality on those queries while adding latency to all of them."
100x
price spread between the cheapest usable LLM and the most capable frontier model in 2026, making workload-to-model matching the highest-impact cost lever available
37-46%
cost-per-query reduction achieved by routing layers that direct 60-70% of traffic to smaller models, with quality maintained on calibrated thresholds
98%
of enterprises now actively manage AI spending, up from 31% two years ago; AI cost management is the #1 skill teams are adding in 2026 (FinOps Foundation)

The 5 routing approaches compared: where each one saves money and where each one breaks

Most teams discover these approaches by trying them in order, learning what breaks only after a cost or quality incident. The table maps the failure modes so you can skip the expensive lessons.

Routing approachHow it worksWhere it saves moneyWhere it failsProduction risk
Single frontier modelAll queries sent to one model regardless of complexityNothing; predictable but expensive at scaleFrontier cost on simple tasks; cost per query does not decrease as volume growsCritical
Static task tiersManually assign task types to model tiers by prompt template or endpointEasy wins on predictable, well-defined task classesBreaks on novel query types; requires ongoing manual maintenance as tasks evolveHigh
Rule-based classifierRegex or keyword rules route queries to tiersZero-latency overhead; saves on well-scoped domainsBrittle on natural language variation; misses context and ambiguityHigh
Embedding classifierLightweight model scores query complexity before routing80-88% routing accuracy at under 4ms overhead; scales to mixed query distributionsRequires labeled training data from production traffic; accuracy degrades on out-of-distribution queries without retrainingModerate
Production-validated routing layerClassifier plus calibrated quality thresholds, fallback logic, and monitoring37-46% cost reduction at volume; quality maintained through threshold calibration and escalationTakes 4-6 weeks to calibrate; requires production traffic data before thresholds are reliableLower

The research on query routing classifiers puts accuracy at 80-91% depending on approach. Embedding-based classifiers score around 80% at sub-millisecond latency. WideMLP classifiers reach 88% accuracy at under 4ms. LLM-as-judge routers achieve up to 91% accuracy but add 60-600ms per query depending on whether the judge runs locally or via API. That latency overhead is often the reason teams abandon LLM-based routing: a router that costs 600ms to decide adds more to user experience than the cheaper model saves on the bill. The right classifier choice depends on your latency budget, not just your accuracy target.

Not sure where your AI model routing gaps are?

10decoders audits your current AI architecture, identifies which query classes are costing you frontier prices for commodity work, and builds the routing layer your production system needs. We design, calibrate, and monitor routing strategies for mid-market healthcare and fintech teams running multiple models at volume.

Book a Free AI Assessment →

Why quality thresholds matter more than model selection

Most routing discussions focus on which models to put in the routing pool. The harder decision is where to set the quality threshold that determines which model handles a given query. Set it too low and cheap-model errors reach users. Set it too high and almost everything routes to the expensive model, producing a routing layer that adds overhead without meaningful savings. The threshold is not a static configuration: it needs calibration against your actual task distribution and measurement against your actual quality definition, which is rarely the same as a public benchmark score.

A production quality threshold for routing needs to answer a specific question: for this task class, on this query population, what is the minimum acceptable response quality, and at what accuracy does the cheaper model fall below it? Published benchmarks cannot answer that question. It requires labeled evaluation data from your own traffic. The teams that skip this calibration step set their thresholds based on intuition and then discover the hard way that the cheap model's quality on their domain is different from its quality on MT Bench. A conservative threshold that leans toward over-routing to the expensive model is a better starting position than one that optimizes aggressively for cost before the quality baseline is known.

One practical approach is to start with 100% of traffic on the frontier model, log all queries and responses, build a small labeled set of 200-400 examples where human reviewers rated quality, then train an embedding classifier on that labeled set and calibrate thresholds against the same labels. That sequence takes 4-6 weeks for a team with one engineer on it, but it produces a router with known accuracy and a quality threshold grounded in real data. A router built this way catches its own degradation: when accuracy on the labeled set drops below the target, it is a signal that the classifier needs retraining, not that the routing strategy is wrong.

Stage 1
Single-model baseline

All queries on one model

Frontier model handles everything. Predictable quality, predictable cost, no routing overhead. Cost per query does not decrease as volume grows.

Stage 2
Static task tiers

Manual task-to-model assignment

Defined task types assigned to model tiers by prompt template or endpoint. Saves on well-scoped tasks; breaks on novel queries. No quality monitoring in place.

Stage 3
Production routing layer

Calibrated classifier with monitoring

Embedding classifier routes by complexity. Quality thresholds calibrated on production-derived labels. Fallback escalation and routing accuracy tracked in production.

The model routing readiness checklist for production teams

Before deploying a routing layer on a production AI system, each item below should be in place. These are the controls that determine whether routing reduces costs or introduces new failure modes your users will encounter first.

Model routing pre-deployment checklist
Query complexity taxonomy written before classifier training.Define what "simple" and "complex" mean for your specific task types before choosing a classifier approach. A support ticket triage system has a different complexity distribution than a contract analysis tool. Classifier training on a poorly defined taxonomy produces a classifier that cannot distinguish what actually matters for your workload.
Routing classifier benchmarked on production traffic, not public benchmarks.A classifier that reaches 88% accuracy on a published benchmark may reach 60% on your query distribution if your domain language differs significantly from the training data. Run accuracy measurement on a sample of your actual production queries before deploying the routing layer to live traffic.
Quality threshold calibrated per task class, not globally.A single quality threshold applied across all task types will be too aggressive on some and too conservative on others. Calibrate minimum acceptable quality separately for classification tasks, generation tasks, and reasoning tasks. Each class has a different tolerance for cheap-model errors and a different cost of getting it wrong.
Fallback escalation path defined and tested before go-live.When the cheap model produces a response below the quality threshold (if you have an output scorer) or when the classifier confidence is low, the query needs a defined path to the expensive model. That path should be tested under load before the routing layer goes live. An untested fallback path that fails under concurrent requests defeats the point of having one.
Classifier latency measured against your user experience budget.An LLM-based router adds 60-600ms depending on whether it runs locally or via API. An embedding classifier adds under 4ms. If your p95 latency target is 300ms, a remote LLM router consuming 600ms of it has a net negative effect regardless of the cost savings. Choose the classifier approach based on your latency budget, then optimize for accuracy within that constraint.
Routing decisions logged per query in production.Each query's routing decision, the classifier's confidence score, the model it was sent to, and the response latency should be logged. Without this data, you cannot diagnose routing degradation, cannot retrain the classifier on updated distribution, and cannot attribute cost savings to specific query classes or product areas.
Classifier accuracy tracked separately from end-to-end quality.End-to-end quality metrics can look stable while routing accuracy degrades, if the quality issues hit a small enough subset of queries. Track routing accuracy on a held-out labeled set at least weekly. A drop in routing accuracy is an early warning that the query distribution has shifted and the classifier needs retraining, before that shift shows up in user-facing quality reports.
"The teams getting 37-46% cost reduction from routing are not using smarter models. They are using cheaper models on the right queries, with a classifier that knows the difference."

What to do this week

Pull your current query distribution before touching the architecture

Sample 500 recent queries from your production logs and classify each one by complexity: single-step classification, short-form generation, multi-step reasoning, or retrieval-dependent synthesis. Calculate the percentage of each. If more than 50% of your queries fall into the first two categories, you have a workload that a well-tuned routing layer can serve cheaply without touching the hard queries. That distribution analysis costs one engineer a day and tells you whether routing is worth building before you commit to building it.

Measure the cost of your current setup per query class

Track your actual API spend by query type, not just in aggregate. A single frontier model handling 10,000 classification queries per day at $2.50 per million input tokens costs roughly $250 per day on that query class alone. The same queries on a $0.15 per million input token model cost $15 per day. Run that math across a 90-day period and the routing investment pays for itself before you finish building it. Framing cost savings by query class makes the case concrete rather than directional.

Start with static task tiers before building a classifier

If your product already has distinct task types separated by prompt template or API endpoint, assign each task type to a model tier manually before building a classifier. Static tiers produce real cost savings with no classifier overhead and no training data requirement. They break when query types are ambiguous or mixed, but for many enterprise products with well-defined task boundaries, static tiers capture 60-70% of the available routing savings before any learned classifier is needed. Build the classifier for the remainder once you have production data on where the static tiers are breaking.

Build the labeled dataset before the classifier

A routing classifier is only as good as the labels it was trained on. Before writing any classifier code, build a labeled dataset of 200-400 production queries with human-assigned complexity ratings and expected quality thresholds. That dataset is the most valuable artifact in the routing project: it defines what the classifier is trying to learn, it is used to measure routing accuracy, and it is used to calibrate quality thresholds. Teams that skip the labeling phase and train on synthetic or proxy labels get classifiers that look accurate on the training distribution and break on production queries that fall outside it.

Let 10decoders design your AI model routing strategy

We audit your current AI spend by query class, identify which workloads are running on models more expensive than the task requires, and build the routing layer your production volume needs. From classifier design to quality threshold calibration and production monitoring, 10decoders handles the architecture decisions most teams defer until the bill is already large.