Why enterprises overspend on LLMs and why routing alone does not fix it
The default pattern in enterprise AI is to pick one frontier model and route everything through it. It is the lowest-friction way to get something working, and for a pilot or early deployment it is often the right call. The problem is that frontier models charge frontier prices on every query, including the simple ones. A question like "summarize this paragraph" or "classify this support ticket into one of five categories" costs the same as a question requiring multi-step reasoning over a 50-page contract. The work is not the same, so paying the same rate for both is waste at scale.
The case for routing is simple on paper: cheaper model for easy queries, expensive model for hard ones, savings without quality loss. Execution is harder. A misclassifying routing layer degrades quality on the queries it gets wrong. A classifier that takes 600ms to decide on every request produces a net latency penalty regardless of the model savings. And a routing setup with no quality monitoring drifts silently as the query distribution shifts. The teams that get real savings from routing spend as much time on the classifier and the monitoring as they spend on model selection itself.
The RouteLLM research demonstrated 85% cost savings while retaining 95% of GPT-4 quality on MT Bench benchmarks, which is the number that typically gets cited in routing discussions. The more useful number for production planning is the range that well-tuned enterprise routing layers have actually delivered on real workloads: 37-46% cost-per-query reduction, which is meaningful at volume but assumes the routing classifier is accurate, the quality thresholds are calibrated, and the fallback logic works. Five architecture decisions determine whether you reach that range or build an expensive layer that underperforms a single well-chosen model.
"A routing layer that misclassifies 20% of queries does not save money. It degrades quality on those queries while adding latency to all of them."
The 5 routing approaches compared: where each one saves money and where each one breaks
Most teams discover these approaches by trying them in order, learning what breaks only after a cost or quality incident. The table maps the failure modes so you can skip the expensive lessons.
| Routing approach | How it works | Where it saves money | Where it fails | Production risk |
|---|---|---|---|---|
| Single frontier model | All queries sent to one model regardless of complexity | Nothing; predictable but expensive at scale | Frontier cost on simple tasks; cost per query does not decrease as volume grows | Critical |
| Static task tiers | Manually assign task types to model tiers by prompt template or endpoint | Easy wins on predictable, well-defined task classes | Breaks on novel query types; requires ongoing manual maintenance as tasks evolve | High |
| Rule-based classifier | Regex or keyword rules route queries to tiers | Zero-latency overhead; saves on well-scoped domains | Brittle on natural language variation; misses context and ambiguity | High |
| Embedding classifier | Lightweight model scores query complexity before routing | 80-88% routing accuracy at under 4ms overhead; scales to mixed query distributions | Requires labeled training data from production traffic; accuracy degrades on out-of-distribution queries without retraining | Moderate |
| Production-validated routing layer | Classifier plus calibrated quality thresholds, fallback logic, and monitoring | 37-46% cost reduction at volume; quality maintained through threshold calibration and escalation | Takes 4-6 weeks to calibrate; requires production traffic data before thresholds are reliable | Lower |
The research on query routing classifiers puts accuracy at 80-91% depending on approach. Embedding-based classifiers score around 80% at sub-millisecond latency. WideMLP classifiers reach 88% accuracy at under 4ms. LLM-as-judge routers achieve up to 91% accuracy but add 60-600ms per query depending on whether the judge runs locally or via API. That latency overhead is often the reason teams abandon LLM-based routing: a router that costs 600ms to decide adds more to user experience than the cheaper model saves on the bill. The right classifier choice depends on your latency budget, not just your accuracy target.
Not sure where your AI model routing gaps are?
10decoders audits your current AI architecture, identifies which query classes are costing you frontier prices for commodity work, and builds the routing layer your production system needs. We design, calibrate, and monitor routing strategies for mid-market healthcare and fintech teams running multiple models at volume.
Book a Free AI Assessment →Why quality thresholds matter more than model selection
Most routing discussions focus on which models to put in the routing pool. The harder decision is where to set the quality threshold that determines which model handles a given query. Set it too low and cheap-model errors reach users. Set it too high and almost everything routes to the expensive model, producing a routing layer that adds overhead without meaningful savings. The threshold is not a static configuration: it needs calibration against your actual task distribution and measurement against your actual quality definition, which is rarely the same as a public benchmark score.
A production quality threshold for routing needs to answer a specific question: for this task class, on this query population, what is the minimum acceptable response quality, and at what accuracy does the cheaper model fall below it? Published benchmarks cannot answer that question. It requires labeled evaluation data from your own traffic. The teams that skip this calibration step set their thresholds based on intuition and then discover the hard way that the cheap model's quality on their domain is different from its quality on MT Bench. A conservative threshold that leans toward over-routing to the expensive model is a better starting position than one that optimizes aggressively for cost before the quality baseline is known.
One practical approach is to start with 100% of traffic on the frontier model, log all queries and responses, build a small labeled set of 200-400 examples where human reviewers rated quality, then train an embedding classifier on that labeled set and calibrate thresholds against the same labels. That sequence takes 4-6 weeks for a team with one engineer on it, but it produces a router with known accuracy and a quality threshold grounded in real data. A router built this way catches its own degradation: when accuracy on the labeled set drops below the target, it is a signal that the classifier needs retraining, not that the routing strategy is wrong.
All queries on one model
Frontier model handles everything. Predictable quality, predictable cost, no routing overhead. Cost per query does not decrease as volume grows.
Manual task-to-model assignment
Defined task types assigned to model tiers by prompt template or endpoint. Saves on well-scoped tasks; breaks on novel queries. No quality monitoring in place.
Calibrated classifier with monitoring
Embedding classifier routes by complexity. Quality thresholds calibrated on production-derived labels. Fallback escalation and routing accuracy tracked in production.
The model routing readiness checklist for production teams
Before deploying a routing layer on a production AI system, each item below should be in place. These are the controls that determine whether routing reduces costs or introduces new failure modes your users will encounter first.
"The teams getting 37-46% cost reduction from routing are not using smarter models. They are using cheaper models on the right queries, with a classifier that knows the difference."
What to do this week
Pull your current query distribution before touching the architecture
Sample 500 recent queries from your production logs and classify each one by complexity: single-step classification, short-form generation, multi-step reasoning, or retrieval-dependent synthesis. Calculate the percentage of each. If more than 50% of your queries fall into the first two categories, you have a workload that a well-tuned routing layer can serve cheaply without touching the hard queries. That distribution analysis costs one engineer a day and tells you whether routing is worth building before you commit to building it.
Measure the cost of your current setup per query class
Track your actual API spend by query type, not just in aggregate. A single frontier model handling 10,000 classification queries per day at $2.50 per million input tokens costs roughly $250 per day on that query class alone. The same queries on a $0.15 per million input token model cost $15 per day. Run that math across a 90-day period and the routing investment pays for itself before you finish building it. Framing cost savings by query class makes the case concrete rather than directional.
Start with static task tiers before building a classifier
If your product already has distinct task types separated by prompt template or API endpoint, assign each task type to a model tier manually before building a classifier. Static tiers produce real cost savings with no classifier overhead and no training data requirement. They break when query types are ambiguous or mixed, but for many enterprise products with well-defined task boundaries, static tiers capture 60-70% of the available routing savings before any learned classifier is needed. Build the classifier for the remainder once you have production data on where the static tiers are breaking.
Build the labeled dataset before the classifier
A routing classifier is only as good as the labels it was trained on. Before writing any classifier code, build a labeled dataset of 200-400 production queries with human-assigned complexity ratings and expected quality thresholds. That dataset is the most valuable artifact in the routing project: it defines what the classifier is trying to learn, it is used to measure routing accuracy, and it is used to calibrate quality thresholds. Teams that skip the labeling phase and train on synthetic or proxy labels get classifiers that look accurate on the training distribution and break on production queries that fall outside it.
Let 10decoders design your AI model routing strategy
We audit your current AI spend by query class, identify which workloads are running on models more expensive than the task requires, and build the routing layer your production volume needs. From classifier design to quality threshold calibration and production monitoring, 10decoders handles the architecture decisions most teams defer until the bill is already large.
