Model routing: the AI cost skill to hire for in 2026
Model routing: the AI cost skill to hire for in 2026
Model routing means sending each AI request to the cheapest model that can handle it well, open-weight models included, instead of sending everything to a frontier model. In 2026, Uber, AT&T, and Pinterest reported inference cost figures between 34% and 92% lower, each measured differently, from levers including routing, caching, effort caps, open-weight models, and post-training. If you’re hiring for AI platform, applied AI, or DevOps roles, this is now a skill to screen for.
Picking an AI model used to be a vendor decision. You signed with a provider and the bill was whatever the bill was. In 2026, it’s an engineering decision, and it shows up on the P&L.
The evidence landed in August. Uber’s engineering team reported cost per session down 52% from its June peak, while weekly agentic requests grew 9.4× between February and August. AT&T cut the cost of coding and some other advanced AI tasks by as much as 56%. Pinterest’s CEO told investors its open models run at under 8% of the cost of comparable closed ones. The Pragmatic Engineer pulled all three together in its September 10 Pulse.
The budget side agrees. The FinOps Foundation’s State of FinOps 2026 survey of 1,192 practitioners found that 98% now manage AI spend, up from 31% two years earlier. AI cost management was the skillset they most wanted to add in the next 12 months.
This guide covers what model routing is, why those numbers don’t mean the same thing, which skills produced them, and five screening questions for your next interview. We’re former software engineers who run Data & AI and Cloud & DevOps searches, so we read the source material the way your tech lead would.
Key takeaways
- Model routing sends each request to the cheapest model that meets the quality bar. Uber, AT&T, and Pinterest reported inference cost figures 34–92% lower in 2026, each on a different metric and from several levers, not open-weight models alone.
- The figures don’t share a denominator: Uber reports cost per 1,000 requests (−34%) and cost per session (−52%), AT&T an “up to 56%” cut on specific tasks, and Pinterest a cost-per-transaction ratio against closed models. A candidate who quotes a saving without the denominator is repeating a headline.
- Savings came with trade-offs. AT&T measured a 2% quality drop. Pinterest credits its performance to post-training open models on its own data, which is specialist ML work.
- Routing, caching, and post-training are three different skills. Post-training needs a separate ML search. The first two belong in the platform and applied AI roles you already hire for.
- Don’t open an “AI FinOps engineer” req by default. Uber’s post describes cost levers built into engineering work and doesn’t mention a separate cost team. Add the skill to existing specs and test for it.
What is model routing?
Model routing is a layer between your application and your models that decides which model serves each request. Simple work such as classification, extraction, or a short reply goes to a small or open-weight model. Hard reasoning goes to a frontier model. The router can be a static rule (“summaries go to model A”) or a learned classifier that scores each request first.
Two other levers usually ship alongside routing, and people often lump them in with it:
- Prompt caching: storing the processed prefix of a prompt so later requests that reuse it pay a reduced rate
- Effort and context caps: limiting reasoning depth and context size by default, and raising them only when a task needs it
Keep these separate in your head and in your job spec. They draw on different knowledge, and a candidate can be strong in one and weak in another. If one person will own this layer end to end, the role overlaps heavily with what an applied AI engineer does day to day.
What did Uber, AT&T, and Pinterest actually report?
The reported figures run from 34% to 92%, but each company measured something different. The easy summary is “open models cut the bill in half.” The primary sources say something more useful: three companies, four figures, four denominators, one direction.
| Company | What they measured | Figure | How | Source |
|---|---|---|---|---|
| Uber | Cost per 1,000 model requests | ~34% below peak (Feb–Jul 2026) | Pareto model selection across frontier and open-weight models, prompt caching, effort and context caps | Uber Engineering, Aug 27, 2026 |
| Uber | Cost per session | 52% below June peak | Same | Same |
| AT&T | Cost of coding and some other advanced AI tasks | Up to 56% lower, with 2% lower quality | LiteLLM routers; Nemotron, Llama, and Gemma alongside Anthropic and OpenAI models | The Information via PYMNTS, Aug 20, 2026 |
| Cost per transaction, compared with comparable closed models | Under 8% of closed-model cost (a ratio, not a before/after cut) | Open models, some post-trained on Pinterest data, for example Pinterest Assistant | Q2 2026 earnings call, Aug 4, 2026 |
Why the numbers don’t average
Start with Uber. In a single post, the same company reports a 34% drop in cost per 1,000 requests and a 52% drop in cost per session. The window and the baseline differ, so both can be true at once. Meanwhile Uber says total AI spend has “relatively stabilized since April,” even as weekly agentic requests grew 9.4× and weekly active users 7× between February and August. (The Pragmatic Engineer’s summary says “since March.” Uber’s own post says April, so we cite Uber.)
That’s the first thing to understand about this skill. A cost figure without a denominator is not a result. Per request, per session, per transaction, and total spend can move in different directions in the same quarter. An engineer who owns inference cost knows which one the business cares about.
Uber’s post gives its own definition: Pareto efficient means “cost/completed task, output quality, and model reliability.” Three numbers, tracked together. That sentence is close to a job spec on its own.
What each company traded for the saving
- AT&T accepted a measured quality loss. Mark Austin, the AT&T vice president overseeing AI for employees, put it at 2%. He also said open models trail frontier ones by 6–10 months. AT&T sends about 40% of queries to open models and aims for 60–70%. It evaluated DeepSeek and Moonshot models and didn’t adopt them.
- Pinterest invested in post-training. CEO Bill Ready said that “with open models, we are achieving cost per transaction at less than 8% of the cost of comparable closed proprietary models.” On the same call, he credited better performance for Pinterest’s use cases to post-training open models on its own data. The cost figure comes from open models. The quality that lets Pinterest rely on them comes from ML engineering, not a router config.
- Uber spread the work across many levers. Its post describes a harness that serves “any model, frontier or open-weight, behind one interface,” an internal benchmark built from thousands of real pull requests, prompt caching, a 400K-token context cap, and medium reasoning effort by default. It also notes that longer caches cost more to write: 1.25× for a 5-minute entry, 2× for a 1-hour one.
Why does the model you pick matter so much for cost?
Per-call price gaps between models are wide enough that routing even part of the traffic changes the bill. The Pragmatic Engineer’s September 10 Pulse cited a code-review comparison: the most expensive open-weight model cost $0.30 (USD) per review, against $0.50 for the cheapest frontier model and $2.50 for the most expensive.
Run those prices at an illustrative 10,000 reviews a month. Open-weight: $3,000. Frontier: $5,000 to $25,000. The spread is up to $22,000 a month, or $264,000 a year.
Capturing that spread is routing and evaluation work, not ML research. You need someone who can prove the cheaper model is good enough for that workload and keep proving it as models change. Pinterest’s approach is different. If you want open models that outperform closed ones on your own data, you’re hiring for a much narrower profile.
Routing, caching, and post-training are three different skills
“Reduce our LLM costs” reads like one skill in a job ad. In practice it’s three, and they draw on different backgrounds. This is where hiring goes wrong.
| Skill | What the work looks like | Typical role | Hiring note |
|---|---|---|---|
| Model routing and effort control | Tiering workloads, running evals before switching models, setting reasoning and context defaults, handling provider deprecations | AI platform engineer, applied AI engineer | Requires production LLM experience; demos don’t exercise it |
| Caching and context economics | Structuring prompts so prefixes cache, choosing cache lifetimes, batching chatty calls | Backend or platform engineer | Adjacent to existing backend skills |
| Post-training open models | Fine-tuning on proprietary data, evaluation, serving on vLLM, SGLang, or TensorRT-LLM | ML engineer | Specialist ML work; budget for a separate search |
Uber and AT&T describe mostly the first two rows. Pinterest’s quality claim rests on the third. If your roadmap looks like Pinterest’s, budget for an ML search. Our guide to hiring ML engineers in Romania covers what that involves.
What should a Data & AI or DevOps job spec say?
Write responsibilities as ownership of a number, with a quality gate attached. “Experience with LLMs” and “prompt engineering” describe nothing a candidate can be measured on.
Spec lines that work:
- Own cost per [request / session / transaction] for [product or platform], and report it weekly next to a quality metric.
- Maintain an evaluation suite that must pass before any workload moves to a cheaper or open-weight model.
- Design and operate the routing layer (for example, LiteLLM or an internal gateway) across frontier and open-weight models.
- Set and review default reasoning effort, context limits, and caching policy per workload.
- Own model migrations: track provider deprecations, run regression tests, ship the switch.
- (ML roles only) Post-train and serve open models on vLLM, SGLang, or TensorRT-LLM against a cost-per-transaction target.
Lines to cut: “familiarity with GPT and Claude,” “prompt engineering,” “passion for AI.” None of them tell a candidate what they’ll be accountable for, or tell you what to test.
A note on role design. Uber’s post describes cost as part of engineering work and doesn’t mention a separate FinOps team or cost-owner title. AT&T’s program sits with a vice president responsible for employee AI use. Neither source shows a standalone “AI FinOps engineer” behind the results. For most teams, the better move is to add lines 1–5 to the AI platform or applied AI role you already have.
Reality check: “model routing” is a new label, so don’t expect it on CVs. Screen for the behavior instead, with the questions below.
Five screening questions for inference cost judgment
Ask for the denominator first, then test quality, risk, caching economics, and migrations. Each question targets a failure you can see in the source material above.
| # | Question | Strong answer | Red flag |
|---|---|---|---|
| 1 | “You cut AI cost by 50%. Per what?” | Names the denominator, the baseline, and what total spend did | “Our bill went down” |
| 2 | “Which requests would you never route to an open-weight model?” | Names workloads by risk and error cost, and says how they found out | “Open models are as good now, so all of them” |
| 3 | “How did you measure quality before and after the switch?” | An eval suite, a number, and an acceptance threshold (AT&T’s 2% is the kind of figure you want) | “Users didn’t complain” |
| 4 | “When does a longer prompt cache cost more than it saves?” | Explains that cache writes carry a premium, so short-lived or low-reuse prefixes don’t earn it back | Treats caching as free money |
| 5 | “A provider deprecates the model under your product next month. Walk me through it.” | Regression suite, shadow traffic, staged rollout, a rollback plan | “We’d update the model name” |
Hiring for inference cost in Romania and CEE
Look for engineers who shipped LLM features to production and had to justify the bill. Coursework and demos don’t involve a bill, so they can’t show this judgment. In practice, that points to backend and platform engineers who moved into applied AI, as much as to ML specialists.
The honest friction:
- You’re competing for one of the hardest hires of 2026. Christian & Timbers’ AI-Native Builder Report 2026, citing Lightcast labor data, puts demand for AI-native builders at 3.4 times available supply as of Q1 2026. It also reports that roughly seven in ten of the searches that close do so through direct outreach to passive candidates, not inbound applications. That’s market-wide data, not Romania-specific, but the CEE talent pool competes for the same people.
- Budget for real AI engineering pay. Our applied AI engineer guide puts indicative Romanian pay for AI engineers at roughly €4,900 a month for mid-level to €8,000 or more for lead roles, based on ERI, levels.fyi, and Glassdoor data that disagree on gross versus net. Check the ML and AI ranges in our 2026 Romanian IT salary guide before you set the budget, and add headroom if the plan includes post-training.
- The field moves fast. AT&T’s own estimate is a 6–10 month gap between open and frontier models, and Uber writes that “the frontier shifts every few weeks.” Whoever you hire will re-run evaluations for as long as they’re in the role. Hire for the discipline of measuring, not for knowledge of this quarter’s models.
We argue in AI-native engineering that the scarce skill in AI-native teams is verification: proving AI output is good enough before trusting it. Inference cost is the same judgment, applied to the bill.
Frequently asked questions
What is model routing in AI? Model routing is a layer that sends each request to the model best suited to it on cost and quality. Simple tasks go to small or open-weight models, and hard ones go to frontier models. It usually ships with prompt caching and limits on reasoning effort and context size.
How much can open-weight models cut LLM costs? It depends on the metric and the other levers involved. Uber’s 34% and 52% figures came from several levers, not open-weight models alone. AT&T’s 56% is an “up to” figure on specific tasks. Pinterest’s under-8% is a comparison with closed models, not a before/after cut.
Do open-weight models lower quality compared to frontier models? Sometimes, so measure it rather than assume. AT&T measured a 2% quality decline and estimated open models trail frontier ones by 6–10 months. Pinterest reported better performance for its use cases after post-training open models on its own data.
Do I need an AI FinOps role, or can my platform team own this? For most teams, the platform or applied AI team can own it. Uber’s post describes these levers as part of engineering work and doesn’t mention a separate cost team. Add cost-per-request ownership and an evaluation gate to existing job specs first.
How do I interview an engineer about LLM cost optimization? Start with the denominator: ask what their last cost saving was measured per. The five questions in the screening table above cover quality, routing risk, caching, and migrations.
The bottom line
Model routing turned AI inference cost from a contract term into an engineering result. The August figures from Uber, AT&T, and Pinterest point the same way, but they measure different things and came from different techniques.
Three things to take from this:
- Ask for the denominator. Per request, per session, per transaction, and total spend are different results.
- Hire the skill, not a new title. Routing and caching belong in AI platform and applied AI specs. Post-training is a separate ML search.
- No saving without an eval. AT&T’s 2% quality loss is the kind of number a good engineer brings with them.
If you’re about to hire an engineer to own model routing and inference cost, we can help you write the spec and test the judgment behind the claims. We’re former engineers, so we screen with the questions above. You get a curated shortlist of three to five candidates, not a CV flood. Tell us what you’re hiring for.
Last updated: September 14, 2026. Sources verified September 14, 2026: Uber Engineering, “Running a Software Factory Efficiently at Uber Scale” (Uday Kiran Medisetty, Aug 27, 2026); The Information via PYMNTS (Aug 20, 2026); Pinterest Q2 2026 earnings call transcript, The Motley Fool (call held Aug 4, transcript published Aug 11, 2026); The Pragmatic Engineer, “The Pulse: tech companies move to open AI models” (Sep 10, 2026); FinOps Foundation, State of FinOps 2026; Christian & Timbers, AI-Native Builder Report 2026. Romanian pay ranges are from our applied AI engineer guide and are directional. The per-review code-review prices are as cited by The Pragmatic Engineer. The 10,000-reviews-a-month example is illustrative arithmetic on those prices, not a reported deployment.