LLM inference cost at scale
The uncomfortable property of LLM inference cost is that it grows faster than your traffic does. Users who like a feature use it more, conversations get longer, context accumulates, and agent workloads multiply calls per task. Doubling users rarely doubles the bill, it usually more than doubles it.
Where the money actually goes
- Context length. Every turn of a conversation re-sends history. Long-context features are billed on the whole window, every call.
- Model mix. Most teams route 100% of traffic to one premium model because it was the easiest thing to ship. This is normally the single largest overspend.
- Agent trajectories. One user action can become 20 model calls. Cost per task, not per call, is the unit that matters.
- Retries and validation failures. Invisible in dashboards that only count successful responses.
- Cache misses. Repeated or near-identical prompts paying full price every time.
The five levers that work
- Route by difficulty. Classify the request and send easy work to a cheaper model behind a quality floor. This is the biggest single lever for most teams, commonly a 40–70% reduction with no measurable quality change.
- Cache aggressively. Provider prompt caching for shared prefixes, plus semantic caching for near-duplicate requests. See semantic caching.
- Trim context. Summarize history instead of replaying it. Most conversations do not need the full transcript.
- Enforce structured output. Schema-valid responses on the first attempt remove the retry tax. See structured outputs.
- Set budgets per workload. A per-key budget converts a surprise invoice into a policy decision.
Measure cost per successful task
Per-token cost tells you nothing about whether spend is productive. Per-task cost, total tokens across retries and agent steps, divided by tasks actually completed, is the number that tells you whether an optimization worked. Neural Router returns exact usage and cost on every routing receipt so this is directly computable rather than estimated.