The Real Cost of AI Inference at Scale (And How to Control It)
A classification pipeline looks like a bargain at a fraction of a cent per call. Then volume grows, and the bill grows faster than the volume. There is no bug. It is the gap between the pricing page and what you pay once you count retries, the system prompt sent with every call, and the fact that nobody batches by default.
The gap between the pricing page and your bill
Every major model provider quotes cost per million tokens. GPT-4o has listed at $2.50 per million input tokens, Claude 3.5 Sonnet at $3.00 and Gemini 1.5 Flash at $0.075. These numbers are accurate. They are also nearly useless for production budget planning.
Here is what the pricing page does not show: your system prompt is re-sent on every call unless you are using prompt caching. A 2,000-token system prompt across 9 million monthly calls is 18 billion input tokens before a single user message is processed. At $3.00 per million, that is $54K a month. From a system prompt.
The usual mistake is pricing the happy path (one user message, one response, modest context) and ignoring everything around it. The system prompt overhead. The retry tokens when the model returns malformed JSON. The logging calls that echo the full input into a tracing system. None of these are edge cases, and together they can be most of your token spend.
Add them up and the real token count per call can easily be two or three times the happy-path estimate. Measure it from your API logs before you trust any budget.
Batching vs streaming: which one is killing your budget
Streaming is the right default for user-facing features. When a person is waiting, streaming makes the product feel faster and increases completion rates. But streaming carries real overhead: a persistent connection per request and server-sent event handling. More important, it rules out the request-level batching that cuts costs sharply for background workloads.
Background pipelines almost never need streaming. Nobody watches the tokens arrive in document processing, classification or entity extraction. Using streaming here is a habit, not a requirement. The Batch API on both Anthropic and OpenAI returns results asynchronously at a 50% cost discount. Same models, same quality, half the price. For a pipeline spending $8K a month on background work, that is $4K back.
The rule we apply now: any workload without a human waiting on the response gets the Batch API. Streaming is reserved for chat, copilot features and real-time output. On a mixed workload, the saving is half of whatever share of your spend runs in the background.
Caching: what it actually saves
Prompt caching on Anthropic's API caches the prefix of a conversation, usually the system prompt and any fixed context, and charges 10% of the normal input rate on cache hits. The requirement is that the cached prefix is at least 1,024 tokens and the same content appears at the start of subsequent requests within the cache TTL.
With a stable 3,000-token system prompt and steady traffic, most requests hit the cache once the first few have warmed it. Say 80% of your input tokens are cache hits. You pay full price on the other 20% and a tenth of the price on the rest, so a $3.00-per-million model costs you about $0.84 per million input tokens, before the small premium Anthropic charges on cache writes.
Semantic caching is a different mechanism entirely. Rather than caching at the API level, you store embedding vectors of past queries and responses in a vector database. When a new query arrives, you check for semantic similarity above a threshold (0.93 cosine similarity is a common starting point) and return the cached response if the match is close enough. This skips the model call entirely.
Semantic caching works well for FAQ-style workloads with predictable query patterns. On customer support pipelines where the same questions recur, the hit rate can be high. It is a poor fit for document processing or creative generation where every input is genuinely novel. Know which workload you have before investing in the infrastructure.
Our take
Model routing: the decision that moves the needle most
Routing is the biggest cost lever in most pipelines. Not every query needs your most capable model. A request to classify a document as invoice, receipt, or contract does not need Claude 3.5 Sonnet at $3.00/million input tokens. It needs Claude 3 Haiku at $0.25/million. On a simple classification task the accuracy difference is usually small. The price difference is 12x.
We implement routing as a lightweight classifier that runs before every request. The classifier, itself a small model or a set of rules, assigns each request to one of three tiers: fast/cheap, balanced, or capable. Fast/cheap handles classification, entity extraction, simple transformations. Capable handles long-form generation, multi-step reasoning, synthesis across many documents. Balanced is the middle ground.
Before you trust the saving, measure quality on every task you route to a cheaper model. The routing criteria should come from task type and complexity, written down, not from guesswork.
Routing on token count is a common mistake. A 50-token query that asks for legal reasoning is harder than a 500-token query asking for a date extraction. Route on task type and expected reasoning depth, not input length.
The hidden costs nobody tracks
Retry logic is invisible on most cost dashboards because it looks like legitimate traffic. Without strict output enforcement, models return malformed JSON often enough to matter, and the typical implementation retries up to three times. On a high-volume pipeline, even a 2% error rate with three retries means some of your calls pay for four model calls to get one result.
The fix is constrained decoding. Both Anthropic and OpenAI support structured output modes that force the model to conform to a JSON schema at the generation level. This drops malformed output rates to near zero and eliminates the retry tax. We enable this on every pipeline that has a defined output schema. No exceptions.
Error handling overhead is a separate category. When a pipeline step fails, the default behavior in most frameworks is to log the full input and output for debugging. If your inputs are 4,000 tokens and you are logging them to a downstream system that charges per character, error logging can become a meaningful cost at scale. Set a token budget on logged payloads, for example by truncating inputs to the first 500 tokens in error logs. You keep what you need to debug and stop paying to store the rest.
Observability tooling is worth auditing separately. Tools like LangSmith, Langfuse, and Helicone are genuinely useful but charge per trace or per token logged. A system making 5 million calls a month at $0.0008 per trace pays $4K a month just for observability. Sample your traces in production. A 10% sample still shows systemic issues, at a tenth of the cost.
Watch out
A worked example: the document intake pipeline
Take a document intake system. Every incoming contract, invoice or letter gets classified, has its key fields extracted, and is routed to the right downstream process. It launches at 3,000 documents a day. A year later the business has grown and it is handling 90,000.
The original estimate assumed one classification call at 800 input tokens and 150 output tokens. In reality there are three steps per document: pre-classification, extraction, and a validation call added during QA to catch edge cases. The 2,800-token system prompt isn't cached. A JSON parsing bug nobody caught in testing keeps retries high. And everything streams, because the first prototype was a chat interface.
Every problem in that list has a fix from this article. Cache the system prompt. Move all three steps to the Batch API for the 50% discount. Enforce the JSON schema so retries fall to almost nothing. Route pre-classification to a small model like Haiku, and merge extraction and validation into one call with a richer schema.
None of these changes the throughput or, if you measure as you go, the quality. Together they can cut the monthly bill several times over, and none of them is more than a few days of work.
Where to start
If you are running a pipeline spending more than $2K/month on inference, run this audit in order:
- 01Measure your actual token ratio. Pull real API logs, calculate average input and output tokens per call, and compare to your original estimate. If reality is more than 1.5x your estimate, find out why before optimizing anything else.
- 02Identify every background workload. List every pipeline step. Mark each one as "human waiting" or "background." Migrate every background step to the Batch API.
- 03Enable prompt caching on every pipeline with a system prompt over 1,024 tokens. Structure prompts so all static content comes first. Check your cache hit rate after 48 hours.
- 04Audit retries. If your retry rate is above 3%, fix structured output enforcement before anything else.
- 05Profile each step by task type. Route simple classification and extraction to a smaller model. Measure quality. Keep routing only for steps where quality holds.
None of this is advanced optimization. It corrects defaults that were set for convenience during development and never revisited. Plenty of pipelines have two or more of these problems at once. Finding them is a spreadsheet exercise with API logs. The engineering to fix them is typically one to three days of work.
Related reading