AI EngineeringBuild
The Hidden Cost of Running AI Applications in Production
The token price didn't change between prototype and production. The architecture did. Where AI app costs really come from, and how to cut them.
The bill nobody planned for
In development, the assistant cost almost nothing. A few engineers, a few hundred test questions, a rounding error on the invoice.
Then it shipped. The system prompt had grown to a few thousand tokens and went out with every request. Conversations got longer, and the full history travelled with each turn. Retrieval pulled twenty chunks "to be safe". The agent made three or four model calls per question. A flaky upstream service triggered automatic retries. A nightly eval suite fired hundreds of model calls. Even the logs, full of prompts and responses, became a line item.
The price per token never changed. The architecture did. That is the real story behind LLM cost in production.
The model invoice is only the beginning
A prototype costs roughly one model call per question. A production system costs everything around that call: embeddings, vector search, tool executions, retries, evaluations, the compute that hosts it and the observability that watches it.
A useful mental model, not a billing formula:
AI system cost ~ model inference + retrieval and embeddings + tool calls
+ retries + evaluation + infrastructure and observability
+ wasted context
The metric that exposes it is cost per successful outcome: the total cost of every call, retry and tool step needed to complete one user task. Cost per request can look fine while retries and agent loops quietly multiply it.
Where the cost actually comes from
| Cost driver | What creates it | Engineering lever |
|---|---|---|
| Input tokens | Prompts, history, retrieved docs and tool results sent on every call | Context budgets, prompt caching |
| Output tokens | Long answers, verbose formats, reasoning tokens on some models | Per-feature output limits, concise formats |
| Embeddings | Embedding every query, re-embedding documents | Cache embeddings, re-embed only changed docs |
| Retrieval | Vector database queries, storage, rerankers | Right-size topK, filter before searching |
| Tool calls | External APIs and databases called by agents | Timeouts, caching, trimmed results |
| Agent loops | Several model calls per user task | Iteration caps, smaller models for steps |
| Retries | Timeouts, rate limits, malformed output | Classify errors, cap retries |
| Evaluation | Suite size times run frequency times model size | Tiered suites, batch processing |
| Infrastructure | Compute, serverless invocations, storage, data transfer | Right-sizing, scaling to zero |
| Observability | Logging full prompts, responses and traces | Sampling, retention limits |
Context is a recurring cost
As covered in Understanding Context Windows, the model keeps nothing between calls. The system prompt, the conversation so far and the retrieved documents are sent, and paid for, on every turn.
Here's the multiplication effect with hypothetical prices of $3 per million input tokens and $15 per million output tokens. Assume 10,000 input tokens and 2,000 output tokens per request, and 10,000 requests a month:
Input: 10,000 x 10,000 = 100M tokens x $3/M = $300
Output: 2,000 x 10,000 = 20M tokens x $15/M = $300
Total = $600 / month
Cut input context to 4,000 tokens:
Input: 40M tokens x $3/M = $120 -> total $420 (30% lower)
Notice that with these prices, input and output cost exactly the same. That's why "input tokens dominate" isn't a rule. Long contexts with short answers skew toward input; long generations and reasoning models skew toward output. Plug in your provider's real prices and your real token counts before deciding what to optimise.
RAG doesn't make context free
One RAG question is really a pipeline:
query -> embedding -> vector search -> metadata filter -> rerank -> build context -> LLM
Each step can cost something: an embedding call, a query against a vector database you pay to run, a reranker, and the extra input tokens the retrieved text adds to the model call.
That's why topK = 20 isn't automatically better than topK = 5. Twenty chunks mean more input tokens, more latency and more noise the model has to ignore. But five weak chunks can miss the answer, and a wrong answer costs a follow-up question, which is another full request. The right value depends on your retrieval quality and the task. Choose it with evals on real questions, not intuition.
Agent loops multiply calls
request -> LLM -> search tool -> LLM -> database tool -> LLM -> answer
To the user, that's one question. To your bill, it's three model calls and two tool executions. How many tokens each step uses depends on how your application builds each request: full history every time, trimmed history or a summary. Either way, cost now scales with the number of steps, and the model decides that number at runtime.
Controls that keep it bounded:
- a maximum number of iterations per task
- a timeout on every tool
- smaller, cheaper models for intermediate decisions
- tool results trimmed to what the next step needs
- early termination once the answer is good enough
- model calls per task tracked as a metric, with alerts
Retries: when reliability turns into cost
Retries are good engineering. Networks fail, providers rate-limit, and a transient 503 shouldn't become a user-facing error. The problem is the combination: unbounded retries, on large requests, to expensive models. A request that times out on your side may still have been processed, and billed, on the provider's side, so one user action can quietly become three paid calls.
The rules:
- Retry only transient failures: timeouts, rate limits and 5xx errors.
- Never retry validation or authentication errors. They will fail again.
- Use exponential backoff with jitter, and respect rate-limit headers when the provider sends them.
- Cap the number of retries.
- Make operations with side effects idempotent.
A provider-neutral wrapper that handles both retries and per-feature cost logging:
type Usage = { inputTokens: number; outputTokens: number };
type Price = { inputPerM: number; outputPerM: number }; // from config, never hardcoded
const RETRYABLE = new Set([408, 429, 500, 502, 503, 504]);
const sleep = (ms: number) => new Promise(r => setTimeout(r, ms));
export async function callLLM<T extends { usage: Usage }>(
feature: string, model: string, price: Price,
call: () => Promise<T>, maxRetries = 2,
): Promise<T> {
for (let attempt = 0; ; attempt++) {
try {
const res = await call();
const cost = (res.usage.inputTokens * price.inputPerM
+ res.usage.outputTokens * price.outputPerM) / 1_000_000;
log.info("llm_call", { feature, model, attempt, ...res.usage, cost });
return res;
} catch (err: any) {
const status = err?.status; // undefined = network error
const retryable = status === undefined || RETRYABLE.has(status);
log.warn("llm_call_failed", { feature, model, attempt, status });
if (!retryable || attempt >= maxRetries) throw err; // 400, 401, 422: fail fast
await sleep(2 ** attempt * 500 + Math.random() * 250); // backoff with jitter
}
}
}
Logging failed attempts matters as much as logging successes. They are exactly the calls that never show up in a "cost per request" view.
Evaluation: control it, don't cut it
Eval cost grows with suite size, run frequency, model size, prompt length and the number of prompt or model variants you compare. Deleting evals saves that money and ships the regressions instead.
Control it instead: run a small smoke suite on every commit and the full suite nightly or before release, use a cheaper judge model where quality allows, skip cases whose inputs haven't changed, and use asynchronous batch processing where your provider offers it at a discount.
Levers that actually work
- Prompt caching. Some providers can reuse repeated, identical prompt prefixes and bill them at a lower rate or serve them faster. Availability, eligibility rules, pricing and cache lifetime vary by provider and model. It only helps when the prefix is truly identical, so put stable content such as instructions and tool definitions first, and anything volatile last.
- Model routing. Send easy requests to a cheaper model. It needs a reliable routing signal and evals proving the cheap path is good enough.
- Context budgets. Cap history and retrieved context per request, per feature.
- Response caching. Reuse answers for repeated questions, but only when they are not personalised and not sensitive.
- Output limits. A max-token cap stops runaway answers, but it only saves money when the answer would otherwise have run longer. Set it per feature, and handle truncated responses.
Make LLM cost visible
Log cost for every call, tagged by feature and model, and aggregate it per successful outcome. Put cost next to latency on the same dashboard. A monthly invoice tells you that you spent too much. Per-feature data tells you where: which feature, which prompt, which step.
Then treat a prompt change that doubles cost the way you'd treat one that doubles latency: something caught in code review, not on next month's invoice.
In production, you aren't really paying for tokens. You're paying for decisions about context, retries and steps, and every one of those decisions lives in your code.