AI EngineeringBuild

The Hidden Cost of Running AI Applications in Production

The token price didn't change between prototype and production. The architecture did. Where AI app costs really come from, and how to cut them.

By Faizan Syed6 min read56 views

The bill nobody planned for

In development, the assistant cost almost nothing. A few engineers, a few hundred test questions, a rounding error on the invoice.

Then it shipped. The system prompt had grown to a few thousand tokens and went out with every request. Conversations got longer, and the full history travelled with each turn. Retrieval pulled twenty chunks "to be safe". The agent made three or four model calls per question. A flaky upstream service triggered automatic retries. A nightly eval suite fired hundreds of model calls. Even the logs, full of prompts and responses, became a line item.

The price per token never changed. The architecture did. That is the real story behind LLM cost in production.

The model invoice is only the beginning

A prototype costs roughly one model call per question. A production system costs everything around that call: embeddings, vector search, tool executions, retries, evaluations, the compute that hosts it and the observability that watches it.

A useful mental model, not a billing formula:

text
AI system cost ~ model inference + retrieval and embeddings + tool calls
               + retries + evaluation + infrastructure and observability
               + wasted context

The metric that exposes it is cost per successful outcome: the total cost of every call, retry and tool step needed to complete one user task. Cost per request can look fine while retries and agent loops quietly multiply it.

Where the cost actually comes from

Cost driverWhat creates itEngineering lever
Input tokensPrompts, history, retrieved docs and tool results sent on every callContext budgets, prompt caching
Output tokensLong answers, verbose formats, reasoning tokens on some modelsPer-feature output limits, concise formats
EmbeddingsEmbedding every query, re-embedding documentsCache embeddings, re-embed only changed docs
RetrievalVector database queries, storage, rerankersRight-size topK, filter before searching
Tool callsExternal APIs and databases called by agentsTimeouts, caching, trimmed results
Agent loopsSeveral model calls per user taskIteration caps, smaller models for steps
RetriesTimeouts, rate limits, malformed outputClassify errors, cap retries
EvaluationSuite size times run frequency times model sizeTiered suites, batch processing
InfrastructureCompute, serverless invocations, storage, data transferRight-sizing, scaling to zero
ObservabilityLogging full prompts, responses and tracesSampling, retention limits

Context is a recurring cost

As covered in Understanding Context Windows, the model keeps nothing between calls. The system prompt, the conversation so far and the retrieved documents are sent, and paid for, on every turn.

Here's the multiplication effect with hypothetical prices of $3 per million input tokens and $15 per million output tokens. Assume 10,000 input tokens and 2,000 output tokens per request, and 10,000 requests a month:

text
Input:   10,000 x 10,000 = 100M tokens x $3/M  = $300
Output:   2,000 x 10,000 =  20M tokens x $15/M = $300
Total                                          = $600 / month

Cut input context to 4,000 tokens:
Input:   40M tokens x $3/M = $120   ->   total $420 (30% lower)

Notice that with these prices, input and output cost exactly the same. That's why "input tokens dominate" isn't a rule. Long contexts with short answers skew toward input; long generations and reasoning models skew toward output. Plug in your provider's real prices and your real token counts before deciding what to optimise.

RAG doesn't make context free

One RAG question is really a pipeline:

text
query -> embedding -> vector search -> metadata filter -> rerank -> build context -> LLM

Each step can cost something: an embedding call, a query against a vector database you pay to run, a reranker, and the extra input tokens the retrieved text adds to the model call.

That's why topK = 20 isn't automatically better than topK = 5. Twenty chunks mean more input tokens, more latency and more noise the model has to ignore. But five weak chunks can miss the answer, and a wrong answer costs a follow-up question, which is another full request. The right value depends on your retrieval quality and the task. Choose it with evals on real questions, not intuition.

Agent loops multiply calls

text
request -> LLM -> search tool -> LLM -> database tool -> LLM -> answer

To the user, that's one question. To your bill, it's three model calls and two tool executions. How many tokens each step uses depends on how your application builds each request: full history every time, trimmed history or a summary. Either way, cost now scales with the number of steps, and the model decides that number at runtime.

Controls that keep it bounded:

  • a maximum number of iterations per task
  • a timeout on every tool
  • smaller, cheaper models for intermediate decisions
  • tool results trimmed to what the next step needs
  • early termination once the answer is good enough
  • model calls per task tracked as a metric, with alerts

Retries: when reliability turns into cost

Retries are good engineering. Networks fail, providers rate-limit, and a transient 503 shouldn't become a user-facing error. The problem is the combination: unbounded retries, on large requests, to expensive models. A request that times out on your side may still have been processed, and billed, on the provider's side, so one user action can quietly become three paid calls.

The rules:

  • Retry only transient failures: timeouts, rate limits and 5xx errors.
  • Never retry validation or authentication errors. They will fail again.
  • Use exponential backoff with jitter, and respect rate-limit headers when the provider sends them.
  • Cap the number of retries.
  • Make operations with side effects idempotent.

A provider-neutral wrapper that handles both retries and per-feature cost logging:

ts
type Usage = { inputTokens: number; outputTokens: number };
type Price = { inputPerM: number; outputPerM: number }; // from config, never hardcoded

const RETRYABLE = new Set([408, 429, 500, 502, 503, 504]);
const sleep = (ms: number) => new Promise(r => setTimeout(r, ms));

export async function callLLM<T extends { usage: Usage }>(
  feature: string, model: string, price: Price,
  call: () => Promise<T>, maxRetries = 2,
): Promise<T> {
  for (let attempt = 0; ; attempt++) {
    try {
      const res = await call();
      const cost = (res.usage.inputTokens * price.inputPerM
                  + res.usage.outputTokens * price.outputPerM) / 1_000_000;
      log.info("llm_call", { feature, model, attempt, ...res.usage, cost });
      return res;
    } catch (err: any) {
      const status = err?.status;                              // undefined = network error
      const retryable = status === undefined || RETRYABLE.has(status);
      log.warn("llm_call_failed", { feature, model, attempt, status });
      if (!retryable || attempt >= maxRetries) throw err;      // 400, 401, 422: fail fast
      await sleep(2 ** attempt * 500 + Math.random() * 250);   // backoff with jitter
    }
  }
}

Logging failed attempts matters as much as logging successes. They are exactly the calls that never show up in a "cost per request" view.

Evaluation: control it, don't cut it

Eval cost grows with suite size, run frequency, model size, prompt length and the number of prompt or model variants you compare. Deleting evals saves that money and ships the regressions instead.

Control it instead: run a small smoke suite on every commit and the full suite nightly or before release, use a cheaper judge model where quality allows, skip cases whose inputs haven't changed, and use asynchronous batch processing where your provider offers it at a discount.

Levers that actually work

  • Prompt caching. Some providers can reuse repeated, identical prompt prefixes and bill them at a lower rate or serve them faster. Availability, eligibility rules, pricing and cache lifetime vary by provider and model. It only helps when the prefix is truly identical, so put stable content such as instructions and tool definitions first, and anything volatile last.
  • Model routing. Send easy requests to a cheaper model. It needs a reliable routing signal and evals proving the cheap path is good enough.
  • Context budgets. Cap history and retrieved context per request, per feature.
  • Response caching. Reuse answers for repeated questions, but only when they are not personalised and not sensitive.
  • Output limits. A max-token cap stops runaway answers, but it only saves money when the answer would otherwise have run longer. Set it per feature, and handle truncated responses.

Make LLM cost visible

Log cost for every call, tagged by feature and model, and aggregate it per successful outcome. Put cost next to latency on the same dashboard. A monthly invoice tells you that you spent too much. Per-feature data tells you where: which feature, which prompt, which step.

Then treat a prompt change that doubles cost the way you'd treat one that doubles latency: something caught in code review, not on next month's invoice.

In production, you aren't really paying for tokens. You're paying for decisions about context, retries and steps, and every one of those decisions lives in your code.