Production AIBuild
The Hidden Cost of Running AI Applications in Production
Token prices are the visible part of the bill. Retries, oversized context, evaluation runs and idle infrastructure are the parts that surprise teams.
Input tokens dominate
Teams estimate cost from the length of answers, but most spend is usually input tokens: system prompts, conversation history and retrieved context sent on every turn. A chat that resends a 6,000-token history with each message pays for that history again and again.
Where the money actually goes
- Unbounded conversation history. Summarize or truncate older turns.
- Over-retrieval. Sending 20 chunks "to be safe" multiplies input cost and often hurts answer quality.
- Retries. Timeouts and malformed output trigger retries that double the cost of the worst requests.
- Agent loops. Each tool step is another full model call with the whole context.
- Evaluation and testing. Running a regression suite against a large model on every commit adds up.
Measure cost per feature, not per month
Log input tokens, output tokens, model and feature name for every call. A monthly invoice tells you that you spent too much; per-feature data tells you where. It is common to find that one feature, or one prompt, accounts for most of the bill.
Levers that work
- Prompt caching for long, stable prefixes such as system prompts and tool definitions.
- Model routing: send easy requests to a smaller, cheaper model and reserve the large model for hard ones.
- Context budgets: cap retrieved context and history per request.
- Response caching for repeated, non-personalized questions.
- Output limits: set max tokens deliberately rather than leaving it at the maximum.
Make cost visible to the team
Put cost per request next to latency on your dashboards. Engineers optimize what they can see, and a prompt change that doubles cost should be as visible in review as one that doubles latency.