AI EngineeringBuild

How to Debug an AI Application

Non-deterministic output, hidden prompts and multi-step pipelines make AI bugs hard to reproduce. A systematic approach that turns 'it gave a weird answer' into a fix.

By The AI Build Room2 min read90 views

Start by capturing the exact request

"The bot gave a weird answer" is not reproducible. You need the exact final prompt, model and version, parameters such as temperature, tool calls and their results, and the raw response. If you are not logging these per request, that is the first bug to fix.

Reproduce with everything pinned

Replay the captured prompt against the same model version at temperature 0. If the bad output reproduces, you have a deterministic bug to work on. If it does not, run it 20 times and measure how often it fails; you are now dealing with a reliability problem rather than a single bug.

Bisect the pipeline

AI applications are pipelines. Check each stage's output in order:

  1. Was the user input parsed and routed correctly?
  2. Was the right context retrieved?
  3. Did the assembled prompt contain what you expected, in the order you expected?
  4. Did tool calls receive valid arguments and return the expected results?
  5. Did the model's raw output contain the error, or was it introduced by post-processing?

Most "model" bugs turn out to be in steps 1 to 3 or step 5.

Read the prompt like the model does

Print the fully rendered prompt and read it top to bottom. Contradictory instructions, an empty variable rendered as "undefined", or an example that accidentally teaches the wrong format are common and easy to miss in templates.

Turn every fix into a test

When you fix a failure, add the input to your evaluation set with the expected behavior. Over time this set becomes your regression suite, and the same bug cannot quietly return after the next prompt change or model upgrade.