AI AgentsBuild

From Chatbots to AI Agents: How LLM Applications Became Autonomous

Chatbots generate answers. Agents decide how to reach a goal. What actually changed in the architecture, and when each approach is the right one.

By Faizan Syed10 min read55 views

A chatbot can tell you. An agent can try.

Ask a chatbot why your payment API is returning 500 errors, and you get a clear list of common causes. Useful, and completely disconnected from your system. It hasn't seen your logs, your last deployment or your metrics.

Ask an agent the same question, and it can read the logs, notice a spike after the last deploy, check what changed in that release and come back with a likely cause.

The difference isn't a smarter model. It's the software around the model. The debate about AI agents vs chatbots is really about how much of the path from question to answer you let the model decide.

It helps to see that difference as a set of layers, each adding one capability:

text
Chatbot -> LLM application -> RAG -> Tool calling -> Workflow -> Agent

One correction to the usual story: this is a layering of capabilities, not a strict timeline. Retrieval-augmented generation was described in research in 2020, before ChatGPT launched in late 2022, and the ReAct research pattern of models interleaving reasoning with actions appeared in 2022, before major providers shipped native tool-calling APIs in 2023. The pieces arrived in a messier order. What matters for engineers is what each layer adds, and what it costs.

Stage 1: The chatbot

text
User -> LLM -> Response

The simplest LLM application sends the user's message, plus the conversation so far, to the model and shows the reply. Ask it to explain Kubernetes deployments and it will, well.

Its knowledge has two sources: what the model learned in training, up to its training cutoff, and whatever is in the current conversation. It has never seen your database, your production logs, your internal docs or your ticketing system. It can talk about those systems; it can't look at them, and it can't change anything.

Stage 2: The LLM application

text
User -> Application -> Prompt builder -> LLM -> Validation -> Response

The first real architectural shift was putting code around the model. The application now decides the system instructions, which context and history to include, which model to call, and how to check and format the output before the user sees it.

This is the moment the LLM stops being the application and becomes one component inside it. Everything that follows, from RAG to agents, is a question of how much more you build around that component, and how much control you hand to it.

Stage 3: Retrieval-augmented generation

text
Question -> Retrieve relevant documents -> LLM (question + documents) -> Answer

The next question every team asked was: how do I make the model answer from my data? RAG answers it by searching your documents at request time and putting the relevant passages into the context. Ask "what is our refund policy?" and the model answers from the policy page, not from general knowledge.

RAG adds knowledge, but not control. The flow is fixed by the application: retrieve, then generate, every time.

Stage 4: Tool calling

Tool calling lets the model ask the application to do something. You describe the available tools (name, purpose, parameters), and instead of answering in text, the model can return a structured request:

json
{ "name": "get_order", "arguments": { "orderId": "12345" } }

The exact format differs between providers, but one fact doesn't: the model never executes anything. It proposes a call. Your application decides whether to run it, runs it, and sends the result back so the model can continue.

text
User -> LLM -> tool request -> Application runs the tool -> result -> LLM -> Response

This is the step that lets an LLM application read live systems instead of only talking about them. But tool calling on its own doesn't make something an agent, which brings us to the distinction that matters most.

Stages 5 and 6: Workflows vs agents

This is where most explanations get blurry, so it's worth borrowing precise definitions.

Anthropic's engineering guide Building effective agents (December 2024) groups both under "agentic systems" but separates them by who controls the path. In its words, workflows are systems where LLMs and tools are "orchestrated through predefined code paths", while agents are systems where LLMs "dynamically direct their own processes and tool usage".

OpenAI's A practical guide to building agents draws the same line from the other side: an agent uses an LLM to manage workflow execution and make decisions, and applications that integrate LLMs but don't use them to control workflow execution, such as simple chatbots or sentiment classifiers, are not agents.

So the test is not "does it use tools?" It's "who decides the next step: your code or the model?"

A workflow can include several model calls and tools, but the sequence is written by you:

text
Classify request (LLM) -> Retrieve docs -> Analyse (LLM) -> Validate -> Draft reply (LLM)

An agent receives a goal and a set of tools, and the model chooses the next action based on what it has observed so far:

text
Goal: "Investigate why the payment API is failing"

search incidents -> read payment logs -> (spike after deploy) -> check deployment diff
  -> (config change found) -> inspect config -> write report

Run it again on a different incident and the path changes, because the evidence changes. That flexibility is the whole point, and also the whole problem: you can no longer list every path the system might take.

Which is why Anthropic's guide recommends starting with the simplest solution and adding complexity only when needed. A workflow is easier to test, debug, secure, cost-estimate and monitor. Autonomy is a trade-off, not an upgrade.

One request through every stage

A customer writes: "My order is late. Figure out what happened and tell me my options."

ArchitectureWhat it can do with this request
ChatbotExplain in general terms why orders get delayed, and suggest checking the tracking number
LLM applicationThe same, but in your brand's tone, with your business rules and output checks
RAGQuote your actual shipping and delay policy
Tool callingLook up this customer's real order status
WorkflowRun a fixed sequence: get order, check delivery status, check delay policy, draft a reply
AgentInvestigate: check the carrier, see a regional disruption, find the order was partly refunded, adjust its answer, and list only the options that actually apply

The agent's advantage shows up when the situation doesn't fit the script. If the carrier API is down, it can try another source. If the order turns out to be already refunded, it changes direction. The path depends on what it observes.

But notice the workflow row. For the large majority of "where is my order?" messages, the fixed sequence is faster, cheaper and easier to trust. A practical pattern is a workflow that handles the common path and hands only the unusual cases to an agent.

Choosing the right architecture

The newest layer is not automatically the right one. Pick the simplest architecture that solves the problem:

If the task is...UseExample
General knowledge, explanation, writingChatbot or LLM application"Explain recursion in JavaScript"
Answering from your own documentsRAG"What does our engineering handbook say about production deployments?"
Reading or changing data in one known systemTool calling"What's the status of order 12345?"
A process whose steps are known in advanceWorkflowExtract invoice, validate fields, calculate tax, store it
A goal whose steps depend on what you discoverAgent"Investigate this production incident and find what changed"

The last row is the only one where an agent clearly earns its complexity. An incident might come from a deployment, a database, a configuration change, an external API, a dependency or a traffic spike, and you don't know which until you look. When the right sequence can't be written down in advance, letting the model choose it becomes valuable.

Controlled autonomy, not unlimited autonomy

"Autonomous" sounds like "the AI can do whatever it wants". Production agents should work very differently. The model decides what it wants to do next; your application decides what it is allowed to do.

At its core, every agent is a loop: the model picks an action, the application executes it, the result goes back to the model, and it decides again until it stops. Here's that loop with the boundaries a production system needs:

ts
type Action =
  | { type: "tool"; name: string; args: unknown }
  | { type: "finish"; answer: string };

async function runAgent(goal: string, maxSteps = 8): Promise<string> {
  const history: string[] = [`Goal: ${goal}`];

  for (let step = 0; step < maxSteps; step++) {      // stopping condition 1: step budget
    const action: Action = await decideNextAction(history); // the model proposes
    if (action.type === "finish") return action.answer;     // stopping condition 2: model is done

    if (!isAllowed(action)) {                               // authorisation is code, not a prompt
      history.push(`Denied: ${action.name} is not permitted`);
      continue;
    }
    if (needsApproval(action) && !(await askHuman(action))) { // risky actions wait for a person
      history.push(`Rejected by human: ${action.name}`);
      continue;
    }

    const result = await runToolSafely(action);            // timeouts, error capture, trimmed output
    history.push(`Observed from ${action.name}: ${result}`);
    trace(step, action, result);                           // every step is reconstructable later
  }
  return "Stopped: step limit reached. Escalating to a human.";
}

The model never touches a system directly. Every action passes through permission checks, optional human approval, timeouts and tracing, and the loop always ends.

Why agents are harder than chatbots

With a chatbot, the main question is: was the answer good? With an agent, you also have to ask:

  • Did it choose the right action, and the right tool for it?
  • Did it use the tool safely, with valid arguments?
  • Did it stop at the right time, or burn fifteen model calls going in circles?
  • What happened when a tool failed?
  • Can I reproduce a failed run step by step?
  • Should a human have approved that action?

Answering those means the agent needs everything a serious backend service needs: observability, evaluation, authorisation, state management, error handling, cost limits, clear tool contracts and explicit stopping conditions. Agent frameworks have started building these in. OpenAI's Agents SDK, for example, includes guardrails, sessions and tracing, but the responsibility for using them well stays with you.

The problems that come next

Once an agent works in a demo, a new set of engineering problems appears:

  • Tool selection: how does it choose well between twenty tools?
  • Context: how much information should each step see?
  • Memory: what should persist between runs, and what should be forgotten?
  • Security: what happens when a retrieved document or tool result contains instructions designed to manipulate it?
  • Evaluation: how do you know it actually completed the task correctly?
  • Cost: how do you stop a simple request from turning into dozens of model and tool calls?

Those questions define the next generation of AI engineering, and each gets its own article in this series.

From answers to actions

The story of chatbots becoming agents isn't really a story about models getting smarter. It's a story about software architecture expanding around the model:

ArchitectureWhat you're really asking
ChatbotTell me something.
RAGTell me something, using my information.
Tool callingTell me something after checking my systems.
WorkflowPerform these steps for me.
AgentAchieve this goal with the capabilities you have.

The last step moves us from asking AI to generate an answer to building systems where AI decides how to reach a goal. That's powerful, but autonomy should never be the goal in itself. The goal is a system that handles tasks too varied to script in advance, while staying observable, bounded, reliable and safe.

Don't build an agent because agents are popular. Build one when the problem itself requires dynamic decisions.