AI AgentsBuild
What Is Agentic AI? How AI Agents Actually Work
An agent doesn't just answer. It decides what to do next, acts through tools, checks the result and keeps going until the job is done or it should stop.
Explaining the investigation vs doing it
For most of the short history of LLM applications, the pattern was simple: the user asks, the model answers. It's still the right pattern for a lot of problems.
Now ask an AI system: "Checkout started returning 500 errors after the last deploy. Find out why."
A chatbot will explain how you might investigate: check the logs, look at the deployment, compare configurations. Good advice, and you still have to do all of it yourself.
An AI agent can do the investigation. It can read the logs, see that the errors are database connection failures, look up what changed in the last deployment, find a modified connection setting, compare it with the previous version and report the likely cause, deciding each step based on what the previous one revealed.
That shift, from generating a response to taking controlled actions toward a goal, is what agentic AI means. This article explains what that actually involves, how the loop at its centre works, and where the hype runs ahead of the engineering.
What is agentic AI?
There is no single, universally agreed technical definition. In practice, "agentic AI" describes systems that pursue a goal over multiple steps, usually by calling tools, keeping track of what has happened, and adapting what they do based on the results.
The two most cited engineering guides are close to each other:
- OpenAI (A practical guide to building agents): an agent uses an LLM to manage workflow execution and make decisions, and has access to tools to gather context and take actions. Applications that use an LLM but don't let it control workflow execution, such as simple chatbots or classifiers, are not agents.
- Anthropic (Building effective agents): agents are systems where LLMs dynamically direct their own processes and tool usage, as opposed to workflows, where code orchestrates the LLM through predefined paths.
Combining both into one engineering definition:
An agentic AI system uses a model to decide what to do next, carries out that action through tools, observes the result, and repeats until it reaches a defined stopping condition.
The key word is decide. A standard LLM application turns input into output in one step. An agent runs a loop:
mermaid
flowchart TD
G[Goal] --> D[Model decides next action]
D -->|tool call| A[Application runs the tool]
A --> O[Result added to context]
O --> D
D -->|done| F[Final answer]
D -->|limit reached or approval needed| S[Stop or ask a human]That loop is the single most important idea in agent engineering. Everything else in this series builds around it.
How AI agents actually work: one run, step by step
Take the production debugging agent from the opening. It has three read-only tools: get_deployments, search_logs and get_config. Here is what actually travels between the model and the application:
GOAL Why did checkout start returning 500 errors after the latest deployment?
MODEL -> get_deployments(service="checkout", last=2)
APP <- #847 at 14:32 (current), #846 at 09:10
MODEL -> search_logs(service="checkout", since="14:32")
APP <- 412 errors, 398 are "connection refused: db-primary:5432"
MODEL -> get_config(service="checkout", deployment=847)
APP <- DATABASE_URL host changed to db-primary-new
MODEL -> get_config(service="checkout", deployment=846)
APP <- DATABASE_URL host was db-primary-old
MODEL -> final answer: deployment 847 changed the database host; errors began
at 14:32 and are connection failures to the new host. Recommend rolling
back to 846 or verifying the new host is reachable.
Three things are worth noticing.
Nobody wrote this sequence. The model chose to look at the logs because of the deployment timing, and chose to compare configurations because the logs showed database errors. A different incident would produce a different path.
The model never touched a system. Each line marked MODEL is a structured request. Each APP line is your code running the tool and returning the result, which gets added to the model's context for the next decision.
The model decides when to stop. It ends the run with an answer once the evidence is enough, though in production your code also enforces a hard limit, in case it never gets there.
That is all an agent is mechanically: a model producing decisions inside an execution loop, with an application doing the acting.
An agent is more than an LLM and a prompt
The run above needed five things working together:
| Component | Role in the run above |
|---|---|
| Model | Chose each next step from the evidence so far |
| Instructions | Defined its job and limits: investigate, read-only, report findings |
| Tools | get_deployments, search_logs, get_config: its only way to see the system |
| Context and state | The goal plus every result so far, rebuilt for each model call |
| Orchestration | The loop: run the requested tool, feed back the result, enforce limits, stop |
Two points from that table cause most beginner mistakes.
Instructions describe limits; code enforces them. An instruction like "never restart production services" guides the model, but the reliable protection is not giving the agent a restart tool, or requiring human approval in code before it runs.
Context is not memory. The model remembers nothing between calls. The application decides what goes into each call, which is why context engineering matters so much for agents: every tool result makes the context bigger.
Each component gets a full treatment in Agentic AI Architecture: The Building Blocks of an AI Agent.
What makes a system "agentic"?
Agentic systems usually share five traits:
- Goal-oriented: it's given an objective, not just text to complete.
- Tool use: it can read from or act on external systems.
- Dynamic decisions: the next action depends on previous results.
- Iteration: it cycles through decide, act, observe, decide again.
- Explicit stopping conditions: it knows when to stop.
The last one is the most neglected. A reliable agent stops when the task is complete, when it hits a maximum number of steps, when required information is unavailable, when an action needs human approval, or when a tool failure can't be recovered. Without those conditions, an agent can call tools in circles, burning tokens and money while making no progress.
Agents vs ordinary automation
Traditional automation follows a path you wrote: trigger, step A, step B, step C, done. An agent chooses its path at runtime.
Neither is better in general. If you know the exact sequence of steps, ordinary code is faster, cheaper, easier to test and easier to trust. Anthropic's guidance is explicit on this: start with the simplest solution, and add agentic behaviour only when the task needs it, because agentic systems often trade latency and cost for better task performance.
Don't use an agent when deterministic software is enough.
The question isn't "can I make this agentic?" It's "does dynamic decision-making add enough value to justify the extra complexity?" (The full comparison of chatbots, RAG, workflows and agents is in From Chatbots to AI Agents.)
Do AI agents really "think"?
We describe agents with human words: they reason, plan, think, decide. Those words are convenient, and they can mislead you into designing the wrong system.
From an engineering point of view, it's more useful to describe what actually happens: on each step, the model receives a context and produces an output, which is either text or a structured tool request. Some models also generate intermediate reasoning before answering, which often improves decisions on multi-step tasks. But the agent only "knows" what is in its current context. It doesn't see your whole environment, and it has no understanding beyond what the tools return to it.
That framing has practical consequences. If the agent made a bad decision, the first questions are concrete ones: what was in its context, what did the tool return, how was the tool described? Those are things you can inspect and fix.
Agentic doesn't mean fully autonomous
"Agentic" is often read as "completely autonomous". Production agents almost never are, and shouldn't be. Autonomy is a dial you set per action:
| Risk of the action | Example | Policy |
|---|---|---|
| Low | Read logs, search docs | Agent acts automatically |
| Medium | Create a ticket, draft an email | Agent prepares, policy or human checks |
| High | Restart production, issue a refund, delete data | Human approval required, or the agent can't do it at all |
This matters most when an agent can change production systems, send messages, touch sensitive data or move money. Agent SDKs now ship guardrails and approval mechanisms for exactly this reason, but the policy is yours to design. The goal is controlled autonomy, not unlimited autonomy.
Where agentic AI earns its complexity
Agents fit tasks that are multi-step, hard to script in advance, dependent on live information, and different from one request to the next:
| Domain | What the agent does | Why a fixed script struggles |
|---|---|---|
| Software engineering | Reads code, edits it, runs tests, reads failures, fixes again | The next step depends on which tests fail |
| Customer support | Finds the customer and order, checks policy, resolves or escalates | Each case combines different facts and exceptions |
| DevOps and incidents | Checks metrics, searches logs, inspects deployments, recommends a fix | The cause could be anywhere until evidence narrows it |
| Research | Searches, reads, compares, spots gaps, searches again, reports | What to search next depends on what was found |
Coding is a particularly good fit because results are verifiable: tests either pass or they don't, which gives the agent clear feedback inside the loop.
And where agents are the wrong tool: anything with a fixed sequence (invoice processing, data pipelines, scheduled reports), anything a single model call can answer, and anything where an unpredictable path is unacceptable and the steps can be written down.
The real engineering challenge
A demo agent takes an afternoon. A reliable production agent takes real engineering, because sooner or later you must answer questions like these:
- What happens when it picks the wrong tool, or a tool returns bad data?
- How many steps is it allowed, and what happens when it gets stuck?
- What happens when an API times out halfway through a run?
- Whose permissions does it act with, and which actions need approval?
- How do you stop a retrieved document from injecting instructions into it?
- How do you reproduce a failed run, and what should you log?
- How do you measure whether it succeeded, and what does each task cost?
That's why agentic AI is an engineering discipline, not a prompting technique. The prompt is a small part; the loop, the tools, the limits and the observability around it are most of the work.
The mental model to keep
Forget the idea of an agent as a smarter chatbot. Think of it as a loop:
Goal -> decide -> act -> observe -> decide again -> ... -> stop
Everything else you'll hear about (memory, RAG, MCP, planning, multi-agent systems, guardrails, evaluation, human approval) is built around that loop to make it useful, safe and debuggable.
Agentic AI isn't an LLM that gives better answers. It's an application architecture in which a model helps decide what should happen next, acts through tools, checks the results and keeps working toward a goal, inside limits you define.
Once that loop makes sense, the next question follows naturally: what components do you need to make it reliable?
An agent isn't defined by how autonomous it sounds. It's defined by how well, and how safely, it gets from a goal to controlled actions.