AI AgentsBuild
Agentic AI Architecture: The Building Blocks of an AI Agent
Behind every agent that looks simple is an execution system: a model, tools, context, state, permissions and a loop that knows when to stop.
Simple on the outside
From the outside, an AI agent looks almost trivial. A user describes a goal, the agent works on it for a while, and an answer comes back.
Inside, a lot more is happening. Something is choosing which model to call and what to show it, which tools it may use, whether an action needs a human's approval, what to remember if the process crashes, and when to stop. That machinery is agentic AI architecture, and it matters far more to whether your agent survives production than any prompting trick.
At the centre of every agent is one loop:
Goal -> Model decides -> Act (tool, retrieval, memory) -> Observe result -> Model decides again
-> Continue / Stop / Ask a human
Everything else exists to make that loop useful, safe and debuggable. Here is the map of components this article walks through:
| Component | Its job | What goes wrong without it |
|---|---|---|
| Model + instructions | Decides the next step | Nothing works |
| Tools | Read from and act on real systems | The agent can only talk |
| Context | What the model sees for this step | Wrong or missing information, wasted tokens |
| Memory and retrieval | Knowledge beyond the current step | The agent forgets, or invents |
| Loop and state | Runs steps, survives failures | Endless loops, lost progress after a crash |
| Guardrails and approval | Limits what the agent may do | Unsafe actions |
| Observability | Records what happened | Debugging becomes guesswork |
| Evaluation | Measures whether it works | You find out from users |
The model: the decision engine
The model reads the goal, its instructions, the available tool descriptions, retrieved information and everything that has happened so far, and produces the next output. That output is either a final answer or a structured request to use a tool:
{ "name": "search_logs", "arguments": { "service": "checkout", "minutes": 30 } }
The exact format varies by provider. The architectural fact doesn't: the model only proposes the call. Your application, or the agent runtime you use, validates it, checks permissions, runs it and returns the result. The model decides what it wants to do; the code around it decides what actually happens.
Tools: how the agent touches the real world
Without tools, an agent can only work with what's already in its context. Tools let it search, query, create tickets, run tests or read metrics. Tools can live in your own code or be exposed by external servers through the Model Context Protocol (MCP), an open standard introduced by Anthropic in 2024.
Tool design is not plumbing. The model chooses tools by reading their names and descriptions, so those descriptions are effectively part of its instructions. Anthropic's engineering guidance on writing tools for agents makes the point directly: agents can call the wrong tool or pass the wrong parameters, and vague or overlapping tools make good choices harder. Your tool API is part of the agent's reasoning environment.
A tool should carry more than a function. It needs a schema the runtime validates, a risk level, a timeout and a result format designed for the model to read:
import { z } from "zod";
type Risk = "read" | "write" | "destructive";
interface Tool<I> {
name: string;
description: string; // the model reads this: write it like good API docs
input: z.ZodType<I>; // validated before anything runs
risk: Risk; // drives permission and approval rules
timeoutMs: number;
run: (input: I) => Promise<string>; // returns a short, model-friendly summary
}
const searchLogs: Tool<{ service: string; minutes: number }> = {
name: "search_logs",
description: "Search error logs for one service over the last N minutes (1-60).",
input: z.object({ service: z.string(), minutes: z.number().int().min(1).max(60) }),
risk: "read",
timeoutMs: 5_000,
run: async ({ service, minutes }) => summarise(await logs.errors(service, minutes)),
};
Note what run returns: a summary, not 128 raw log lines. Every token a tool returns lands in the model's context on the next step.
Context, memory and retrieval
These three get blurred together constantly. They are different components with different jobs.
| What it is | Lifetime | Example | |
|---|---|---|---|
| Context | Everything the model sees on this call | One model call | Instructions, the request, recent tool results |
| Memory | Information stored outside the model for later use | Across steps, runs or sessions | Customer preferences, a saved plan, past findings |
| Retrieval | A way to fetch relevant information into the context | On demand | Searching docs, past tickets or memory |
The model itself remembers nothing between calls. So memory only matters when something retrieves it into the context, and context is the only thing the model can actually use.
Context is a budget, not a bucket. Each step should see what's relevant to the current decision, not everything the agent has ever touched. Agents make this harder, because every tool result adds to the context and the loop keeps running. That's why context engineering matters more for agents than for chatbots; see Understanding Context Windows for the mechanics.
Memory needs a policy. The design question isn't "how do we store everything?" It's "what should persist, for how long, and why?" Storing every intermediate result makes retrieval noisier, not smarter.
Retrieval can be fixed or agentic. Classic RAG retrieves once before generating. An agent can decide to search, read the results, then search again with a better query based on what it learned. RAG supplies information; the agent decides when and how to use it. They complement each other.
The loop: where an agent becomes an agent
The loop is the component that turns a model with tools into an agent. Agent SDKs implement it in broadly the same way: call the model; if it returns a final answer, stop; if it requests tools, run them, add the results and call the model again, until it finishes or hits a turn limit.
The minimal version fits in a few lines. The production version needs much more around each turn: a maximum number of steps, tool timeouts, retries only for transient failures, argument validation, permission checks, human approval for risky actions, cancellation, cost limits and a trace of every step. (The companion article From Chatbots to AI Agents shows a bounded loop in TypeScript.)
State: surviving a crash halfway through
An investigation that takes several minutes might search logs, inspect a deployment, query a database and compare configurations before writing a diagnosis. If the process dies after step three, an agent without durable state starts again from zero, repeating every call and every cost.
Persist the run's state after every step:
interface AgentRun {
runId: string;
status: "running" | "waiting_for_approval" | "completed" | "failed";
step: number;
completed: { tool: string; args: unknown; resultSummary: string; idempotencyKey: string }[];
pending?: { tool: string; args: unknown };
startedAt: string;
updatedAt: string;
}
Two details make resumption safe. A waiting_for_approval status lets a run pause for a human for hours and continue later, which in-memory loops can't do. And an idempotency key on every action with side effects means that if the crash happened mid-call, resuming won't create the ticket twice or send the email again.
Long-running agents need explicit state, the same way long-running jobs in any other backend do.
Guardrails and permissions
Give an agent readDatabase, writeDatabase, sendEmail, deleteUser and restartService, and ask whether the model should be able to call all five without anyone checking. The answer is obviously no, and the reason is architectural: the model must never be the final authority on what is allowed. Instructions in a prompt are guidance; permissions belong in code.
Using the risk level from the tool definition, every call passes through one gate:
async function execute(tool: Tool<any>, rawArgs: unknown, user: User): Promise<string> {
const args = tool.input.parse(rawArgs); // malformed arguments never run
if (!user.can(tool.name)) return "Not permitted for this user.";
if (tool.risk === "destructive" && !(await requestApproval(user, tool.name, args))) {
return "Rejected by approver.";
}
return withTimeout(tool.run(args), tool.timeoutMs);
}
Notice user.can. The agent acts on behalf of a user, so it should never hold more access than that user has. An agent with admin credentials serving a regular user is a privilege escalation waiting for the right prompt.
That risk is real. Retrieved documents, web pages and tool results can contain text written to look like instructions, which is known as indirect prompt injection. Treat everything that enters the context from outside as untrusted, and rely on code-level permissions, not the model's judgement, to contain the damage. Guardrails can also check the input before the agent starts and the output before it reaches the user.
Human in the loop
Some actions should wait for a person, however confident the model is:
| Action type | Policy |
|---|---|
| Read | Automatic |
| Write | Policy check, logged |
| Destructive or high-impact | Human approval |
An infrastructure agent that finds an unhealthy server can investigate and prepare the restart; a human approves it. The agent investigates and recommends; people control the high-risk actions. This boundary matters most for production infrastructure, payments, data deletion, customer communication, security changes and deployments.
Observability: making every run reconstructable
When a normal API fails, you read the logs and find the bad line. When an agent fails, the bug is usually a decision: it chose the wrong tool, passed a bad argument, misread a result or stopped too early. To debug that, you need the whole run, not just errors:
Run 8f31 - "Investigate checkout errors" total 9.1s, 3 model calls
09:01:02 run started (user u_42)
09:01:03 model call 1 -> search_logs(service=checkout, minutes=30)
09:01:05 tool result 128 errors, 121 match "timeout calling payments"
09:01:06 model call 2 -> get_deployment(service=payments)
09:01:08 tool result deploy #842 at 08:52, config change: pool size 50 -> 5
09:01:09 model call 3 -> final answer
09:01:11 run completed stop reason: model finished
A useful trace records, for every step: what the model saw, what it decided, the tool and arguments, what came back, latency, tokens and cost, and why the run stopped. Without it, debugging an agent is guesswork.
Evaluation: testing decisions, not just outputs
Traditional tests assert exact outputs. An agent can reach a correct answer by several valid paths, or a wrong one by a plausible path, so you need to evaluate both the result and the trajectory:
- Outcome: was the task completed, and is the final answer correct?
- Trajectory: were the right tools used, with valid arguments, without detours?
- Safety: did it stay within permissions and request approval when required?
- Efficiency: how many steps, how long, how much did it cost?
- Resilience: what did it do when a tool failed?
Build a small set of realistic tasks with expected behaviour ("for this incident, it should inspect logs, find the deployment and identify the config change") and run them whenever you change a prompt, a tool description or the model. Evaluation belongs in the architecture from the start, not after the first production failure.
Single agent or multi-agent?
Once one agent grows complicated, splitting it into a supervisor with specialist agents is tempting:
Supervisor
/ | \
Research Code Security
agent agent agent
It can work well. Anthropic reported in 2025 that its multi-agent research system, a lead agent delegating to parallel subagents, outperformed a single agent by 90.2% on its internal research evaluation, especially for broad questions that split into independent directions.
The same report shows the price. In Anthropic's data, agents used about 4x more tokens than chat interactions, and multi-agent systems about 15x more. It also noted that most coding tasks involve fewer truly parallelisable parts than research, and that agents were not yet good at coordinating with each other in real time.
So the trade-off is concrete: more agents can mean better results on genuinely parallel, breadth-first work, in exchange for more tokens, more latency, more state to coordinate and more ways to fail. OpenAI's practical guide to building agents gives the same advice from the other direction: get as much as possible out of a single agent before splitting it into several. Start with one well-equipped agent, and split only when the evidence says the work divides cleanly.
Putting it together
mermaid
flowchart TD
U[User] --> API[Agent API<br/>auth, user identity]
API --> O[Orchestrator<br/>loop, step limits, state]
O --> M[LLM<br/>decides next step]
M -->|tool request| G[Guardrails<br/>validate, permissions]
G -->|read or write| T[Tools]
G -->|destructive| H[Human approval]
H --> T
M -->|needs knowledge| R[Retrieval and memory]
T --> OB[Observation]
R --> OB
OB --> O
O -->|finished| Resp[Response]
O -.-> TR[(Traces and state store)]This isn't a framework. It's a map of responsibilities. Frameworks combine or rename these boxes, but every production agent needs something doing each job: deciding, acting, limiting, remembering, recording and stopping.
Build the smallest version first
It's tempting to start with multiple agents, RAG, long-term memory, MCP servers, twenty tools and long-running execution, all before proving the problem needs any of it. Start here instead:
One model + 2-3 well-described tools + clear instructions
+ a bounded loop + permission checks + a trace of every run
Then add memory, retrieval, more tools or more agents when evaluation shows a specific gap. That matches the guidance both Anthropic and OpenAI publish: use the simplest architecture that works, and add complexity only when the task benefits from it.
The design questions that matter more than the framework
Don't start with "which agent framework should I use?" Start with "what does this system need to decide dynamically?" Then:
- What is the goal, and what does "done" look like?
- What information does each decision need?
- Which tools are required, and what is each one's risk level?
- Which actions need human approval?
- Whose permissions does the agent act with?
- What state must survive a crash?
- What makes the loop stop?
- How are failures recovered?
- How is every run traced?
- How will you know it works?
Answer those and the framework choice becomes a detail.
The engine and everything around it
At the centre of every agent is a small loop: decide, act, observe, decide again, stop. On its own it's a demo. The model, tools, context, memory, state, guardrails, approvals, traces and evals around it are what make it something you can put in front of users.
Good agent architecture isn't about giving AI more autonomy. It's about giving it exactly the capabilities, context and permissions a task needs, and nothing more.