What Is Agentic AI? How AI Agents Actually Work
An agent doesn't just answer. It decides what to do next, acts through tools, checks the result and keeps going until the job is done or it should stop.
Practical AI engineering — how AI systems are designed, built, debugged and shipped. Guides, architecture, experiments and production lessons.
Every build in this room is a variation of this path.
Explore the room
Eight areas of practical AI engineering. Each one filters the builds below.
Latest builds
Experiments, engineering notes and production lessons. Filter by topic, or search above.
Build by problem
Start from what you're building. Each problem opens the builds that solve it, in reading order.
Everything in the room13 builds
An agent doesn't just answer. It decides what to do next, acts through tools, checks the result and keeps going until the job is done or it should stop.
Behind every agent that looks simple is an execution system: a model, tools, context, state, permissions and a loop that knows when to stop.
Chatbots generate answers. Agents decide how to reach a goal. What actually changed in the architecture, and when each approach is the right one.
What Netflix-scale video streaming teaches us about streaming data, from LLM tokens to live events.
The token price didn't change between prototype and production. The architecture did. Where AI app costs really come from, and how to cut them.
A context window is the model's whole workspace for one request. Here's what fills it, why models miss what's buried inside, and how to manage it in real applications.
Lambda is attractive for spiky AI traffic, but timeouts, cold starts and response streaming change how you design the request path.
Token prices are the visible part of the bill. Retries, oversized context, evaluation runs and idle infrastructure are the parts that surprise teams.
Same prompt, different answers. Sampling, prompt drift, model updates and hidden context changes all contribute. Here is how to make output predictable enough to ship.
Ingestion, chunking, embedding, retrieval, reranking, generation and evaluation: a reference architecture for RAG that survives real users.
Non-deterministic output, hidden prompts and multi-step pipelines make AI bugs hard to reproduce. A systematic approach that turns 'it gave a weird answer' into a fix.
You cannot prompt your way out of prompt injection. Limit what a compromised model can do with permissions, isolation and human confirmation.
Caching can cut AI latency and cost dramatically, or serve wrong and leaked answers. What to cache, how to key it and when to invalidate.