Agentic Patterns and the Many Faces of RAG
Translated from the Spanish original. Read in Spanish
When you start building AI-powered apps for the first time, everything looks simple: a prompt, a model, an answer. And it works — until it doesn’t. Until the model hallucinates instead of looking things up, picks the wrong tool with enviable confidence, or treats a complex question exactly like a trivial one.
What you need next is no longer a prompt. It’s a pattern — a recurring way of wiring an LLM into a system so it behaves predictably on real-world input.
I put together an animated reference of these patterns at agents.zerogap.uk. This post is its written companion.

Start with the Single Agent
The starting point I recommend is the simplest one: a model with a system prompt and a defined set of tools. The agent reads the user’s request, decides which tool to call, calls it and synthesises an answer.
You don’t need anything more elaborate until the single agent stops being enough. The signs are concrete: it hallucinates instead of looking things up, picks the wrong tool with confidence, or answers quick and slow questions the same way (badly, for one of them).
That last symptom is what the router gate pattern fixes. A lightweight classifier sits in front of the system and decides: this question goes to the single agent, that one goes to a heavier multi-step pipeline.
RAG Isn’t a Pattern — It’s a Family
The most common reason a single agent isn’t enough is that the answer lives in your documents, not in the model’s weights. That’s where Retrieval-Augmented Generation (RAG) comes in.
Naive RAG. Embed the query, find the nearest chunks by cosine similarity, put them in the prompt. It’s in every tutorial. It works when the corpus is small and well written. It fails when users ask about exact identifiers, codes or jargon — vector similarity smooths them out.
Keyword RAG. BM25 or similar lexical search. The opposite failure mode: excellent for exact-match queries (“error code TRX-4012”), worse at understanding intent.
Hybrid RAG. Run both and combine them. Almost always better than either one on its own in production.
Reranking RAG. Retrieve a broad first pass, then run a cross-encoder reranker over the top candidates. The reranker reads the query and the document together, catching relevance that embedding similarity missed.
Agentic RAG. The agent decides whether to retrieve, what to retrieve and when to stop. Retrieval becomes a tool the model can call repeatedly. Slower, but it handles multi-hop questions Naive RAG can’t.
Corrective RAG (CRAG). The agent retrieves, then grades what it retrieved. If the context is weak, it rewrites the query, falls back to web search or declines to answer.
Graph RAG. Retrieve from a knowledge graph instead of (or alongside) vector chunks. When relationships matter — who reports to whom, which products depend on which suppliers — graph traversal gives answers that pure chunk similarity can’t reconstruct.
Tree-Index RAG (Vectorless). Hierarchical summaries. Large documents are summarised at several levels; retrieval can land on a leaf chunk or climb to a section summary. Useful when context windows are tight.
CAG (Cache-Augmented Generation). When your corpus fits in the context window of a long-context model, skip retrieval and load everything with prompt caching. Often the right answer for narrow domains with stable knowledge bases. Cheaper than people expect once caching is on.
The decision tree isn’t “which one is best” — it’s “which failure mode am I trying to avoid”. Pick the simplest one that passes your hardest query, then add complexity when a real query breaks it.
Orchestration: When One Agent Isn’t Enough
Beyond retrieval, the next family of patterns is about how multiple steps or multiple agents connect to each other.
- Sequential — agent A → agent B → agent C. Each step’s output feeds the next.
- Parallel — fan out independent subtasks, combine the results.
- Conditional — branch on the output of a previous step. Standard control flow applied to LLM calls.
- Review-and-critique — one agent produces, another reviews, the first one fixes. Slower, but it consistently raises quality.
- Coordinator — a manager agent delegates to specialists and assembles their work. The classic “team of experts” shape.
- Hierarchical decomposition — break a big task into subtasks, each of which can be broken down further.
- Dynamic prompting — the prompt is built at runtime from the user’s context, retrieved facts and previous turns. Less a pattern than a discipline: stop hardcoding prompts.
- Swarm — many lightweight agents in parallel with minimal coordination.
See Them in Motion
These patterns are easier to internalise when you can see them. Every one of them at agents.zerogap.uk is animated: you watch the query flow to the agent, the tool calls fire, the retrieval happen, the critique come back.
If you’re building with LLMs and have wondered which pattern your problem really needs — go in and play. The animations make the trade-offs obvious in a way diagrams alone never quite manage.
