AXLE · RAG & Context Engineering
Home / Week 7 / Study material
Week 7 · Study Material

Agentic RAG: Retrieval-Aware Workflows

Your Week 6 router made one decision, once, with a fixed script behind it. This week the model takes the wheel: it decides whether to search, what to search for, whether what came back is enough, and whether to go again. The honest finding this week is often that the agent loses — and knowing exactly when is the professional skill.

Part 1 — Retrieval as a tool, not a pipeline stage

Every system you've built so far shares a shape: always retrieve, once, then answer. The query is a passive input; the pipeline is a fixed track. An agent inverts the relationship — retrieval becomes a tool the model chooses to invoke, with arguments it writes itself.

That single change unlocks four behaviors your pipeline can't express:

The cost is control. A pipeline's behavior is inspectable and its latency is predictable; an agent's is neither. You are trading determinism for adaptability — which is exactly why this week ends with a measurement, not a celebration.

Part 2 — The ReAct loop

ReAct (Reason + Act) is the pattern nearly all agent frameworks implement, usually with far more ceremony than it needs. The loop:

┌─────────────────────────────────────────────┐
│  THOUGHT   what do I know? what's missing?  │
│  ACTION    search("...")  or  answer("...") │
│  OBSERVE   results appended to the trace    │
└──────────────── repeat, max N ──────────────┘

Three implementation realities the tutorials gloss over:

The trace is the state. There's no hidden memory — the accumulated thought/action/observation text is everything the model knows. Manage that text and you manage the agent.

The stop condition is load-bearing. Without a hard iteration cap, agents loop: search, get the same results, decide they're insufficient, search again. Always cap (3–5 hops), and always define a fallback answer when the cap is hit.

Small models drift from output formats. An 8B model will eventually emit prose where you expected ACTION: search(...). Parse defensively and treat unparseable output as "answer now" rather than crashing. Robustness here is most of the engineering.

Part 3 — Self-correction: draft, critique, re-retrieve

A second, complementary loop. Instead of deciding before answering whether context suffices, the agent answers, then audits itself:

  1. Draft an answer from the retrieved context.
  2. Critique it: which claims aren't supported by the sources? (You already built this in Week 5 — it's your faithfulness judge, pointed inward.)
  3. Re-retrieve for the unsupported claims specifically.
  4. Revise with the enlarged context.

This measurably raises faithfulness and roughly doubles cost and latency. Two warnings worth stating to students plainly: models are weaker at critiquing their own output than others' (self-preference bias, from Week 5), and more than one revision cycle rarely pays. One cycle, measured — not a philosophy.

Part 4 — Context engineering for agents

Loops accumulate text fast: four hops × four chunks each = sixteen chunks in the window, most of them from queries that turned out to be wrong turns. Left unmanaged, the agent drowns in its own history — and "lost in the middle" (Week 6) hits hardest at exactly the moment the agent needs to reason over everything it gathered.

Three strategies, in increasing order of effort:

StrategyHowTrade-off
WindowingKeep only the last N observations verbatimSimplest; can drop the one useful early hit
SummarizingCompress old observations into a running noteCheap on tokens; lossy — details vanish and can't be cited
ScratchpadExtract facts found so far into a structured list; keep raw text only for the current hopBest quality; most code

This is the heart of the course title. Retrieval decides what could enter the window; context engineering decides what does — and in an agent, that decision is made repeatedly, under a growing budget, which is what makes it hard.

Lab — build the agent, then interrogate it

Step 1 · The tool interface

Create agent_tools.py — thin wrappers over what you already built:

from retrieve import retrieve_chunks

def search(query, k=3):
    chunks = retrieve_chunks(query, k=k)
    return [{"doc": c["doc"], "text": c["text"][:900]} for c in chunks]

TOOL_SPEC = """Available actions:
  SEARCH: <query>   — search the knowledge base
  ANSWER: <answer>  — give the final answer with [doc] citations"""

Step 2 · The ReAct loop

Create agent.py:

import ollama, re
from agent_tools import search, TOOL_SPEC

SYSTEM = """You answer questions using a knowledge base.

{tools}

Rules:
- Reply with EXACTLY ONE line: either "SEARCH: ..." or "ANSWER: ...".
- Search when you lack information. Answer when the observations suffice.
- Never answer from prior knowledge; only from observations.
- If after searching the knowledge base still lacks the answer, ANSWER: I don't know."""

def run_agent(question, max_hops=4, verbose=True):
    trace = [f"QUESTION: {question}"]
    used_chunks = []
    for hop in range(max_hops):
        prompt = (SYSTEM.format(tools=TOOL_SPEC) + "\n\n" + "\n".join(trace)
                  + "\n\nYour next line:")
        line = ollama.chat(model="llama3.1:8b",
                           messages=[{"role": "user", "content": prompt}]
                           )["message"]["content"].strip().splitlines()[0]
        if verbose: print(f"  hop {hop+1}: {line[:100]}")

        m = re.match(r"SEARCH:\s*(.+)", line, re.I)
        if m:
            q = m.group(1).strip().strip('"')
            results = search(q)
            used_chunks.extend(results)
            obs = "\n".join(f"  [{r['doc']}] {r['text'][:400]}" for r in results) or "  (nothing found)"
            trace += [f"THOUGHT: I need to search for: {q}", f"OBSERVATION:\n{obs}"]
            continue

        m = re.match(r"ANSWER:\s*(.+)", line, re.I | re.S)
        if m:
            return m.group(1).strip(), {"hops": hop + 1, "chunks": used_chunks, "trace": trace}

        # defensive: model drifted from the format — treat the text as the answer
        return line, {"hops": hop + 1, "chunks": used_chunks, "trace": trace, "drift": True}

    return "I don't know.", {"hops": max_hops, "chunks": used_chunks,
                             "trace": trace, "hit_cap": True}

Run it on a multihop question from your gauntlet and watch the hops print. The first time an agent writes its own second query and finds what it was missing is the moment this week clicks.

Step 3 · Sufficiency check and self-correction

Create self_correct.py:

import ollama
from generate import generate
from retrieve import retrieve_chunks
from generate import PROMPT, format_sources

def unsupported_claims(answer_text, chunks):
    ctx = "\n\n".join(c["text"] for c in chunks)[:5000]
    r = ollama.chat(model="llama3.1:8b", messages=[{"role": "user", "content":
        f"Context:\n{ctx}\n\nAnswer:\n{answer_text}\n\n"
        f"List any claims in the answer NOT supported by the context, one per line. "
        f"If all claims are supported, reply exactly: NONE"}])
    text = r["message"]["content"].strip()
    return [] if text.upper().startswith("NONE") else \
           [c.strip("-• ") for c in text.splitlines() if len(c.strip()) > 10]

def corrected_answer(question):
    draft, chunks = generate(question)
    gaps = unsupported_claims(draft, chunks)
    if not gaps:
        return draft, {"revised": False, "gaps": []}
    for gap in gaps[:2]:                      # re-retrieve for what was unsupported
        chunks.extend(retrieve_chunks(gap, k=2))
    for i, c in enumerate(chunks, 1):
        c["id"] = f"S{i}"
    r = ollama.chat(model="llama3.1:8b", messages=[{"role": "user", "content":
        PROMPT.format(sources=format_sources(chunks), question=question)}])
    return r["message"]["content"].strip(), {"revised": True, "gaps": gaps}

Step 4 · Context management

Add windowing to run_agent: keep the question and the last two observations verbatim, and replace older observations with a one-line summary. Re-run your gauntlet. Did quality hold? Did token use drop? Break it deliberately — window down to one observation and watch a multihop question fail because hop 1's finding was evicted. That controlled failure is the lesson; students who have caused it never forget why context management is a design decision.

Step 5 · The head-to-head — this week's real deliverable

Create compare.py: run your full golden set through three systems — Week 6 pipeline, agent, self-correcting pipeline — recording per item: correctness (Week 5 judge), faithfulness, latency, and token count. Then break results down by query type, because the aggregate hides the finding:

Query typeExpect
Simple factualPipeline usually wins — same answer, a fraction of the cost
MultihopAgent wins — this is what the loop is for
UnanswerableAgent often worse: it searches repeatedly before admitting defeat, burning tokens; sometimes it talks itself into an answer
Ambiguous / vagueAgent wins when it rewrites the query well; loses when it wanders

Write the deployment recommendation: which traffic goes to which system, with numbers. "Route multihop to the agent, everything else to the pipeline — the agent costs 3.4× the tokens for +0.02 faithfulness on simple queries" is a professional conclusion. "Agents are better" is not.

Troubleshooting
  • Agent never searches: it's answering from parametric memory. Strengthen the rule ("You have NO prior knowledge of this domain") and make the first hop a forced search.
  • Agent loops on the same query: add the previous queries to the prompt with "Do not repeat these searches", and keep the cap at 3–4.
  • Format drift (prose instead of SEARCH:/ANSWER:): expected with 8B models — the defensive branch handles it. Count drift events; it's a legitimate metric for your comparison table.
  • Everything is slow: agents make 3–8 LLM calls per question. Test on 15 golden items, not 50, while iterating.
Week 7 checkpoint — done when
  • agent.py runs a ReAct loop with a hard hop cap and defensive parsing
  • Sufficiency check and one self-correction cycle implemented
  • Context management added, and its failure mode demonstrated deliberately
  • Three-system comparison broken down by query type: quality, latency, tokens
  • A written deployment recommendation with numbers behind it
  • Committed: git commit -m "Week 7: agentic RAG + head-to-head comparison"
This weekend

Challenge 7: Agent vs. Pipeline — the rigorous head-to-head, with the stretch goal of building the router you just recommended and measuring blended cost and quality. Reading and videos on the Week 7 page.