Watch a RAG system work out an answer, step by step.

RAGretrieval-augmented generation

A way to have a language model answer from a given set of documents rather than from memory. For each question, a search first retrieves the most relevant passages; the model then generates its answer from those passages alone and cites them.

One RAG system over 35 Wikipedia articles, with Claude writing the answers. Every step runs live below, with its real timing, tokens and cost; change a setting and ask again to see what changes.

The whole pipeline

Every question

  1. 01Check the request
  2. 02Embed the question
  3. 03Place it on the map
  4. 04Keyword search (BM25)05Vector search
  5. 06Fuse the rankings
  6. 07Rerank
  7. 08Choose the context
  8. 09Write the prompt
  9. 10Generate the answer
  10. 11Check the citations
  11. Answer

Each box links to its step below and fills in as your question runs. Switch the agent on and Claude plans the searches first, then repeats the search steps once per part of the question.

Sheet 1 ·Ask the corpus a question

Starting the lab… The server sleeps when nobody is using it and takes a few seconds to wake.

Or try one of these:

  • (not in the corpus)

The working

11 steps, in the order they run