Loading
0xA0Lesson 11 of 11

Capstone: build a RAG pipeline

Chunk, index, retrieve and answer with citations - and abstain when the sources don’t cover the question.

35 min 5-question quiz 2 code exercises
By the end of this lesson you can
  • Chunk documents into citable sentences
  • Retrieve the best chunk for a question
  • Answer with a citation, or abstain below a threshold

Time to build a whole (tiny) RAG pipeline. To keep it runnable anywhere, the “generator” is extractive: it answers with the best-matching sentence itself, plus its citation. Swapping in a language model later changes only the last step - the indexing, retrieval, citation and abstention logic stays the same.

Two exercises:

  1. Index: split each document into sentences with stable IDs like returns#2.
  2. Answer: score sentences against the question, cite the best one, and say “I don’t know” when the best score is too low.
stem.py
def stem(word):
    return word[:-1] if word.endswith("s") and len(word) > 3 else word
print([stem(word) for word in ["costs", "electronics", "is", "days"]])
Output
['cost', 'electronic', 'is', 'day']

Key takeaways

  • A RAG pipeline is chunk → index → retrieve → generate with citations.

  • Stable chunk IDs make citations possible.

  • A score threshold lets the system abstain instead of guessing.

Lesson quiz

5 questions · pass with 4 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Step 1: index sentences

+25 XP

Each line is doc|text. Split the text into sentences with re.split(r"(?<=[.!?])\s+", text) and print each as doc#N: sentence, numbering from 1 within each document.

  • Two documents
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Step 2: answer with a citation

+25 XP

Lines before --- are documents (doc|text); lines after are questions. Index sentences as in step 1. For each question:

  • take its words (lowercase [a-z0-9]+), stem them with stem, and drop STOP words;
  • score each sentence by how many of those question words appear among its stemmed words;
  • pick the highest score (the first sentence wins ties).

Print question -> sentence [doc#N], or question -> I don't know - no source covers that. if the best score is below 2.

  • Answer, answer, abstain
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: