Skip to content
Ishaan Reddy

Cookbook · RAG

Project: a RAG chatbot

project3 minragembeddingsretrievalvector-search

What you're building

A chatbot that answers questions about a specific document collection (a codebase's docs, a company's internal wiki, a set of PDFs), without any pretraining or fine-tuning at all. Retrieval-Augmented Generation (RAG) works entirely at inference time: relevant passages are found in your documents and stuffed into the prompt, letting an existing pretrained model (see transformers, inference) answer using content it never saw during training.

Why this is a different project than the others in this section

Train your first model and distill an open model are both about changing a model's weights. RAG changes nothing about the model; it changes what goes into the prompt. This is often the right first thing to try when the actual goal is "answer questions about my specific documents," since it needs no training run at all, just an existing model and a retrieval step in front of it.

The stages, in order

1. Chunk your documents. Split each document into passages small enough to fit several at once inside the model's context window, with some overlap between adjacent chunks so a fact split across a chunk boundary isn't lost entirely in either chunk.

2. Embed each chunk. Each chunk gets converted into a vector using a text-embedding model (a different model from the one that will answer questions, and different from the model's own internal token embeddings, see transformers for that distinction), such that chunks with similar meaning end up with similar vectors.

3. Store the vectors in a retrievable index. A vector database (or even a flat in-memory index for a small collection) that supports "find the k closest vectors to this query vector" efficiently.

4. At query time, embed the user's question the same way, and retrieve the closest chunks. This is the "retrieval" in Retrieval-Augmented Generation: finding which of your documents are relevant to this specific question, out of potentially thousands of chunks.

5. Build a prompt combining the retrieved chunks and the question, and generate an answer. A simple template: "Context: [retrieved chunks]. Question: [user's question]. Answer using only the context above." The "using only the context above" instruction matters: it's what keeps the model from quietly falling back on its own possibly-outdated or incorrect pretraining knowledge instead of your actual documents.

6. Evaluate whether retrieval, not just generation, is working. If answers are wrong, check first whether the right chunks were even retrieved (a retrieval failure) before assuming the generation step is at fault. These are separate failure modes with separate fixes: bad retrieval needs better chunking, a better embedding model, or more chunks retrieved per query; bad generation with correct retrieval needs a better prompt or a more capable model.

A constraint that's easy to miss until it bites

Embedding models and the language model doing the answering are usually different models with different, incompatible vector spaces, an embedding from one has no meaningful relationship to a vector from the other. There's no shortcut here: retrieval needs its own dedicated embedding model, chosen and evaluated on its own terms (how well does it retrieve the right chunk for a given query), not just whatever the answering model happens to produce internally.

Where to look further