Skip to content
Ishaan Reddy

Cookbook · Transformers

Transformers

4 mintransformersarchitecturedeep-learning

The idea, in one analogy

Imagine reading a sentence one word at a time, and before guessing the next word, you're allowed to glance back at every word you've already read and decide how much attention each one deserves. Take "The trophy didn't fit in the suitcase because it was too big": to figure out what "it" refers to, you weigh "trophy" and "suitcase" very differently depending on the rest of the sentence. That weighing act is "attention," and a transformer is a stack of layers that do this weighing, over and over, building up a richer understanding of the sequence at each layer.

What's inside one

A GPT-style transformer (decoder-only, the kind used for text generation) is built from four pieces, stacked.

Token and position embeddings come first: each token (see tokenization) gets converted to a vector of numbers. Attention has no built-in sense of order, so a separate "position" vector is added so the model knows token 3 came before token 7.

Then causal self-attention: each position looks back at itself and everything before it, decides how much each earlier position matters, and mixes their information together. This is the mechanism that lets a model use context ("it" referring back to "trophy" or "suitcase" depending on the rest of the sentence). The full mechanics, including the formula and a worked example, live in attention.

After attention mixes information across positions, an MLP (feed-forward block) processes each position independently: typically expand to 4x the width, bend the values through a simple non-straight-line function (this is what breaks the whole network out of being one giant linear operation, letting it represent curves and thresholds, not just straight-line relationships), then project back down. A lot of the model's "knowledge" is thought to live here.

Last, residual connections and normalization. Attention and MLP outputs are added back to their input (not replacing it) at every step, and normalized, which is what makes it possible to stack dozens of these layers without training collapsing.

A "block" is one attention plus one MLP, with their residual connections. A GPT model is N of these blocks stacked, followed by a final projection back to vocabulary size, turning the model's internal representation into a probability distribution over what token comes next.

The full pipeline, start to finish

The tokenizer step below doesn't just produce numbers out of nowhere: it first splits the raw text into pieces (see tokenization for how those pieces are decided), and each piece is what gets looked up as an ID.

PieceThecatsatonthemat
Token ID464379733323192622603

Note the leading space is usually part of the piece itself (GPT-2's tokenizer would split this as The, cat, sat, on, the, mat, six pieces for six words here since each is common enough to get its own token, but a rarer or longer word can split into multiple pieces, see tokenization's BPE walkthrough for an example where that happens).

Where to look further

  • attention: the full attention mechanism, causal vs. bidirectional, and a worked numeric example.
  • inference: how a trained transformer generates text, one token at a time.
  • Karpathy's nanoGPT: the clearest minimal implementation to read alongside this.
  • This project's own from-scratch build: slm-from-scratch Phase 1 and src/phase1/model.py.