LLM Cookbook
A modular, tool-agnostic reference for training language models from scratch — transformers, tokenization, data engineering, fine-tuning, distillation, scaling laws, and evaluation, structured to look things up rather than read cover to cover.
Transformers
- AttentionDerives the query-key-value attention mechanism, contrasts causal masking against bidirectional attention, and walks a worked numeric example of masked scores turning into a softmax distribution.
- TransformersWalks through the GPT-style decoder-only transformer architecture — embeddings, causal self-attention, MLP blocks, and residual connections — and traces the full pipeline from raw text to next-token probabilities.
Evaluation
- BenchmarksExplains how multiple-choice LLM benchmarks compute loglikelihood scores, why a benchmark's chance floor determines whether a raw percentage means anything, and how error analysis decides whether targeted iteration will help.
- Human evaluationCovers the three common human-eval setups — pairwise preference, rubric-based rating, and production A/B testing — and why inter-rater agreement (Cohen's kappa) determines how much weight a rating task's scores deserve.
- LLM-as-judgeExplains the pointwise, pairwise, and reference-guided patterns for using a strong LLM as an evaluation grader, and the documented length, position, and self-enhancement biases that make it a proxy rather than ground truth.
LLMs
- Data engineeringCovers the language-identification, quality-filtering, deduplication, and re-tokenization pipeline that turns raw web scrape into a corpus worth pretraining on.
- PretrainingWalks through the forward/backward/optimizer training loop behind next-token prediction, what cross-entropy loss means in practice, and why checkpoint-resume correctness depends on ordered optimizer state.
- Scaling lawsExplains the Chinchilla compute-optimal model-size-to-dataset-size ratio and how to run a small, controlled scaling-law comparison of your own.
- TokenizationExplains why byte-pair encoding is the standard middle ground between character- and word-level vocabularies, with a worked BPE training example and a measured comparison of token counts across tokenizers.
- Project: train your first model from scratchprojectA staged walkthrough of training a small language model from randomly-initialized weights, covering architecture and tokenizer choice, corpus filtering, the training loop, and reading the resulting loss curve honestly.
Inference
- DeploymentDistinguishes serving from deployment and walks through where a model runs — managed API, self-hosted server, or on-device — plus the application-layer concerns of streaming, rate limiting, and cost modeling.
- Inference (generation)Covers why autoregressive generation is sequential rather than parallel, surveys the main sampling strategies (greedy, temperature, top-k, top-p, beam search), and explains how the KV cache avoids recomputing attention at every step.
- QuantizationExplains the affine quantization formula behind rounding model weights to lower-bit representations, contrasts post-training quantization with quantization-aware training, and covers the accuracy tradeoff at 8-bit, 4-bit, and sub-4-bit precision.
- ServingCovers the request-level techniques a production inference stack uses to stay fast under unpredictable traffic — continuous batching, PagedAttention KV cache management, and speculative decoding.
FineTuning
- Project: distill a small model from an open Llama modelprojectA staged walkthrough of knowledge distillation from an open Llama teacher, covering the tokenizer-compatibility check that gates soft labels, measuring teacher generation throughput, and training the student on hard or soft labels.
- DistillationContrasts hard-label and soft-label knowledge distillation, works through what 'dark knowledge' means in a teacher's output distribution, and covers the vocabulary-compatibility and teacher-inference-cost constraints that shape a distillation pipeline.
- LoRA (Low-Rank Adaptation)Explains how LoRA freezes a pretrained weight matrix and trains a small low-rank adapter alongside it, including the underlying math and where in a transformer the adapter is typically applied.
- QLoRAExplains how QLoRA quantizes a frozen base model's weights to 4-bit NF4 alongside a full-precision LoRA adapter, and the double quantization and paged optimizer tricks that make the training-time memory savings work.
- SFT dataset formatsCovers the instruction/response and multi-turn chat-message shapes used for supervised fine-tuning data, and why loss must be masked to the assistant's response tokens rather than the whole sequence.
- Supervised fine-tuning (SFT)Explains why supervised fine-tuning is cheap relative to pretraining, how it differs from full fine-tuning versus parameter-efficient alternatives, and the catastrophic forgetting failure mode to guard against.
DPO
RLHF
- GRPO (Group Relative Policy Optimization)Explains how GRPO replaces PPO's learned critic with group-relative reward normalization across sampled responses, and when a verifiable checker makes that trade-off worthwhile.
- Preference dataset formatsCovers the prompt/chosen/rejected triple used by DPO and RLHF, how those comparisons are typically collected (human annotation, RLAIF, rejection sampling), and the length-bias trap to check for before training on one.
- RLHF (Reinforcement Learning from Human Feedback)Explains the two-stage RLHF recipe — training a reward model on human preference comparisons, then optimizing the policy against it with PPO under a KL penalty — and why that machinery is expensive relative to DPO and GRPO.