Cookbook · LLMs
Project: train your first model from scratch
project3 minpretrainingtraining-looptokenizationscaling-laws
What you're building
A small (tens of millions of parameters) language model, trained from randomly-initialized weights on a modest text corpus, on hardware you already have. The goal isn't a useful model, undertrained small models produce grammatical but incoherent text, it's a working, first-hand tour of every stage in pretraining with your own numbers instead of someone else's.
The stages, in order
1. Pick a model size and architecture. Start small: a handful of transformer blocks (see transformers), a modest embedding dimension, a context length of 512-1024 tokens. Bigger comes later; the first goal is a full working loop, not a competitive model.
2. Pick a tokenizer. Using an existing pretrained tokenizer (GPT-2's, for instance) is the fastest path to a working loop. Training your own (see tokenization) is a good second project once the pipeline works end to end, since it adds a real extra step (corpus preparation, merge training) on top of everything else.
3. Get and filter a corpus. A few hundred million tokens of general web text is enough to see real training dynamics. Basic filtering (see data filtering) matters even at small scale: exact deduplication alone often measurably helps.
4. Write (or adapt) the training loop. Forward pass, cross-entropy loss, backward pass, optimizer step (see pretraining's worked example for what to expect from the loss curve). Get checkpointing working and tested (resume from a saved checkpoint at least once) before a long run, not after a crash.
5. Run it, and watch the loss curve. Expect a loss around 10-11 at initialization (random guessing, see pretraining's reference scale) falling toward single digits as training progresses. A loss that isn't falling at all after a reasonable number of steps usually means a bug (a learning rate that's off by an order of magnitude, a data-loading issue), not a model or data problem, worth debugging before assuming the setup is fine and just needs more time.
6. Generate some text and read it honestly. At this scale, expect locally grammatical sentences with no long-range coherence, this is normal and expected for the token budget, not a failure.
Decisions worth making deliberately
How much compute you have (a single consumer GPU, a few hours, versus a rented cluster) should set your model size and token budget before you start, not get discovered partway through a run that turns out to take a week. Scaling laws covers how to reason about that tradeoff, and how to run your own small controlled comparison (does more data help more than more parameters, at your specific budget) instead of guessing.
Where to look further
slm-from-scratch: a completed real-world example of exactly this project, including the scaling-law follow-up experiments, at 75M-150M parameters on a single consumer GPU.- Karpathy's nanoGPT: a clean, minimal reference implementation to build from or compare against.
- Benchmarks: once you have a trained checkpoint, how to measure what it learned instead of just reading a handful of samples.