Skip to content
Ishaan Reddy

Projects

SLM From Scratch

Training small language models (75M-500M params) from scratch on a single consumer GPU — a hands-on exploration of transformer training, scaling behavior, and what's actually achievable without a GPU cluster.

research— Phases 1-3 complete: core training, tokenizer/data engineering, and three controlled scaling-law runs on a single consumer GPU.2026-05-31 – present
  • Python
  • PyTorch 2.6
  • Hugging Face Transformers
  • tokenizers
  • datasets
  • Weights & Biases

Problem

Understanding how language models are trained from papers and blog posts only goes so far — it tells you what works, not why the decisions in a real training pipeline matter. SLM From Scratch is a hands-on exploration of that gap: building a transformer, its training loop, and its data pipeline from scratch, and actually running it — on hardware anyone could plausibly own, not a GPU cluster.

Motivation

SLM From Scratch exists because I wanted to move beyond understanding the theory of language models and actually experience the engineering involved in training one. Reading papers explains what works, but building the training pipeline yourself teaches you why those decisions matter. The project naturally became the practical companion to the LLM Cookbook — one explains the concepts, the other validates them through implementation.

The hardware constraint — a single RTX 5070 Ti (16GB VRAM) — was intentional, not incidental. Modern LLM research often assumes access to large GPU clusters, but I wanted to explore how much could realistically be learned on consumer hardware: what many enthusiasts, students, and independent developers can actually afford. Demonstrating that meaningful experimentation is possible within those constraints was part of the project's purpose from the beginning.

Architecture

The project is structured as a 5-phase plan. Its first three phases — the core training pipeline, tokenizer and data engineering, and controlled scaling-law experiments — are complete. Phase 1 established the working transformer, training loop, and generation pipeline:

  • src/model.py — a GPT-style transformer built from scratch: causal self-attention, MLP blocks, transformer blocks, assembled into a full GPT model.
  • data/prepare.py — tokenizes OpenWebText (~38GB raw, 8M documents) into train.bin/ val.bin using GPT-2's BPE tokenizer, yielding roughly 8.5B tokens.
  • src/train.py — the training loop: cosine learning-rate schedule with warmup, AdamW (betas 0.9/0.95, weight decay 0.1), gradient accumulation, gradient clipping, mixed precision, torch.compile, Flash Attention via PyTorch's SDPA, and Weights & Biases logging.
  • src/generate.py — top-k sampling from a trained checkpoint.

Two initial model configurations were trained end-to-end on this pipeline: a 50M-parameter model (6 layers, 6 heads, 384-dim embeddings) and a 77M-parameter model (8 layers, 8 heads, 512-dim embeddings), each for 5,000 steps (~80M tokens) on the RTX 5070 Ti. Phase 2 then added a custom 32K BPE tokenizer and filtered data pipeline. Phase 3 re-tokenized the full ~9.14B-token OpenWebText corpus and completed three controlled runs: a 75M configuration at 500M and 2B tokens, and a 150M configuration at 500M tokens.

Core Features

  • A transformer built from first principles: attention, MLP, and transformer blocks, not a library abstraction.
  • A full training loop with the practical details that make real training work: mixed precision, torch.compile, Flash Attention/SDPA, gradient accumulation, gradient clipping, checkpointing, and live metric logging to Weights & Biases.
  • An OpenWebText data pipeline, tokenized with GPT-2's BPE tokenizer into a multi-billion-token training set.
  • Text generation via top-k sampling from any trained checkpoint.
  • Five documented training runs across the initial 50M/77M experiments and Phase 3's controlled 75M/150M scaling comparisons.

Technical Decisions

The single-GPU constraint shaped the entire project, deliberately. Rather than renting cloud compute to train larger models faster, the project stayed on one RTX 5070 Ti throughout — the point wasn't to train the biggest model possible, it was to find out what's actually learnable about transformer training within hardware most people could realistically own.

The project is also structured as five explicit phases rather than one open-ended effort: build the core pipeline first (Phase 1), then use it to run controlled experiments (scaling laws, distillation, evaluation) once that foundation is solid — rather than trying to explore all of it in parallel before any of it worked reliably.

Engineering Challenges

Getting the training pipeline working reliably was the first milestone, and it had several parts that all had to work together before any experiment was meaningful: mixed precision, Flash Attention/SDPA, torch.compile, gradient accumulation, checkpointing, and evaluation.

Once that foundation existed, the more interesting challenge became interpreting the experimental results rather than producing them. The 77M model underperforming the 50M model at equal token count — despite having more capacity — is a genuine finding, not a bug: it reinforced that simply increasing parameter count doesn't help if the training budget (tokens seen) stays fixed. That result made scaling-law discussions concrete in a way reading about them doesn't.

Phase 3 also exposed a genuine training-instability spike in the 75M/2B-token run: validation loss jumped from roughly 4.4 to 5.9 around iterations 50,000-53,000, then recovered without a checkpoint resume and continued downward. Keeping that transient visible in the results, rather than smoothing it away, was important for interpreting the experiment honestly.

Screenshots

Validation loss plotted against training tokens for three Phase 3 scaling-law runs using 75M and 150M model configurations
Phase 3 validation loss versus tokens seen. The 75M/2B-token run's temporary instability is visible before it recovers and continues downward.
Sample text generated by the 50M parameter model after 5,000 training steps
50M model output after 5,000 steps.
Sample text generated by the 77M parameter model after 5,000 training steps
77M model output, same training budget as the 50M run.

Lessons Learned

The biggest lesson is that training language models is far less mysterious than it initially appears. Once the infrastructure is in place, the process becomes a series of engineering trade-offs rather than magic. Reading papers explains the algorithms, but implementing the pipeline teaches the practical considerations — memory constraints, checkpointing, optimizer behavior, data quality, debugging numerical issues — that are difficult to appreciate without running the experiments yourself.

Before running experiments, scaling laws felt like theoretical guidance. After training models myself, they became concrete engineering constraints. Phase 3 also complicated the simple lesson from Phase 1: at 500M tokens, the 150M configuration achieved lower validation loss than the 75M configuration despite its lower tokens-per-parameter ratio. Parameters, tokens, compute budget, optimization, and the observed curves all have to be considered together rather than reduced to a single heuristic.

Current Status

Phases 1-3 are complete and documented: the core transformer/training pipeline, a custom 32K BPE tokenizer and filtered data pipeline, and three controlled scaling-law runs. Phase 3's checkpoints are published on Hugging Face, with plots generated directly from the training logs rather than hand-copied metrics. The broader five-phase plan continues from this experimental foundation into distillation and evaluation, alongside the related LLM Cookbook reference.

Tech Stack

Python, PyTorch 2.6 with torch.compile and SDPA-based Flash Attention, Hugging Face transformers/tokenizers/datasets for the data pipeline, and Weights & Biases for experiment tracking.

Background Reading: LLM Cookbook documents the theory and engineering concepts used throughout this project — tokenization, scaling laws, distillation, and evaluation are all covered there in depth. This project is the hands-on validation of those ideas.

Gallery

Validation loss plotted against training tokens for three Phase 3 scaling-law runs using 75M and 150M model configurations
Phase 3 scaling experiments: validation loss across 75M/500M-token, 75M/2B-token, and 150M/500M-token runs.
Sample text generated by the 50M parameter model after 5,000 training steps
50M model output — grammatically plausible, not yet semantically coherent across sentences.
Sample text generated by the 77M parameter model after 5,000 training steps
77M model output, same training budget as the 50M run.
Phase 3 controlled runs
3
Best Phase 3 validation loss
3.7578 (150M config, 500M tokens)
Largest Phase 3 token budget
2B (75M config)