Skip to content
Ishaan Reddy

Cookbook · FineTuning

Project: distill a small model from an open Llama model

project3 mindistillationquantizationllamafine-tuning

What you're building

A small student model trained to imitate a larger, openly-available Llama model's outputs (see distillation), rather than trained from scratch on raw text. This project surfaces the two constraints distillation always runs into in practice, tokenizer compatibility and teacher-generation cost, using real, checkable numbers instead of taking them on faith.

The stages, in order

1. Check tokenizer compatibility before anything else. If you want soft-label distillation (matching the teacher's full probability distribution, not just its top answer), the student's tokenizer needs to match the teacher's, or you need a token-alignment scheme. Llama models use their own vocabulary; a student built around a different tokenizer (a custom one, or a different family's) rules out soft labels entirely until that's resolved. Check this first, it's a five-minute check that avoids discovering the constraint mid-project.

2. Decide hard-label or soft-label distillation. If the tokenizers don't match and reconciling them isn't worth it, hard-label distillation (train the student on the teacher's generated text as if it were ground truth) still works with no compatibility requirement, at the cost of the extra signal soft labels would have carried (see distillation's "dark knowledge" example for what that extra signal looks like).

3. Measure teacher throughput before committing to a data-generation plan. Autoregressive generation is decode-bound (see inference), meaningfully slower than a single forward pass. Generate a small batch first and measure actual tokens/second on your hardware; multiply that out against how many teacher tokens your planned dataset needs before assuming a "generate 500M tokens of teacher output" plan is feasible in the time you have.

4. Generate the teacher dataset. For hard labels: prompts in, teacher completions out, saved as an SFT-style dataset (see SFT dataset formats). For soft labels: cache the teacher's top-k logits per position rather than the full vocabulary distribution (the long tail beyond the top few hundred tokens carries very little of the KL signal, and storing the full vocabulary's logits for every position gets expensive fast).

5. Train the student, using the hybrid loss from distillation if you have soft labels, or ordinary SFT (see SFT) if you're using hard labels only.

6. Evaluate the student against both the teacher and a same-size model trained without distillation, if you have the compute for the comparison. Distillation's actual value only shows up relative to that baseline, not in isolation.

The decision this project is built to surface

Before scoping a soft-label distillation pipeline, verify: same tokenizer as the teacher (or an alignment plan), and a measured (not assumed) teacher throughput number. Projects that skip these two checks tend to discover them the expensive way, partway through building the pipeline instead of before starting it.

Where to look further

  • Distillation: the core mechanism, the vocabulary-compatibility constraint, and the teacher-cost tradeoff this project is built around.
  • slm-from-scratch's Phase 4: a real investigation that hit exactly the tokenizer-mismatch constraint described above, with the actual measured teacher throughput numbers.