Skip to content
Ishaan Reddy

Cookbook · GPUProgramming

Training frameworks and efficiency techniques

3 mintraining-frameworksdistributed-traininggradient-checkpointingmixed-precision

Picking a framework

Once you're not trying to learn the training loop itself (see pretraining), hand-rolling it stops being the point. Common options, roughly in order of how much they abstract away:

  • Raw PyTorch: full control, most learning value, most code to maintain yourself.
  • Hugging Face Trainer: a maintained training loop covering most standard cases (mixed precision, checkpointing, logging, distributed training) with a config-driven interface.
  • Unsloth: custom Triton kernels (fused RoPE, MLP, attention) and smart sequence packing that can make fine-tuning (LoRA/QLoRA especially) noticeably faster and lower-memory on consumer GPUs. Built on top of the Hugging Face ecosystem, not a replacement for it.
  • axolotl: a YAML-config-driven wrapper around the Hugging Face stack, popular for fine-tuning runs without writing training code at all.

Efficiency techniques that apply regardless of framework

Gradient checkpointing trades compute for memory: instead of storing every intermediate activation from the forward pass for use in the backward pass, it stores only a subset and recomputes the rest when needed. This can cut activation memory substantially at the cost of a slower backward pass (often 20-30% slower), and is usually the first thing to enable when a training run doesn't fit in memory.

Mixed precision (see pretraining's knobs section) does the forward/backward math in bf16 or fp16 instead of fp32, roughly halving memory for activations and increasing throughput on hardware with dedicated low-precision compute units, while keeping a higher-precision copy of weights for the actual optimizer update.

Distributed training splits training across multiple GPUs (or machines) when a single device isn't enough:

  • Data parallelism: every device holds a full copy of the model and processes a different slice of the batch, then gradients are synchronized across devices. Simplest to reason about; doesn't help if the model itself doesn't fit on one device.
  • Tensor parallelism: individual weight matrices are split across devices, so a single layer's computation is itself distributed. Needed when a model is too large for one device's memory even at batch size 1.
  • Pipeline parallelism: different layers of the model live on different devices, and micro-batches flow through them like an assembly line. Reduces per-device memory at the cost of some idle time ("bubble") while later stages wait for earlier ones.

Large training runs typically combine several of these at once (data + tensor + pipeline parallelism together), which is what frameworks like DeepSpeed and Megatron-LM are built to orchestrate.

Where to look further

  • DeepSpeed: Microsoft's distributed training library, including ZeRO (a memory-optimization technique for data parallelism that shards optimizer state, gradients, and even parameters across devices).
  • Megatron-LM: NVIDIA's framework for large-scale tensor and pipeline parallelism.
  • Pretraining: the training loop these techniques all optimize around.