Skip to content
Ishaan Reddy

Cookbook · FineTuning

LoRA (Low-Rank Adaptation)

4 minfine-tuninglorapefttraining

The idea, in one analogy

Imagine editing a massive, finely-tuned machine instead of rebuilding it. Rather than rewiring every one of its million internal connections, you bolt on a small control panel with a handful of dials, and only those dials get adjusted. The machine's original wiring stays untouched; the dials just nudge its behavior. LoRA (Low-Rank Adaptation) does this for a pretrained model's weights: instead of updating a full weight matrix during fine-tuning, it freezes the matrix entirely and trains a small "adapter" bolted alongside it.

Why this matters

Full fine-tuning (see SFT) updates every parameter in a model, which for a large model means holding gradients and optimizer state for the whole thing, memory that scales with model size regardless of how narrow the actual behavior change is. LoRA freezes the pretrained weights entirely and trains a small number of added parameters instead. The base model's knowledge stays untouched; only a small, cheap-to-train adapter changes. This is what makes fine-tuning a 7B+ parameter model feasible on a single consumer GPU, since the memory-hungry part (gradients and optimizer state) only has to cover the tiny adapter, not the full model.

Where the adapter goes

LoRA is applied per weight matrix, not to the model as a whole. In a transformer (see transformers), the usual targets are the attention projections (W_Q, W_K, W_V, the output projection) and sometimes the MLP's linear layers. Applying it to more matrices increases the adapter's capacity (and its parameter count) at the cost of some of the memory savings; applying it to fewer keeps things cheaper but limits how much the model's behavior can shift.

What you get at the end

A trained LoRA adapter is a small, separate set of weights, often tens of megabytes, that can be merged into the base model's weights (W' = W + BA, computed once) for zero-overhead inference, or kept separate and swapped at load time so one base model can serve many different fine-tuned "personalities" without duplicating the whole model per task.

Where to look further