Things I have designed and built.
Training small language models (75M-500M params) from scratch on a single consumer GPU — a hands-on exploration of transformer training, scaling behavior, and what's actually achievable without a GPU cluster.