Build and train your own small language model from scratch
Made by Fareed Khan
View profile on GitHub (opens in a new tab)This is a hands-on tutorial that walks through building a language model entirely from scratch in PyTorch, with no shortcuts through libraries like transformers, trl or peft that normally do the heavy lifting.
It starts from the Transformer architecture in the original "Attention is All You Need" paper and goes all the way through modern post-training steps — supervised fine-tuning, a reward model, and reinforcement learning methods like DPO, PPO and GRPO — so the reader ends with a small model that can hold a conversation, trainable on a single GPU. Every code block is preceded by a plain explanation and followed by the output you should expect.
You could use it to…
- Build and train a small language model from scratch
- Watch pretraining, fine-tuning and RL happen in code you wrote
- Follow every step from raw text to a chatting model
Can I use this?
You'll need Python, PyTorch and one GPU · Setup needed
Worth knowing The models this produces on a single GPU are small (the author shows a 13-million-parameter example) and meant for learning the mechanics, not for matching production chatbots; the material assumes comfort reading and writing Python. MIT licence. BuildTube has not run or verified this software.
Why it's here
Not trending this week
Last seen on the GitHub trending list on 28 September 2026. It stays on BuildTube so you can still find it.
- 11,991 people have starred it on GitHub.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on GitHubBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?