Picture published by the creator on GitHub.
Run AI language models locally on almost any computer
Made by ggml
View profile on GitHub (opens in a new tab)llama.cpp is the software most local AI chat tools are quietly built on: a plain C/C++ program with no external dependencies that runs language models (and, increasingly, models that also read images) on ordinary hardware — a laptop CPU, an old GPU, a phone — instead of requiring a data-center card.
It supports Apple Silicon, AMD, Intel and RISC-V chips as first-class targets, and squeezes models down with quantization (storing each number in fewer bits) at several compression levels so larger models fit in less memory. A built-in web interface and a command-line chat mode are both included, and Ollama, LM Studio and many other local-AI apps use it as their engine underneath.
You could use it to…
- Chat with a downloaded AI model entirely on your own computer
- Run a model on almost any hardware, from a phone to a GPU
- Feed it a photo through its built-in web interface
Can I use this?
You'll need A computer and a downloaded model file · Setup needed
Worth knowing Getting a good result still depends on which model you download and how heavily you compress it — more compression means a smaller file but a less accurate model, and there is no single answer for the right tradeoff. This project runs models; it does not provide its own model weights, safety filtering, or any hosted service. MIT licence. BuildTube has not run or verified this software.
Why it's here
Not trending this week
Last seen on the GitHub trending list on 28 September 2026. It stays on BuildTube so you can still find it.
- 130,485 people have starred it on GitHub.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on GitHubBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?