Run GGUF AI models on an AMD, Intel or Nvidia GPU behind a local API
Made by Vibra-Ingenn
View profile on GitHub (opens in a new tab)Janus is a small local AI server written in Go: give it a GGUF model file (a compressed model format that runs on ordinary hardware) and it answers through an OpenAI-compatible API, the same request format OpenAI uses, so tools like Cursor or Cline can talk to it instead of a cloud service.
It runs the model through llama.cpp using Vulkan, a graphics interface supported by AMD, Intel and Nvidia cards, or on the CPU. Its GitHub description calls it an API router for AI models; the README describes a local model server. No Python or Docker is needed.
Can I use this?
You'll need Windows 10/11 (main target), Linux or macOS, Go 1.22+, and a GGUF model · Vulkan GPU recommended
Worth knowing There is no ready-made download: you build it yourself with Go. The single binary still needs a llama.cpp library file beside it, which the Windows build script downloads. Windows is the main platform; macOS runs on the CPU. Models are 2 to 8 GB each and the first reply can take 10 to 60 seconds while the model loads. It is a new, small project (created 24 September 2026). BuildTube has not run or verified this software.
Why it's here
- 129 people have starred it on GitHub.
- Shown on Hacker News (Show HN) on 1 October, with 106 points.
Numbers as of 6 October 2026, from BuildTube’s latest daily refresh.
Behind it
See the code on GitHubBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?