All insights
Small BusinessOpen SourceTry This

One Binary, Any AI Model: Zero Setup Required

Jesse Burcsik·October 2, 2026·3 min read

A new open-source project landed on Hacker News this morning with a refreshingly short pitch: one file, drop in a model, and you have a private AI server running on your own hardware.

What's happening

Janus is a single Go binary that runs any GGUF model locally and exposes an OpenAI-compatible API. No Python environment to manage, no Docker, no Ollama required. You download the binary, drop a model file in the models/ folder, and run it. Anything that already talks to OpenAI (your scripts, Cursor, Cline, custom apps) can now talk to your own machine instead.

The project handles GPU acceleration through Vulkan, which covers AMD, Intel, and Nvidia cards without the CUDA lock-in that has historically made local AI harder for non-Nvidia hardware. No GPU at all? It falls back to CPU. It also supports hot model swapping (switch models without restarting the server) and auto-detects chat templates for different model families, including thinking models.

For a small team, the practical result is a private AI layer that runs on hardware you already own, processes nothing in the cloud, and costs nothing per query. Models are freely available on Hugging Face and range from under 2GB (fine for drafting and summarizing) to 20GB+ for more demanding reasoning tasks.

Try this this week

  • Go to github.com/Vibra-Ingenn/Janus and grab the latest release for your OS. The Windows build is prebuilt and ready to run.
  • Pick a small model to start: Phi-3.5 Mini Q4 or Mistral 7B Q4 are both under 5GB and fast on most machines. Download the .gguf file from Hugging Face.
  • Drop the model file into the models/ folder beside the Janus binary.
  • Run janus.exe (or ./janus on Linux). It starts an OpenAI-compatible API server at localhost:8080.
  • Test it: curl http://localhost:8080/v1/chat/completions with a simple JSON payload and you should see a reply in seconds.
  • In any OpenAI-compatible tool or script, swap the API base URL to http://localhost:8080. Your prompts stay on your machine, full stop.

The bigger picture

Running your own private AI used to mean setting up serious infrastructure. Now it means downloading one file. For a small business or nonprofit handling anything remotely sensitive (client records, donor info, patient notes), switching from a cloud API to a local one is both a privacy win and a budget win. The smallest working thing, running on what you already have, keeps getting more attainable.

Written by Jesse Burcsik

I am a web developer in Ottawa, Ontario who builds AI systems for small businesses. I write these up because I am usually testing the thing in my own work first. Here is what I have built.

Like this? Let's put it to work.

I help small teams turn ideas like these into working systems, without a six month roadmap.

Book a call