Short version: choose the model by the memory you have, not by leaderboard. On a 16-32GB machine, a Qwen3 model between 8B and 30B is the default recommendation in September 2026. Run it with Ollama and put Open WebUI in front of it. Licences differ between models, so check before you build anything commercial.

The models worth trying
Qwen3 family (Apache-2.0). Sizes from 1.7B up to 235B, which means there is a sensible size for almost any machine. The practical choice for most local users is the 8B, 14B or 30B range. Newer Qwen3.x releases continue the family.
gpt-oss (Apache-2.0). OpenAI's open-weight models. The smaller one is realistic on well-equipped desktops; the 120B version is for workstation-class hardware and is also offered on hosted services such as Antigravity's free tier.
Phi-4-mini (Microsoft). A small model that runs on the machine you already have, without a GPU. Good for scripts, summaries and quick questions.
Also in the 2026 field: Gemma 4 from Google, DeepSeek V4-Flash, Mistral Small 4, and very large releases such as Moonshot's Kimi K3, which needs server hardware rather than a laptop. Leaderboard positions change monthly; memory requirements do not.
The tools to run them

- **Ollama** is the easiest start:
ollama runpulls and serves a model with an API other tools can call. - **llama.cpp** is the engine under much of the ecosystem, for maximum control.
- LM Studio gives you a desktop app and model browser if you prefer a GUI.
- **vLLM** is for serving many users on GPUs, built for throughput.
- **Open WebUI** is the ChatGPT-style interface you put in front of any of them.
A sensible starting setup
- Install Ollama.
- Then pull a Qwen3 model sized for your RAM.
- Add Open WebUI.
- Point it at Ollama for a chat interface you can share.
- Connect your editor.
- Cline or Continue can use the same local model for coding help.
- Keep a hosted model for hard tasks.
- Local models still trail frontier models on large multi-file changes.
Our self-hosted LLM guide walks through that setup end to end.