by ggml-org

llama.cpp pricing, free plan and limits

Free MIT-licensed C/C++ engine that runs LLMs locally. 130,000+ GitHub stars, GGUF models, OpenAI-compatible server. A developer tool.

  • edge ai platforms
  • Windows
  • Mac
llama.app homepage showing the headline Your AI on your computer, a Download for Windows 11 button and the note Free, Open source, Works offline
Vendor homepage, October 2026

Last updated: 2026-10-07

llama.cpp is an MIT-licensed C/C++ engine that runs language and vision models on local hardware, with more than 130,000 GitHub stars as of 7 October 2026. It supports 1.5-bit to 8-bit quantization and backends for CUDA, Metal, Vulkan, HIP and SYCL, and its llama serve command exposes an OpenAI-compatible API.

About llama.cpp

llama.cpp is an open-source inference engine written in plain C and C++ that runs language and vision models on your own hardware. It is a developer tool: you install a command line program, point it at a model file and either chat in the terminal or start a local server. The code is MIT licensed, the GitHub repository holds more than 130,000 stars as of 7 October 2026, and it sits underneath many of the friendlier apps in this category, including Ollama and LM Studio.

The design goal is to run a model with minimal setup on a wide range of machines. The engine has no required dependencies, treats Apple silicon as a first-class target through Metal and Accelerate, and ships backends for CUDA on Nvidia, HIP on AMD, Vulkan, SYCL for Intel GPUs and several others. Models are stored in the GGUF format and can be quantized from 8-bit down to 1.5-bit to fit into less memory, and a model larger than your video memory can be split between GPU and CPU. A single command such as llama cli -hf followed by a Hugging Face repository name downloads and runs a model from Hugging Face.

For applications, llama serve starts an OpenAI-compatible HTTP server with a built-in web chat page, so tools such as Aider, LangChain and LlamaIndex can talk to a local model by changing one base URL. In 2026 the llama.cpp team and Hugging Face also released Llama, a menu bar and system tray app for Mac and Windows that wraps the engine, picks settings for your machine and exposes the same local API. The vendor lists it as a 1 MB download.

Who it suits: developers building on local models, people who want the newest model support and the most control over settings, and anyone who needs to squeeze a model onto limited memory. If you want a polished model browser or document chat without a terminal, Jan, GPT4All, AnythingLLM and Open WebUI are friendlier starting points, and our guide to local coding models covers which models are worth running.

There is no price. Everything is free and open source, and the real cost is the computer and electricity you run it on. Pre-built binaries, Docker images, Winget and a Linux install script are available, or you can build from source.

Screenshots

llama.app homepage showing the headline Your AI on your computer, a Download for Windows 11 button and the note Free, Open source, Works offline
Vendor homepage, October 2026

Pricing

No paid plans and no accounts. The only cost is your own hardware and power. app and the GitHub repository 2026-10-07.

Key Features

  • Runs on almost any hardware: Plain C/C++ with no required dependencies and backends for CUDA, HIP, Metal, Vulkan, SYCL and CPU instruction sets from AVX512 to ARM NEON.
  • Quantization from 8-bit to 1.5-bit: Integer quantization in GGUF files cuts memory use so larger models fit on ordinary machines, at some cost in output quality.
  • OpenAI-compatible local server: llama serve exposes OpenAI-style endpoints with streaming, tool calling and a built-in web chat page, so existing apps only need a new base URL.
  • Hybrid CPU and GPU inference: Splits a model across video memory and system memory so models larger than your VRAM still run, more slowly.
  • One-command model download: llama cli -hf pulls a GGUF model straight from a Hugging Face repository and starts a chat, with no separate model manager.
  • Llama menu bar and tray app: A 1 MB Mac and Windows app from the llama.cpp team and Hugging Face that picks settings per machine and unloads models after 5 minutes idle.

Pros

  • Free under the MIT licence with no account, so the only cost is the machine it runs on.
  • Supports more than a dozen hardware backends, from Nvidia and AMD GPUs to Apple silicon and Snapdragon.
  • Exposes every setting and is the engine many friendlier apps are built on, so skills transfer to Ollama and LM Studio.
  • The built-in OpenAI-style server means editors and coding agents work with a local model and no API bill.

Cons

  • It is a command line developer tool; there is no model browser or document chat in the core project, so non-technical users need a wrapper app.
  • Speed and quality depend on your hardware and the quantization you pick, and low-bit models lose accuracy against full precision.
  • Settings are numerous and the project moves quickly, so commands and flags change between releases and older guides go stale.
  • No security attestation such as SOC 2 or ISO 27001 applies, since it is software you run yourself, so regulated teams must own their own compliance.

Data Handling

Training-data policy
Inference only. The engine runs models on the user's own machine, and the vendor site says the Llama app works offline.

Frequently Asked Questions

How much does llama.cpp cost?

Nothing. llama.cpp and the new Llama desktop app are free and open source under the MIT licence, with no paid tiers, accounts or usage fees. You pay for the computer and electricity that run the model, and for any API you choose to connect it to.

Can you use llama.cpp without a GPU?

Yes. The engine runs on CPU alone, using AVX, AVX2, AVX512 or ARM NEON instructions, and can split a large model between CPU and GPU when video memory runs short. Expect modest speed on a laptop without a graphics card, and plan on smaller or more heavily quantized models.

What are the best alternatives to llama.cpp?

Ollama adds a model library and background service on top of the same engine, LM Studio offers a desktop model browser, and Jan and GPT4All give a simple chat window. Open WebUI and AnythingLLM add team workspaces and document chat. Choose llama.cpp when you want direct control or the newest model support.

Is llama.cpp faster than Ollama?

Both use the same core engine, so raw speed on one model is usually close. The difference is control: llama.cpp exposes every setting and tends to receive new model support first, while Ollama hides most settings behind defaults and a model library. Measure on your own hardware before assuming a gap.

How do you get llama.cpp running on a PC?

On Windows run the install script from llama.app in PowerShell, or use Winget or the pre-built binaries on the releases page. Then run llama cli with a Hugging Face model repository to download and chat, or llama serve to start the local API. Mac and Windows users can instead install the Llama app.

Top Alternatives

  • Ollama: Pick llama.cpp for direct control and the newest model support; pick Ollama for a model library and background service that hides the settings.
  • LM Studio: Choose llama.cpp for scripting and a free open-source engine; choose LM Studio for a polished desktop model browser.
  • Jan: Pick llama.cpp when you build on the engine; pick Jan for a simple offline chat app.
  • GPT4All: Choose llama.cpp for hardware reach and a server API; choose GPT4All for a minimal installer with LocalDocs.
  • Open WebUI: Pick llama.cpp as the engine and pair it with Open WebUI when a team needs a browser interface on top.

More AI Tools on HokAI

Visit llama.cpp Official Website