llama.cpp is an open-source inference engine, written in C/C++, for running large language models and vision-language models with minimal setup. It runs on a wide range of hardware: CPUs, Apple silicon and GPUs from NVIDIA, AMD and others through backends such as CUDA, Vulkan and Metal. Quantisation lets models fit into limited memory. With llama-server it provides an OpenAI-compatible API and a built-in web interface. llama.cpp is a software library with accompanying tools that you run yourself, not a hosted service.
The project was started by Georgi Gerganov and is released under the MIT licence. In 2026 the ggml team behind it joined Hugging Face, which states that llama.cpp will remain fully open source and community driven, with the team keeping technical leadership. Because inference runs entirely on your own hardware, prompts and documents need not leave your infrastructure, and no account with a vendor is required. For a European organisation this makes it a building block for self-hosted AI. The licence terms and origin of the model weights you choose remain a separate point of attention.