5,756 open-source and SaaS tools, with GitHub stats refreshed every day.

llama.cpp

Open source

A dependency-free C/C++ engine for running large language models locally on CPUs and GPUs, with quantization, a CLI and an HTTP server.

llama.app
llama.cpp homepage screenshot
GitHub stars
130k
Last commit
today
Repository age
3 years
Version
v0.5.0
Licence
MIT
Self-hosted
Yes

About llama.cpp

llama.cpp is an open-source library and set of tools for running large language model inference, including vision-language models, with minimal setup. The aim is strong performance across many kinds of hardware, whether local or in the cloud. It is a plain C/C++ implementation without external dependencies, built on the ggml library, and is licensed under MIT.

The project optimizes for many platforms: Apple silicon through ARM NEON, Accelerate and Metal, x86 with AVX, AVX2, AVX512 and AMX, and several RISC-V extensions. Quantization from 1.5-bit to 8-bit integers reduces memory use and speeds up inference. GPU backends include CUDA for NVIDIA, HIP for AMD, MUSA, Vulkan and SYCL, and hybrid CPU plus GPU inference lets you partly accelerate models larger than available VRAM.

It provides a command-line interface and llama-server, which offers a REST API and a built-in web UI. You can install it from pre-built binaries, Docker or by building from source. Other local-AI tools such as Ollama build on it, which makes it relevant to developers embedding inference in their own applications and hobbyists running models on modest hardware.

Key features

  • Dependency-free C/C++ inference engine
  • Quantization from 1.5-bit to 8-bit
  • CUDA, HIP, Metal, Vulkan and SYCL backends
  • CPU plus GPU hybrid inference
  • llama-server with REST API and web UI
  • Optimized for Apple silicon and x86

Good fit for

  • →Running LLMs on laptops and desktops
  • →Embedding inference in C/C++ applications
  • →Serving local models over HTTP
Built with
C++
Tags
llm
inference
cpp
quantization
local-ai
ggml
gpu
self-hosted

Open-source alternatives to llama.cpp

See all

SaaS alternatives to llama.cpp

See all