Ollama
A tool for downloading and running open large language models locally through a command line and REST API, with Docker support and many integrations.
- GitHub stars
- 182k
- Last commit
- today
- Latest release
- v0.35.0
- Licence
- MIT
- Self-hosted
- Yes

Ollama makes it straightforward to run open large language models on your own machine. You install it on macOS, Windows or Linux, or use the official Docker image, then pull a model and chat with it from the command line. It is written in Go and released under the MIT license, and it builds on the llama.cpp project for inference.
Besides the CLI, Ollama provides a REST API for managing and running models, with official Python and JavaScript client libraries. Models can be imported and customized through a Modelfile, and an online library lists what is available. It can be connected to coding agents and assistants such as Claude Code, OpenCode, Codex and OpenClaw, and many community chat interfaces, including Open WebUI and LibreChat, work with it.
Because models run on your own hardware, Ollama can serve as a self-hosted alternative to cloud assistants such as ChatGPT and Perplexity when privacy, offline use or cost control matter. It is aimed at developers and hobbyists experimenting with open models, and at teams that want to wire local models into their own applications.
Key features
- Run open language models locally from the CLI
- REST API for managing and running models
- Official Python and JavaScript libraries
- Modelfile for customizing and importing models
- Official Docker image
- Integrations with coding agents and chat interfaces
Pricing: Running models locally is free, and the Free plan includes starter cloud credits. Pro costs $20, Max $100 and Team $500 per month with included usage credits; Enterprise is custom.

