About LocalAI
LocalAI is an open-source AI engine for running models of many kinds, including language, vision, voice, image and video models, on your own hardware. A GPU is not required. It is written in Go and released under the MIT license, was started by Ettore Di Giacinto and is looked after by the LocalAI team.
Its design is a small core with separate backends that are pulled on demand, so nothing unused gets installed. Backends wrap engines such as llama.cpp, vLLM, whisper.cpp, stable-diffusion and MLX, and you can write your own in any language against an open interface. It can run on CPU alone, through Vulkan, or on NVIDIA, AMD, Intel and Apple Silicon hardware.
APIs are drop-in compatible with OpenAI, Anthropic and ElevenLabs across every backend. Multi-user features include API key authentication, user quotas and role-based access, and built-in agents support tool use, RAG, MCP and skills. The project stresses that data stays on your infrastructure. It suits developers and teams who want a private, self-hosted replacement for cloud AI APIs.
Key features
- Runs LLM, vision, voice, image and video models
- Composable backends pulled on demand
- OpenAI, Anthropic and ElevenLabs API compatibility
- Runs on CPU, NVIDIA, AMD, Intel and Apple Silicon
- API key authentication, quotas and role-based access
- Built-in agents with tools, RAG and MCP
Good fit for
- →Private self-hosted replacement for cloud AI APIs
- →Running models on CPU-only servers
- →Shared AI service with per-user quotas
- Built with
- Go
- Tags
- local-ai
- llm
- self-hosted
- openai-compatible
- golang
- image-generation
- speech
- agents
LocalAI: questions and answers
- What is LocalAI used for?
- LocalAI is an open-source AI engine that runs LLMs, vision, voice, image and video models on your own hardware behind OpenAI-compatible APIs, with no GPU required. It is a good fit for private self-hosted replacement for cloud AI APIs, running models on CPU-only servers and shared AI service with per-user quotas.
- Is LocalAI open source?
- Yes. LocalAI is open source under the MIT licence. Its source code is on GitHub at mudler/LocalAI and is written mainly in Go.
- Is LocalAI free?
- Yes. LocalAI is open source, so the software itself is free to use.
- Can I self-host LocalAI?
- Yes. LocalAI can be self-hosted on your own server or infrastructure; there is no official hosted version.
- What is LocalAI an alternative to?
- LocalAI is an open-source alternative to ChatGPT, Replicate, Together AI and Fireworks AI. Other open-source alternatives to ChatGPT include Ollama and llamafile.
- Is LocalAI actively maintained?
- Yes. The most recent commit to LocalAI was on 2 October 2026, and the latest release is v4.10.0, published on 17 September 2026. The project has 49k stars on GitHub.
Open-source alternatives to LocalAI
See all
vLLM
AI Infrastructure
A high-throughput and memory-efficient inference and serving engine for LLMs
Apache-2.0vs Amazon Bedrock★ 93k
Ollama
AI Infrastructure
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other model
MITvs ChatGPT★ 182k
SGLang
AI Infrastructure
SGLang is a high-performance serving framework for large language models and multimodal mo
Apache-2.0vs Amazon Bedrock★ 37k
llama.cpp
AI Infrastructure
LLM inference in C/C++
MITvs OpenAI API Platform★ 130k
beta9
AI Infrastructure
Ultrafast serverless GPU inference, sandboxes, and background jobs
AGPL-3.0vs Modal★ 1.8k
llamafile
AI Infrastructure
Distribute and run LLMs with a single file.
OSSvs ChatGPT★ 26k
SaaS alternatives to LocalAI
See all
ChatGPT
AI Tools
ChatGPT is OpenAI's conversational AI assistant for answering questions, writing, coding and analysis through web and mobile apps.
SaaS
Replicate
AI Infrastructure
API platform for running open-source AI models in the cloud
SaaS
Together AI
AI Infrastructure
Cloud platform for running, fine-tuning and training open and custom AI models
SaaS
Fireworks AI
AI Infrastructure
Inference platform for running and fine-tuning generative AI models at low latency
SaaS
Groq
AI Infrastructure
AI inference cloud and API built on custom LPU hardware for fast model serving
SaaS
DeepInfra
AI Infrastructure
Pay-per-use API for running open-source AI models
SaaS

