7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

4 alternatives ranked by real activity

Open-source xAI API alternatives

A curated, ranked list of the 4 best open-source alternatives to xAI API.

The best open-source alternative to xAI API is Ollama. If that doesn't suit you, other good options are llama.cpp, vLLM and LocalAI.

xAI API alternatives are mainly AI infrastructure tools. 4 of them shipped code in the last 30 days, 4 can be self-hosted, and 4 use a permissive licence.

Last updated October 3, 2026 · ranked by GitHub stars, growth and recent commits

Ollama

A tool for downloading and running open large language models locally through a command line and REST API, with Docker support and many integrations.

GitHub stars
182k
Last commit
today
Latest release
v0.35.0
Licence
MIT
Self-hosted
Yes
ollama.comOllama homepage screenshot

Ollama makes it straightforward to run open large language models on your own machine. You install it on macOS, Windows or Linux, or use the official Docker image, then pull a model and chat with it from the command line. It is written in Go and released under the MIT license, and it builds on the llama.cpp project for inference.

Besides the CLI, Ollama provides a REST API for managing and running models, with official Python and JavaScript client libraries. Models can be imported and customized through a Modelfile, and an online library lists what is available. It can be connected to coding agents and assistants such as Claude Code, OpenCode, Codex and OpenClaw, and many community chat interfaces, including Open WebUI and LibreChat, work with it.

Because models run on your own hardware, Ollama can serve as a self-hosted alternative to cloud assistants such as ChatGPT and Perplexity when privacy, offline use or cost control matter. It is aimed at developers and hobbyists experimenting with open models, and at teams that want to wire local models into their own applications.

Key features

  • Run open language models locally from the CLI
  • REST API for managing and running models
  • Official Python and JavaScript libraries
  • Modelfile for customizing and importing models
  • Official Docker image
  • Integrations with coding agents and chat interfaces

Pricing: Running models locally is free, and the Free plan includes starter cloud credits. Pro costs $20, Max $100 and Team $500 per month with included usage credits; Enterprise is custom.

llama.cpp

A dependency-free C/C++ engine for running large language models locally on CPUs and GPUs, with quantization, a CLI and an HTTP server.

GitHub stars
130k
Last commit
today
Latest release
v0.5.0
Licence
MIT
Self-hosted
Yes
llama.appllama.cpp homepage screenshot

llama.cpp is an open-source library and set of tools for running large language model inference, including vision-language models, with minimal setup. The aim is strong performance across many kinds of hardware, whether local or in the cloud. It is a plain C/C++ implementation without external dependencies, built on the ggml library, and is licensed under MIT.

The project optimizes for many platforms: Apple silicon through ARM NEON, Accelerate and Metal, x86 with AVX, AVX2, AVX512 and AMX, and several RISC-V extensions. Quantization from 1.5-bit to 8-bit integers reduces memory use and speeds up inference. GPU backends include CUDA for NVIDIA, HIP for AMD, MUSA, Vulkan and SYCL, and hybrid CPU plus GPU inference lets you partly accelerate models larger than available VRAM.

It provides a command-line interface and llama-server, which offers a REST API and a built-in web UI. You can install it from pre-built binaries, Docker or by building from source. Other local-AI tools such as Ollama build on it, which makes it relevant to developers embedding inference in their own applications and hobbyists running models on modest hardware.

Key features

  • Dependency-free C/C++ inference engine
  • Quantization from 1.5-bit to 8-bit
  • CUDA, HIP, Metal, Vulkan and SYCL backends
  • CPU plus GPU hybrid inference
  • llama-server with REST API and web UI
  • Optimized for Apple silicon and x86

Pricing: Free and open source under the MIT license.

vLLM

A Python library and server for fast, memory-efficient LLM inference and serving, with PagedAttention, continuous batching and broad quantization support.

GitHub stars
93k
Last commit
today
Latest release
v0.30.0
Licence
Apache-2.0
Self-hosted
Yes
vllm.aivLLM homepage screenshot

vLLM is an open-source library for running and serving large language models efficiently. It originated at UC Berkeley's Sky Computing Lab and is now developed by a large community of contributors from academia and industry. It is written in Python and released under the Apache-2.0 license.

Speed comes from techniques such as PagedAttention for attention key-value memory, batching incoming requests continuously, chunked prefill, prefix caching, CUDA and HIP graphs, and optimized attention and mixture-of-experts kernels. It supports many quantization formats, including FP8, INT8, INT4, GPTQ and AWQ, as well as speculative decoding and disaggregated prefill and decode.

On the usability side, vLLM integrates with Hugging Face models and supports parallel sampling, beam search, streaming outputs, structured outputs and tool calling, plus several parallelism modes (tensor, pipeline, data, expert and context) for distributed inference. It targets NVIDIA, AMD and TPU hardware and suits teams that serve open models in production and need high throughput.

Key features

  • PagedAttention and continuous batching
  • Quantization including FP8, INT8, GPTQ and AWQ
  • Speculative decoding support
  • Tensor, pipeline and expert parallelism
  • OpenAI-compatible API server
  • Hugging Face model integration
  • Streaming and structured outputs

Pricing: Free and open source under the Apache-2.0 license.

LocalAI

An open-source AI engine that runs LLMs, vision, voice, image and video models on your own hardware behind OpenAI-compatible APIs, with no GPU required.

GitHub stars
49k
Last commit
today
Latest release
v4.10.0
Licence
MIT
Self-hosted
Yes
localai.ioLocalAI homepage screenshot

LocalAI is an open-source AI engine for running models of many kinds, including language, vision, voice, image and video models, on your own hardware. A GPU is not required. It is written in Go and released under the MIT license, was started by Ettore Di Giacinto and is looked after by the LocalAI team.

Its design is a small core with separate backends that are pulled on demand, so nothing unused gets installed. Backends wrap engines such as llama.cpp, vLLM, whisper.cpp, stable-diffusion and MLX, and you can write your own in any language against an open interface. It can run on CPU alone, through Vulkan, or on NVIDIA, AMD, Intel and Apple Silicon hardware.

APIs are drop-in compatible with OpenAI, Anthropic and ElevenLabs across every backend. Multi-user features include API key authentication, user quotas and role-based access, and built-in agents support tool use, RAG, MCP and skills. The project stresses that data stays on your infrastructure. It suits developers and teams who want a private, self-hosted replacement for cloud AI APIs.

Key features

  • Runs LLM, vision, voice, image and video models
  • Composable backends pulled on demand
  • OpenAI, Anthropic and ElevenLabs API compatibility
  • Runs on CPU, NVIDIA, AMD, Intel and Apple Silicon
  • API key authentication, quotas and role-based access
  • Built-in agents with tools, RAG and MCP

Pricing: Free and open source under the MIT license.

xAI API alternatives: questions

What is the best open-source alternative to xAI API?
Ollama is the top-ranked open-source alternative to xAI API on Enlisted: A tool for downloading and running open large language models locally through a command line and REST API, with Docker support and many integrations. Other strong options are llama.cpp, vLLM and LocalAI.
Are these xAI API alternatives free?
All 4 are open source, so the code is free to use under its licence, and all of them can be self-hosted on your own server or computer.
How is this list of xAI API alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 4 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all