7,380 open-source and SaaS tools, with GitHub stats refreshed every day.

4 alternatives ranked by real activity

Open-source Together AI alternatives

A curated, ranked list of the 4 best open-source alternatives to Together AI.

The best open-source alternative to Together AI is vLLM. If that doesn't suit you, other good options are Unsloth, LocalAI and SGLang.

Together AI alternatives are mainly AI infrastructure tools. 4 of them shipped code in the last 30 days, 4 can be self-hosted, and 4 use a permissive licence.

Last updated October 3, 2026 · ranked by GitHub stars, growth and recent commits

vLLM

A Python library and server for fast, memory-efficient LLM inference and serving, with PagedAttention, continuous batching and broad quantization support.

GitHub stars
93k
Last commit
today
Latest release
v0.30.0
Licence
Apache-2.0
Self-hosted
Yes
vllm.aivLLM homepage screenshot

vLLM is an open-source library for running and serving large language models efficiently. It originated at UC Berkeley's Sky Computing Lab and is now developed by a large community of contributors from academia and industry. It is written in Python and released under the Apache-2.0 license.

Speed comes from techniques such as PagedAttention for attention key-value memory, batching incoming requests continuously, chunked prefill, prefix caching, CUDA and HIP graphs, and optimized attention and mixture-of-experts kernels. It supports many quantization formats, including FP8, INT8, INT4, GPTQ and AWQ, as well as speculative decoding and disaggregated prefill and decode.

On the usability side, vLLM integrates with Hugging Face models and supports parallel sampling, beam search, streaming outputs, structured outputs and tool calling, plus several parallelism modes (tensor, pipeline, data, expert and context) for distributed inference. It targets NVIDIA, AMD and TPU hardware and suits teams that serve open models in production and need high throughput.

Key features

  • PagedAttention and continuous batching
  • Quantization including FP8, INT8, GPTQ and AWQ
  • Speculative decoding support
  • Tensor, pipeline and expert parallelism
  • OpenAI-compatible API server
  • Hugging Face model integration
  • Streaming and structured outputs

Pricing: Free and open source under the Apache-2.0 license.

Read more about vLLMWebsite GitHub

Unsloth

An open-source app and framework for running and fine-tuning LLMs and diffusion models locally, with GGUF and MLX support and an OpenAI-compatible API.

GitHub stars
77k
Last commit
today
Latest release
v0.1.902-beta
Licence
Apache-2.0
Self-hosted
Yes
unsloth.aiUnsloth homepage screenshot

Unsloth is an open-source framework with a desktop app for running and training AI models on your own hardware. The same tool covers inference of local language, image, audio and embedding models as well as fine-tuning them, and it runs on Windows, Linux (including WSL) and macOS, with support for NVIDIA, AMD and Intel GPUs, CPUs and a Vulkan backend.

On the running side it loads GGUF and MLX models, serves them through an OpenAI-compatible API, and lets coding agents such as Claude Code and Codex use local models with tool calling and code execution. It also offers private web search, deep research and retrieval, plus access from other devices over a LAN or a secure remote link. For training it supports LoRA, QLoRA, full fine-tuning, pretraining and reinforcement learning methods such as GRPO and DPO, with export to formats including GGUF.

The project claims fine-tuning that is roughly twice as fast with about 70 percent less VRAM, which was its original draw. Unsloth is Apache-2.0 licensed, can be installed from native packages or a Docker image, and has notebooks and documentation for getting started.

Key features

  • Run GGUF and MLX models locally
  • Fine-tune LLMs with LoRA and QLoRA
  • Reinforcement learning with GRPO and DPO
  • OpenAI-compatible API for local models
  • Native desktop app and Docker image
  • Export trained models to GGUF

Pricing: Free and open source under the Apache-2.0 license.

Read more about UnslothWebsite GitHub

LocalAI

An open-source AI engine that runs LLMs, vision, voice, image and video models on your own hardware behind OpenAI-compatible APIs, with no GPU required.

GitHub stars
49k
Last commit
today
Latest release
v4.10.0
Licence
MIT
Self-hosted
Yes
localai.ioLocalAI homepage screenshot

LocalAI is an open-source AI engine for running models of many kinds, including language, vision, voice, image and video models, on your own hardware. A GPU is not required. It is written in Go and released under the MIT license, was started by Ettore Di Giacinto and is looked after by the LocalAI team.

Its design is a small core with separate backends that are pulled on demand, so nothing unused gets installed. Backends wrap engines such as llama.cpp, vLLM, whisper.cpp, stable-diffusion and MLX, and you can write your own in any language against an open interface. It can run on CPU alone, through Vulkan, or on NVIDIA, AMD, Intel and Apple Silicon hardware.

APIs are drop-in compatible with OpenAI, Anthropic and ElevenLabs across every backend. Multi-user features include API key authentication, user quotas and role-based access, and built-in agents support tool use, RAG, MCP and skills. The project stresses that data stays on your infrastructure. It suits developers and teams who want a private, self-hosted replacement for cloud AI APIs.

Key features

  • Runs LLM, vision, voice, image and video models
  • Composable backends pulled on demand
  • OpenAI, Anthropic and ElevenLabs API compatibility
  • Runs on CPU, NVIDIA, AMD, Intel and Apple Silicon
  • API key authentication, quotas and role-based access
  • Built-in agents with tools, RAG and MCP

Pricing: Free and open source under the MIT license.

Read more about LocalAIWebsite GitHub

SGLang

SGLang is an open-source framework for serving language, vision-language and diffusion models, built for fast inference at scale.

GitHub stars
37k
Last commit
today
Latest release
v0.5.21
Licence
Apache-2.0
Self-hosted
Yes
sglang.ioSGLang homepage screenshot

SGLang is an inference framework, released as open source, for language models, vision-language models and diffusion models. It is optimized for agentic workloads, reinforcement-learning rollouts and large-scale serving. The project includes SGLang Diffusion, a built-in engine for image and video generation that ships in the same repository and Python package. It is released under the Apache-2.0 license.

Getting started takes either a prebuilt Docker image or a Python install using uv, followed by a launch command for the chosen model. A cookbook helps pick a model and hardware pair and generates a ready-to-run command. Supported hardware spans NVIDIA and AMD GPUs, Google TPUs, Intel GPUs and CPUs, Apple Silicon through Metal and MLX, and Huawei Ascend NPUs, with more integrations in progress. The wider SGLang ecosystem adds educational projects and community events.

Key features

  • Serving for LLMs, vision-language and diffusion models
  • Docker image and uv-based Python install
  • Cookbook with ready-to-run launch commands
  • Runs on NVIDIA, AMD, TPU, Intel and Ascend hardware
  • Built-in image and video generation engine
  • Tuned for agentic and RL rollout workloads

Pricing: Free and open source under the Apache-2.0 license.

Read more about SGLangWebsite GitHub

Together AI alternatives: questions

What is the best open-source alternative to Together AI?
vLLM is the top-ranked open-source alternative to Together AI on Enlisted: A Python library and server for fast, memory-efficient LLM inference and serving, with PagedAttention, continuous batching and broad quantization support. Other strong options are Unsloth, LocalAI and SGLang.
Are these Together AI alternatives free?
All 4 are open source, so the code is free to use under its licence, and all of them can be self-hosted on your own server or computer.
How is this list of Together AI alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 4 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all