7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

vLLM

Open source

A Python library and server for fast, memory-efficient LLM inference and serving, with PagedAttention, continuous batching and broad quantization support.

vllm.ai
vLLM homepage screenshot
GitHub stars
93k
Last commit
today
Repository age
3 years
Version
v0.30.0
Licence
Apache-2.0
Self-hosted
Yes

About vLLM

vLLM is an open-source library for running and serving large language models efficiently. It originated at UC Berkeley's Sky Computing Lab and is now developed by a large community of contributors from academia and industry. It is written in Python and released under the Apache-2.0 license.

Speed comes from techniques such as PagedAttention for attention key-value memory, batching incoming requests continuously, chunked prefill, prefix caching, CUDA and HIP graphs, and optimized attention and mixture-of-experts kernels. It supports many quantization formats, including FP8, INT8, INT4, GPTQ and AWQ, as well as speculative decoding and disaggregated prefill and decode.

On the usability side, vLLM integrates with Hugging Face models and supports parallel sampling, beam search, streaming outputs, structured outputs and tool calling, plus several parallelism modes (tensor, pipeline, data, expert and context) for distributed inference. It targets NVIDIA, AMD and TPU hardware and suits teams that serve open models in production and need high throughput.

Key features

  • PagedAttention and continuous batching
  • Quantization including FP8, INT8, GPTQ and AWQ
  • Speculative decoding support
  • Tensor, pipeline and expert parallelism
  • OpenAI-compatible API server
  • Hugging Face model integration
  • Streaming and structured outputs

Good fit for

  • →Serving open LLMs in production
  • →Maximizing GPU throughput for inference
  • →Distributed inference across multiple GPUs
Built with
Python
Tags
llm-serving
inference
pytorch
gpu
quantization
python
model-serving
self-hosted

vLLM: questions and answers

What is vLLM used for?
vLLM is a Python library and server for fast, memory-efficient LLM inference and serving, with PagedAttention, continuous batching and broad quantization support. It is a good fit for serving open LLMs in production, maximizing GPU throughput for inference and distributed inference across multiple GPUs.
Is vLLM open source?
Yes. vLLM is open source under the Apache-2.0 licence. Its source code is on GitHub at vllm-project/vllm and is written mainly in Python.
Is vLLM free?
Yes. vLLM is open source, so the software itself is free to use.
Can I self-host vLLM?
Yes. vLLM can be self-hosted on your own server or infrastructure; there is no official hosted version.
What is vLLM an alternative to?
vLLM is an open-source alternative to Amazon Bedrock, Amazon SageMaker, Replicate and Together AI. Other open-source alternatives to Amazon Bedrock include SGLang and Dify.
Is vLLM actively maintained?
Yes. The most recent commit to vLLM was on 2 October 2026, and the latest release is v0.30.0, published on 22 September 2026. The project has 93k stars on GitHub.

Open-source alternatives to vLLM

See all

SaaS alternatives to vLLM

See all