vLLM
A Python library and server for fast, memory-efficient LLM inference and serving, with PagedAttention, continuous batching and broad quantization support.
- GitHub stars
- 93k
- Last commit
- today
- Latest release
- v0.30.0
- Licence
- Apache-2.0
- Self-hosted
- Yes

vLLM is an open-source library for running and serving large language models efficiently. It originated at UC Berkeley's Sky Computing Lab and is now developed by a large community of contributors from academia and industry. It is written in Python and released under the Apache-2.0 license.
Speed comes from techniques such as PagedAttention for attention key-value memory, batching incoming requests continuously, chunked prefill, prefix caching, CUDA and HIP graphs, and optimized attention and mixture-of-experts kernels. It supports many quantization formats, including FP8, INT8, INT4, GPTQ and AWQ, as well as speculative decoding and disaggregated prefill and decode.
On the usability side, vLLM integrates with Hugging Face models and supports parallel sampling, beam search, streaming outputs, structured outputs and tool calling, plus several parallelism modes (tensor, pipeline, data, expert and context) for distributed inference. It targets NVIDIA, AMD and TPU hardware and suits teams that serve open models in production and need high throughput.
Key features
- PagedAttention and continuous batching
- Quantization including FP8, INT8, GPTQ and AWQ
- Speculative decoding support
- Tensor, pipeline and expert parallelism
- OpenAI-compatible API server
- Hugging Face model integration
- Streaming and structured outputs
Pricing: Free and open source under the Apache-2.0 license.



