7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

Paddler

Open source

An open-source load balancer and serving platform for self-hosted LLMs, built on llama.cpp and delivered as a single Rust binary.

paddler.intentee.com
Paddler homepage screenshot
GitHub stars
1.7k
Last commit
2 days ago
Repository age
2 years
Version
v4.1.0
Licence
Apache-2.0
Self-hosted
Yes

About Paddler

Paddler is an open-source load balancer and serving platform for large language models and vision-language models that you run on your own infrastructure. It lets teams run inference, deploy and scale models without sending data to a closed-source model provider, which helps with privacy, reliability and predictable costs.

It uses a built-in llama.cpp engine for inference and applies load balancing designed for LLM requests. Agents can be added dynamically, so it can work with autoscaling tools, and request buffering allows scaling up from zero hosts. Models can be swapped at runtime, and a web admin panel provides management, monitoring and testing, along with observability metrics. It runs on CPU and GPU.

Paddler ships as a single binary written in Rust and is licensed under Apache-2.0. The README positions it as a simpler alternative to projects such as llm-d and Docker Model Runner, with fewer moving parts and deployments built around the ggml ecosystem. It targets product teams needing embeddings and inference, LLMOps teams, and organizations with strict compliance needs.

Key features

  • LLM-aware load balancing
  • Built-in llama.cpp inference engine
  • Dynamic agents for autoscaling integration
  • Request buffering to scale from zero
  • Runtime model swapping
  • Web admin panel with metrics

Good fit for

  • →Serving private LLMs across several servers
  • →Keeping sensitive data off third-party APIs
  • →Replacing per-token pricing with fixed infrastructure
Built with
Rust
Tags
llm
load-balancer
llama-cpp
inference
llmops
self-hosted
rust
ai-infrastructure
apache-2.0

Paddler: questions and answers

What is Paddler used for?
Paddler is an open-source load balancer and serving platform for self-hosted LLMs, built on llama.cpp and delivered as a single Rust binary. It is a good fit for serving private LLMs across several servers, keeping sensitive data off third-party APIs and replacing per-token pricing with fixed infrastructure.
Is Paddler open source?
Yes. Paddler is open source under the Apache-2.0 licence. Its source code is on GitHub at intentee/paddler and is written mainly in Rust.
Is Paddler free?
Yes. Paddler is open source, so the software itself is free to use.
Can I self-host Paddler?
Yes. Paddler can be self-hosted on your own server or infrastructure; there is no official hosted version.
What are some alternatives to Paddler?
Similar open-source tools in the AI Infrastructure category include Ollama, llama.cpp and SGLang. SaaS products in the same category include Cloudflare Workers AI, AI Stats and api-hub.ai.
Is Paddler actively maintained?
Yes. The most recent commit to Paddler was on 1 October 2026, and the latest release is v4.1.0, published on 19 July 2026. The project has 1.7k stars on GitHub.

Open-source alternatives to Paddler

See all

SaaS alternatives to Paddler

See all