beta9
Open-source runtime for serverless AI workloads that deploys and scales GPU inference, sandboxes, and background jobs from a Python interface.
- GitHub stars
- 1.8k
- Last commit
- yesterday
- Latest release
- gateway-0.1.797
- Licence
- AGPL-3.0
- Self-hosted
- Yes
- Hosted version
- Available

Beam, whose open-source engine lives in the beta9 repository, is a fast runtime for serverless AI workloads. It gives developers a Pythonic interface to deploy and scale AI applications without managing infrastructure. The code is written in Go and released under the AGPL-3.0 license, and topics include GPU, LLM inference, autoscaling, and functions as a service.
Features include sub-second container cold starts using a custom runtime, scheduler, and embedded caching, parallel fan-out to hundreds of containers, hot reloading, webhooks and scheduled jobs, scale-to-zero by default, and mounted distributed storage volumes. GPU support runs on the Beam cloud, with cards such as the RTX 4090 and H100 mentioned, or on your own GPUs.
The project can be self-hosted or used through the hosted platform, which requires creating an account. Beam suits machine learning engineers and startups that want to serve models, run batch jobs, or sandbox untrusted code without operating a Kubernetes GPU cluster themselves.
Key features
- Serverless GPU inference
- Sub-second container cold starts
- Scale-to-zero workloads
- Parallel fan-out to many containers
- Webhooks and scheduled jobs
- Volume storage and bring-your-own GPUs
Pricing: The Developer plan is $0 per month plus usage and Team is $89 per month plus usage; Growth is custom. Usage is billed by the millisecond, with H100 machines from $1.83 per hour.
