6,598 open-source and SaaS tools, with GitHub stats refreshed every day.

9 alternatives ranked by real activity

Open-source Weights & Biases alternatives

A curated, ranked list of the 9 best open-source alternatives to Weights & Biases.

The best open-source alternative to Weights & Biases is Langfuse. If that doesn't suit you, other good options are MLflow, Opik, Arize Phoenix and Laminar.

Weights & Biases alternatives are mainly AI Infrastructure tools. 6 of them shipped code in the last 30 days, 9 can be self-hosted, and 6 use a permissive licence.

Last updated October 2, 2026 · ranked by GitHub stars, growth and recent commits

Langfuse

An open-source LLM engineering platform for tracing, evaluating and improving AI applications, with prompt management, datasets and a playground, self-hosted or cloud.

GitHub stars
35k
Last commit
today
Latest release
v4.50.0
Self-hosted
Yes
Hosted version
Available
langfuse.comLangfuse homepage screenshot

Langfuse is an open-source LLM engineering platform that helps teams develop, monitor, evaluate and debug AI applications together. You instrument your app, and Langfuse records traces of LLM calls and surrounding logic such as retrieval, embeddings and agent actions, so complex runs and user sessions can be inspected and debugged.

Prompt management lets you centrally version and iterate on prompts, with caching on server and client so that changes do not add latency to your app. Evaluation features cover LLM-as-a-judge, code evaluators, user feedback, manual labeling and custom pipelines through the API and SDKs. Datasets provide test sets and benchmarks, there is an LLM playground for trying prompts, and integrations exist for OpenAI, LangChain and LlamaIndex. It is built on the ClickHouse database, and the Langfuse team has been part of ClickHouse since January 2026.

You can use Langfuse Cloud or self-host it, which the maintainers say takes minutes. The repository license is listed as 'Other' because it combines open-source and enterprise components, so check the terms. It suits teams shipping production LLM features who need visibility into quality, latency and cost.

Key features

  • Tracing of LLM calls, retrieval and agent actions
  • Prompt versioning with caching
  • LLM-as-a-judge and custom evaluations
  • Datasets for tests and benchmarks
  • Interactive LLM playground
  • Integrations with OpenAI, LangChain and LlamaIndex

Pricing: Core has a $29 monthly base that includes 100k units, then $8 per 100k units with volume discounts. Hobby and Pro tiers exist but their prices are not shown in the text; Enterprise is by sales.

Read more about LangfuseWebsite GitHub

MLflow

MLflow is an open-source platform for debugging, evaluating and monitoring AI agents, LLM applications and machine learning models.

GitHub stars
28k
Last commit
today
Latest release
v3.16.1
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
mlflow.orgMLflow homepage screenshot

MLflow is an open-source AI engineering platform for agents, LLM applications and ML models. It helps teams debug, evaluate, monitor and optimize AI applications in production while keeping control of costs and access to models and data. It is written in Python, released under the Apache-2.0 license, and has been developed since 2018, with topics covering MLOps, model management and LLMOps.

For LLM and agent work it offers production observability with traces, plus evaluation, prompt management and prompt optimization, and an AI Gateway that governs costs and model access. It works with Python, TypeScript and JavaScript, Java and other languages, and integrates with OpenTelemetry and MCP. Getting started means starting an MLflow server, enabling logging in your code and running it, then exploring traces and metrics in the web UI on port 5000. A setup wizard can connect to an MLflow server or a Databricks workspace and add tracing with a coding agent.

Key features

  • Tracing and observability for LLM apps and agents
  • Evaluation tools for models and agents
  • Prompt management and prompt optimization
  • AI Gateway for cost and model access control
  • OpenTelemetry and MCP integration
  • Web UI for exploring traces and metrics

Pricing: Free and open source under the Apache-2.0 license; the README also mentions connecting to a Databricks workspace.

Read more about MLflowWebsite GitHub

Opik

Opik is an open-source platform from Comet for tracing, evaluating and monitoring LLM applications, RAG systems and AI agents.

GitHub stars
22k
Last commit
today
Latest release
2.2.88
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
comet.comOpik homepage screenshot

Opik is an open-source LLM observability and evaluation platform built by Comet. It covers the application lifecycle starting with early traces during development and ending with production monitoring, for teams that build LLM apps and AI agents. The code is Python, licensed under Apache-2.0, and the README says the full platform is free to self-host. The project documentation lives on the Comet site.

Capabilities include deep tracing of LLM calls and agent activity including complete trace trees for agents with several steps and tool calls, and evaluation with datasets, experiments and LLM-as-a-judge metrics for tasks like hallucination detection, moderation and RAG assessment. An Agent Optimizer SDK improves prompts and agents, dashboards and online evaluation rules support production monitoring, guardrails help with safe AI practices, and a PyTest integration tests LLM pipelines on each commit. Integrations include frameworks such as LangChain, LlamaIndex and OpenAI clients.

Key features

  • Tracing for LLM calls and agent steps
  • Datasets and experiments for evaluation
  • Evaluation metrics using LLM-as-a-judge
  • Prompt and agent optimization SDK
  • Production dashboards and online evaluation
  • PyTest integration for CI checks

Pricing: Free to self-host, with a free hosted tier of 25k spans. Pro includes 100k spans with extra spans at $5 per 100k; Enterprise is custom.

Read more about OpikWebsite GitHub

Arize Phoenix

Open-source AI observability platform from Arize for tracing, evaluating and troubleshooting LLM applications and agents.

GitHub stars
12k
Last commit
today
Latest release
arize-phoenix-v20.19.0
Self-hosted
Yes
Hosted version
Available

Arize Phoenix is an open-source AI observability platform for experimentation, evaluation and troubleshooting. It helps teams building applications on large language models understand what their systems are doing at runtime, measure quality and debug problems before and after release.

Tracing relies on OpenTelemetry-based instrumentation to record an LLM application's runtime behavior. Evaluation uses LLMs to benchmark an application's performance with response and retrieval evals, and datasets can be versioned. The topics point to integrations with OpenAI, Anthropic, LangChain, LlamaIndex and smolagents, as well as work on agents and prompt engineering.

Phoenix is written in Python and published under a license listed in the repository. For managed production workflows, Arize also offers a separate product called Arize AX. The README is available in English and Simplified Chinese, and the project documentation is hosted on the Arize site.

Key features

  • OpenTelemetry-based tracing for LLM apps
  • LLM-assisted response and retrieval evals
  • Versioned datasets for experiments
  • Integrations with common LLM frameworks
  • Support for agent workflows

Pricing: Phoenix is open source; Arize offers a managed product, Arize AX, for production workflows.

Read more about Arize PhoenixWebsite GitHub

Laminar

Open-source observability platform for AI agents with OpenTelemetry-based tracing, evaluations, dashboards, and plain-English alerts on agent behavior.

GitHub stars
3.3k
Last commit
today
Latest release
v0.2.5
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
laminar.shLaminar homepage screenshot

Laminar is an open-source observability platform built for AI agents. It records what agents do during each run so that teams can debug failures, track behavior, and measure quality. The project is released under the Apache-2.0 license and the backend is written largely in Rust, with a TypeScript front end.

Tracing is based on OpenTelemetry, and a single line of SDK setup can automatically trace libraries such as the Vercel AI SDK, Browser Use, Stagehand, LangChain, OpenAI, Anthropic, and Gemini. Signals let you describe a behavior, such as an agent getting stuck in a loop, in plain English and be notified in Slack when it occurs. Evals run through an SDK and CLI locally or in CI, with a UI for comparing results, and SQL queries and dashboards cover traces, metrics, and events.

Coding agents can reach the data through MCP and CLI access to investigate issues themselves. Laminar can be self-hosted locally with Docker Compose, or used as the managed platform at laminar.sh. It suits teams building production agents who need tracing, evaluation datasets, and annotation in one place.

Key features

  • OpenTelemetry-native tracing SDK
  • Plain-English signals with Slack alerts
  • Evals via SDK and CLI with comparison UI
  • SQL queries over traces, spans, and metrics
  • Custom dashboards for traces and events
  • MCP and CLI access for coding agents
  • Datasets and data annotation tools

Pricing: A free plan includes 1 GB of data. Starter costs $30 and Pro $150 per month with unlimited seats and per-GB overage; Enterprise is custom with an on-premise option.

Read more about LaminarWebsite GitHub

Pezzo

An open-source, cloud-native LLMOps platform for managing prompts, tracking versions, observing AI calls, caching responses and collaborating on LLM features.

GitHub stars
3.3k
Last commit
1 mo ago
Latest release
v0.9.2
Licence
Apache-2.0
Self-hosted
Yes

Pezzo is an open-source LLMOps platform aimed at developers who build features on top of large language models. It gives teams one place to design and manage prompts, track versions, monitor and troubleshoot AI operations, and deliver prompt changes to applications without redeploying code.

The platform provides prompt management, observability and caching, with client libraries for Node.js, Python and LangChain. Teams can publish a new prompt version and have applications pick it up instantly, inspect requests and responses to find problems, and reduce cost and latency through caching. Topics on the repository also reference OpenAI, prompt engineering and monitoring.

Pezzo is built with TypeScript, Node.js and NestJS and depends on open-source infrastructure including PostgreSQL, ClickHouse, Redis and Supertokens, which can be started with Docker Compose. It is licensed under Apache-2.0 and designed to be run on your own infrastructure, with documentation covering architecture, tutorials and provider recipes. It suits product and engineering teams that want a self-managed prompt and observability layer.

Key features

  • Prompt management with versioning
  • Instant delivery of prompt changes
  • Observability of LLM requests
  • Response caching to cut cost and latency
  • Clients for Node.js, Python, and LangChain
  • Docker Compose deployment

Pricing: Free and open source under the Apache-2.0 licence.

Read more about PezzoWebsite GitHub

TraceRoot

TraceRoot is an open-source observability layer for AI agents that turns production traces into detector findings, datasets and evaluations.

GitHub stars
790
Last commit
today
Latest release
v1.1.2
Self-hosted
Yes
Hosted version
Available
traceroot.aiTraceRoot homepage screenshot

TraceRoot describes itself as an open-source self-improving layer for AI agents. It captures traces from production agent systems and turns them into actionable feedback and evaluations, closing a loop with your coding agent so that problems found in production lead to improvements. The project is a Y Combinator S25 company.

Its capabilities include tracing of LLM calls, tool use and agent steps with OpenTelemetry-compatible Python and TypeScript SDKs, covering inputs, outputs, latency, tokens and cost. Detectors screen incoming traces for behaviors such as hallucinations, tool failures and logic errors, with configurable sampling and judge models. Datasets and evaluations let you version test cases, score runs and compare candidate versions, while dashboards and threshold alerts track quality, latency and cost.

A CLI exports traces and findings into your coding workflow, and an in-app AI assistant can explore traces with access to your source code and GitHub context, using a hosted model or your own key. TraceRoot offers a hosted cloud and a self-hosting guide, is written in TypeScript, and its license is listed as Other.

Key features

  • OpenTelemetry-compatible agent tracing
  • Detectors that flag failures in traces
  • Versioned datasets and evaluations
  • Dashboards and threshold alerts
  • CLI for exporting traces and findings
  • In-app AI assistant with code context
Read more about TraceRootWebsite GitHub

UpTrain

UpTrain is a free, Apache-licensed, open-source platform for evaluating and improving generative AI applications with preconfigured checks and root-cause analysis on failures.

GitHub stars
2.4k
Last commit
2 yr ago
Latest release
v0.7.1
Licence
Apache-2.0
Self-hosted
Yes
uptrain.aiUpTrain homepage screenshot

UpTrain is an open-source platform aimed at teams building generative AI applications who need a structured way to evaluate output quality rather than eyeballing responses manually. It targets common LLM reliability problems such as hallucination and jailbreak attempts, alongside more general prompt and output quality checks.

It provides grades for more than 20 preconfigured checks covering language, code and embedding-based use cases, and when a check fails, it performs root cause analysis on the failure case and gives insights on how to resolve it, rather than just reporting a pass/fail score. Its topics also cover experimentation, prompt engineering support, and comparisons to tools such as OpenAI Evals, positioning it within the broader LLMOps evaluation space.

UpTrain is written in Python and released under the Apache-2.0 license, so it can be run and self-hosted as part of an existing ML or LLM pipeline. It is aimed at developers integrating evaluation directly into their generative AI development and monitoring workflow.

Key features

  • 20-plus preconfigured LLM evaluation checks
  • Hallucination and jailbreak detection
  • Root cause analysis on failing checks
  • Embedding, language and code use-case coverage
  • Self-hosted LLMOps evaluation pipeline

Pricing: Free and open source under the Apache-2.0 license.

Read more about UpTrainWebsite GitHub

labml

labml is a free, MIT-licensed open-source toolkit for monitoring deep learning training progress and hardware usage remotely, including from a mobile phone.

GitHub stars
2.3k
Last commit
1 yr ago
Latest release
v0.4.132
Licence
MIT
Self-hosted
Yes
labml.ailabml homepage screenshot

labml is an experiment-tracking and monitoring toolkit aimed at machine learning practitioners who want to check on long-running training jobs, including from their phone, without staying tied to a terminal or a single workstation. It integrates with common deep learning frameworks and tracks experiment metadata alongside training metrics.

With as little as two lines of code, it can track git commit information, configuration values and hyperparameters for an experiment, alongside the usual loss and metric curves, and it works with PyTorch, PyTorch Lightning, Keras, TensorFlow and fastai based on its listed topics. A single command also monitors hardware usage on any machine. An optional self-hosted experiments server, backed by MongoDB and optionally fronted by Nginx, provides the web interface for viewing experiments remotely, including from a mobile browser.

labml is written in Python and released under the MIT license. It is installed via pip, both for the monitoring client embedded in training scripts and for the self-hosted server component, and its documentation covers the Python API, custom visualizations and configuration management in more depth.

Key features

  • Remote experiment and hardware monitoring
  • Two-line integration into training scripts
  • Tracks git commit, config and hyperparameters
  • Works with PyTorch, Keras, TensorFlow and fastai
  • Mobile-friendly web dashboard
  • Self-hosted MongoDB-backed server

Pricing: Free and open source under the MIT license.

Read more about labmlWebsite GitHub

Weights & Biases alternatives: questions

What is the best open-source alternative to Weights & Biases?
Langfuse is the top-ranked open-source alternative to Weights & Biases on Enlisted: An open-source LLM engineering platform for tracing, evaluating and improving AI applications, with prompt management, datasets and a playground, self-hosted or cloud. Other strong options are MLflow, Opik, Arize Phoenix and Laminar.
Are these Weights & Biases alternatives free?
All 9 are open source, so the code is free to use under its licence, and 9 of them can be self-hosted on your own server. 6 also offer a paid or managed cloud version if you'd rather not host it yourself.
How is this list of Weights & Biases alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 6 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all