7,380 open-source and SaaS tools, with GitHub stats refreshed every day.

7 alternatives ranked by real activity

Open-source Braintrust alternatives

A curated, ranked list of the 7 best open-source alternatives to Braintrust.

The best open-source alternative to Braintrust is Langfuse. If that doesn't suit you, other good options are Promptfoo, Opik, Arize Phoenix and LangWatch.

Braintrust alternatives are mainly AI infrastructure tools. 6 of them shipped code in the last 30 days, 7 can be self-hosted, and 5 use a permissive licence.

Last updated October 3, 2026 · ranked by GitHub stars, growth and recent commits

Langfuse

An open-source LLM engineering platform for tracing, evaluating and improving AI applications, with prompt management, datasets and a playground, self-hosted or cloud.

GitHub stars
35k
Last commit
today
Latest release
v4.50.0
Self-hosted
Yes
Hosted version
Available
langfuse.comLangfuse homepage screenshot

Langfuse is an open-source LLM engineering platform that helps teams develop, monitor, evaluate and debug AI applications together. You instrument your app, and Langfuse records traces of LLM calls and surrounding logic such as retrieval, embeddings and agent actions, so complex runs and user sessions can be inspected and debugged.

Prompt management lets you centrally version and iterate on prompts, with caching on server and client so that changes do not add latency to your app. Evaluation features cover LLM-as-a-judge, code evaluators, user feedback, manual labeling and custom pipelines through the API and SDKs. Datasets provide test sets and benchmarks, there is an LLM playground for trying prompts, and integrations exist for OpenAI, LangChain and LlamaIndex. It is built on the ClickHouse database, and the Langfuse team has been part of ClickHouse since January 2026.

You can use Langfuse Cloud or self-host it, which the maintainers say takes minutes. The repository license is listed as 'Other' because it combines open-source and enterprise components, so check the terms. It suits teams shipping production LLM features who need visibility into quality, latency and cost.

Key features

  • Tracing of LLM calls, retrieval and agent actions
  • Prompt versioning with caching
  • LLM-as-a-judge and custom evaluations
  • Datasets for tests and benchmarks
  • Interactive LLM playground
  • Integrations with OpenAI, LangChain and LlamaIndex

Pricing: Core has a $29 monthly base that includes 100k units, then $8 per 100k units with volume discounts. Hobby and Pro tiers exist but their prices are not shown in the text; Enterprise is by sales.

Promptfoo

Open-source CLI and library for evaluating prompts, agents, and RAG apps, and for red-teaming LLM applications to find security weaknesses.

GitHub stars
26k
Last commit
today
Latest release
0.123.1
Licence
MIT
Self-hosted
Yes
promptfoo.devPromptfoo homepage screenshot

Promptfoo is an open-source command-line tool and library for testing applications built on large language models. Developers describe test cases and checks in declarative configuration files, run them against one or more models, and review the results, which replaces manual trial and error with repeatable evaluations. It is written in TypeScript and released under the MIT license.

Beyond quality testing, Promptfoo includes red teaming and vulnerability scanning for AI systems, probing prompts, agents, and retrieval-augmented generation (RAG) pipelines for security and safety weaknesses. It can compare models from providers such as OpenAI, Anthropic, Azure, Bedrock, and Ollama side by side, and it works with any LLM API or programming language. Its documentation also covers a code-scanning feature that reviews pull requests for LLM-related security and compliance issues.

The project is designed to run locally: the maintainers state that evaluations execute on the developer's own machine, so prompts only leave it when a chosen model provider receives them. It can be installed with Homebrew or pip, or run without installing via npx, and it plugs into CI/CD so checks run automatically on each change. The README notes that Promptfoo is now part of OpenAI while remaining open source and MIT licensed.

Key features

  • Declarative test configs for prompts and models
  • Side-by-side comparison of multiple LLM providers
  • Red teaming and vulnerability scanning for AI apps
  • Evaluations run locally on your own machine
  • CI/CD integration for automated checks
  • Pull request code scanning for LLM risks
  • Live reload and result caching
  • Shareable results and security reports

Pricing: The Community edition is free forever. Enterprise and On-Premise plans have custom pricing arranged through a demo or by contacting the vendor.

Opik

Opik is an open-source platform from Comet for tracing, evaluating and monitoring LLM applications, RAG systems and AI agents.

GitHub stars
22k
Last commit
today
Latest release
2.2.88
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
comet.comOpik homepage screenshot

Opik is an open-source LLM observability and evaluation platform built by Comet. It covers the application lifecycle starting with early traces during development and ending with production monitoring, for teams that build LLM apps and AI agents. The code is Python, licensed under Apache-2.0, and the README says the full platform is free to self-host. The project documentation lives on the Comet site.

Capabilities include deep tracing of LLM calls and agent activity including complete trace trees for agents with several steps and tool calls, and evaluation with datasets, experiments and LLM-as-a-judge metrics for tasks like hallucination detection, moderation and RAG assessment. An Agent Optimizer SDK improves prompts and agents, dashboards and online evaluation rules support production monitoring, guardrails help with safe AI practices, and a PyTest integration tests LLM pipelines on each commit. Integrations include frameworks such as LangChain, LlamaIndex and OpenAI clients.

Key features

  • Tracing for LLM calls and agent steps
  • Datasets and experiments for evaluation
  • Evaluation metrics using LLM-as-a-judge
  • Prompt and agent optimization SDK
  • Production dashboards and online evaluation
  • PyTest integration for CI checks

Pricing: Free to self-host, with a free hosted tier of 25k spans. Pro includes 100k spans with extra spans at $5 per 100k; Enterprise is custom.

Read more about OpikWebsite GitHub

Arize Phoenix

Open-source AI observability platform from Arize for tracing, evaluating and troubleshooting LLM applications and agents.

GitHub stars
12k
Last commit
today
Latest release
arize-phoenix-v20.19.0
Self-hosted
Yes
Hosted version
Available

Arize Phoenix is an open-source AI observability platform for experimentation, evaluation and troubleshooting. It helps teams building applications on large language models understand what their systems are doing at runtime, measure quality and debug problems before and after release.

Tracing relies on OpenTelemetry-based instrumentation to record an LLM application's runtime behavior. Evaluation uses LLMs to benchmark an application's performance with response and retrieval evals, and datasets can be versioned. The topics point to integrations with OpenAI, Anthropic, LangChain, LlamaIndex and smolagents, as well as work on agents and prompt engineering.

Phoenix is written in Python and published under a license listed in the repository. For managed production workflows, Arize also offers a separate product called Arize AX. The README is available in English and Simplified Chinese, and the project documentation is hosted on the Arize site.

Key features

  • OpenTelemetry-based tracing for LLM apps
  • LLM-assisted response and retrieval evals
  • Versioned datasets for experiments
  • Integrations with common LLM frameworks
  • Support for agent workflows

Pricing: Phoenix is open source; Arize offers a managed product, Arize AX, for production workflows.

LangWatch

LangWatch is an open-source platform for tracing, evaluating and testing LLM applications and AI agents, with prompt management and an AI gateway.

GitHub stars
4.9k
Last commit
today
Latest release
langwatch-3.20.1
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
langwatch.aiLangWatch homepage screenshot

LangWatch is an Apache-2.0 platform for running AI in production. It lets teams trace, test, route and govern LLM calls across the company, covering both the agents they build themselves and the coding assistants their engineers use. The project is aimed at AI engineers and platform teams that need visibility and control over model usage.

Its feature areas include LLM operations such as observability, agent testing, evaluations and prompt management, plus tracking of coding agent sessions with cost per pull request and per team and privacy controls. An AI gateway offers a single endpoint compatible with the OpenAI and Anthropic APIs, with virtual keys, budgets and routing, and an AI governance layer sits on top. Repository topics mention DSPy, datasets, simulation testing and low-code evaluation.

LangWatch is written in TypeScript and can be used through LangWatch Cloud or self-hosted; the README says only Node.js is required to try it locally, with separate guidance for production deployments. It suits teams moving LLM prototypes into monitored, tested production systems.

Key features

  • LLM observability and tracing
  • Agent testing and simulation
  • Evaluations over datasets
  • Managing and testing prompts
  • AI gateway with virtual keys and budgets
  • Coding agent cost tracking per team

Pricing: Free Developer plan with 50k events a month. Growth costs €29 per core-seat per month plus €5 per 100k events beyond 200k; Enterprise is custom. Self-hosting is supported.

Laminar

Open-source observability platform for AI agents with OpenTelemetry-based tracing, evaluations, dashboards, and plain-English alerts on agent behavior.

GitHub stars
3.3k
Last commit
today
Latest release
v0.2.5
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
laminar.shLaminar homepage screenshot

Laminar is an open-source observability platform built for AI agents. It records what agents do during each run so that teams can debug failures, track behavior, and measure quality. The project is released under the Apache-2.0 license and the backend is written largely in Rust, with a TypeScript front end.

Tracing is based on OpenTelemetry, and a single line of SDK setup can automatically trace libraries such as the Vercel AI SDK, Browser Use, Stagehand, LangChain, OpenAI, Anthropic, and Gemini. Signals let you describe a behavior, such as an agent getting stuck in a loop, in plain English and be notified in Slack when it occurs. Evals run through an SDK and CLI locally or in CI, with a UI for comparing results, and SQL queries and dashboards cover traces, metrics, and events.

Coding agents can reach the data through MCP and CLI access to investigate issues themselves. Laminar can be self-hosted locally with Docker Compose, or used as the managed platform at laminar.sh. It suits teams building production agents who need tracing, evaluation datasets, and annotation in one place.

Key features

  • OpenTelemetry-native tracing SDK
  • Plain-English signals with Slack alerts
  • Evals via SDK and CLI with comparison UI
  • SQL queries over traces, spans, and metrics
  • Custom dashboards for traces and events
  • MCP and CLI access for coding agents
  • Datasets and data annotation tools

Pricing: A free plan includes 1 GB of data. Starter costs $30 and Pro $150 per month with unlimited seats and per-GB overage; Enterprise is custom with an on-premise option.

UpTrain

UpTrain is a free, Apache-licensed, open-source platform for evaluating and improving generative AI applications with preconfigured checks and root-cause analysis on failures.

GitHub stars
2.4k
Last commit
2 yr ago
Latest release
v0.7.1
Licence
Apache-2.0
Self-hosted
Yes
uptrain.aiUpTrain homepage screenshot

UpTrain is an open-source platform aimed at teams building generative AI applications who need a structured way to evaluate output quality rather than eyeballing responses manually. It targets common LLM reliability problems such as hallucination and jailbreak attempts, alongside more general prompt and output quality checks.

It provides grades for more than 20 preconfigured checks covering language, code and embedding-based use cases, and when a check fails, it performs root cause analysis on the failure case and gives insights on how to resolve it, rather than just reporting a pass/fail score. Its topics also cover experimentation, prompt engineering support, and comparisons to tools such as OpenAI Evals, positioning it within the broader LLMOps evaluation space.

UpTrain is written in Python and released under the Apache-2.0 license, so it can be run and self-hosted as part of an existing ML or LLM pipeline. It is aimed at developers integrating evaluation directly into their generative AI development and monitoring workflow.

Key features

  • 20-plus preconfigured LLM evaluation checks
  • Hallucination and jailbreak detection
  • Root cause analysis on failing checks
  • Embedding, language and code use-case coverage
  • Self-hosted LLMOps evaluation pipeline

Pricing: Free and open source under the Apache-2.0 license.

Braintrust alternatives: questions

What is the best open-source alternative to Braintrust?
Langfuse is the top-ranked open-source alternative to Braintrust on Enlisted: An open-source LLM engineering platform for tracing, evaluating and improving AI applications, with prompt management, datasets and a playground, self-hosted or cloud. Other strong options are Promptfoo, Opik, Arize Phoenix and LangWatch.
Are these Braintrust alternatives free?
All 7 are open source, so the code is free to use under its licence, and all of them can be self-hosted on your own server or computer. 5 also offer a paid or managed cloud version if you'd rather not host it yourself.
How is this list of Braintrust alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 6 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all