7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

7 alternatives ranked by real activity

Open-source Arize AI alternatives

A curated, ranked list of the 7 best open-source alternatives to Arize AI.

The best open-source alternative to Arize AI is Langfuse. If that doesn't suit you, other good options are MLflow, Opik, Arize Phoenix and LangWatch.

Arize AI alternatives are mainly AI infrastructure tools. 7 of them shipped code in the last 30 days, 7 can be self-hosted, and 4 use a permissive licence.

Last updated October 3, 2026 · ranked by GitHub stars, growth and recent commits

Langfuse

An open-source LLM engineering platform for tracing, evaluating and improving AI applications, with prompt management, datasets and a playground, self-hosted or cloud.

GitHub stars
35k
Last commit
yesterday
Latest release
v4.50.0
Self-hosted
Yes
Hosted version
Available
langfuse.comLangfuse homepage screenshot

Langfuse is an open-source LLM engineering platform that helps teams develop, monitor, evaluate and debug AI applications together. You instrument your app, and Langfuse records traces of LLM calls and surrounding logic such as retrieval, embeddings and agent actions, so complex runs and user sessions can be inspected and debugged.

Prompt management lets you centrally version and iterate on prompts, with caching on server and client so that changes do not add latency to your app. Evaluation features cover LLM-as-a-judge, code evaluators, user feedback, manual labeling and custom pipelines through the API and SDKs. Datasets provide test sets and benchmarks, there is an LLM playground for trying prompts, and integrations exist for OpenAI, LangChain and LlamaIndex. It is built on the ClickHouse database, and the Langfuse team has been part of ClickHouse since January 2026.

You can use Langfuse Cloud or self-host it, which the maintainers say takes minutes. The repository license is listed as 'Other' because it combines open-source and enterprise components, so check the terms. It suits teams shipping production LLM features who need visibility into quality, latency and cost.

Key features

  • Tracing of LLM calls, retrieval and agent actions
  • Prompt versioning with caching
  • LLM-as-a-judge and custom evaluations
  • Datasets for tests and benchmarks
  • Interactive LLM playground
  • Integrations with OpenAI, LangChain and LlamaIndex

Pricing: Core has a $29 monthly base that includes 100k units, then $8 per 100k units with volume discounts. Hobby and Pro tiers exist but their prices are not shown in the text; Enterprise is by sales.

MLflow

MLflow is an open-source platform for debugging, evaluating and monitoring AI agents, LLM applications and machine learning models.

GitHub stars
28k
Last commit
today
Latest release
v3.16.1
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
mlflow.orgMLflow homepage screenshot

MLflow is an open-source AI engineering platform for agents, LLM applications and ML models. It helps teams debug, evaluate, monitor and optimize AI applications in production while keeping control of costs and access to models and data. It is written in Python, released under the Apache-2.0 license, and has been developed since 2018, with topics covering MLOps, model management and LLMOps.

For LLM and agent work it offers production observability with traces, plus evaluation, prompt management and prompt optimization, and an AI Gateway that governs costs and model access. It works with Python, TypeScript and JavaScript, Java and other languages, and integrates with OpenTelemetry and MCP. Getting started means starting an MLflow server, enabling logging in your code and running it, then exploring traces and metrics in the web UI on port 5000. A setup wizard can connect to an MLflow server or a Databricks workspace and add tracing with a coding agent.

Key features

  • Tracing and observability for LLM apps and agents
  • Evaluation tools for models and agents
  • Prompt management and prompt optimization
  • AI Gateway for cost and model access control
  • OpenTelemetry and MCP integration
  • Web UI for exploring traces and metrics

Pricing: Free and open source under the Apache-2.0 license; the README also mentions connecting to a Databricks workspace.

Opik

Opik is an open-source platform from Comet for tracing, evaluating and monitoring LLM applications, RAG systems and AI agents.

GitHub stars
22k
Last commit
today
Latest release
2.2.88
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
comet.comOpik homepage screenshot

Opik is an open-source LLM observability and evaluation platform built by Comet. It covers the application lifecycle starting with early traces during development and ending with production monitoring, for teams that build LLM apps and AI agents. The code is Python, licensed under Apache-2.0, and the README says the full platform is free to self-host. The project documentation lives on the Comet site.

Capabilities include deep tracing of LLM calls and agent activity including complete trace trees for agents with several steps and tool calls, and evaluation with datasets, experiments and LLM-as-a-judge metrics for tasks like hallucination detection, moderation and RAG assessment. An Agent Optimizer SDK improves prompts and agents, dashboards and online evaluation rules support production monitoring, guardrails help with safe AI practices, and a PyTest integration tests LLM pipelines on each commit. Integrations include frameworks such as LangChain, LlamaIndex and OpenAI clients.

Key features

  • Tracing for LLM calls and agent steps
  • Datasets and experiments for evaluation
  • Evaluation metrics using LLM-as-a-judge
  • Prompt and agent optimization SDK
  • Production dashboards and online evaluation
  • PyTest integration for CI checks

Pricing: Free to self-host, with a free hosted tier of 25k spans. Pro includes 100k spans with extra spans at $5 per 100k; Enterprise is custom.

Arize Phoenix

Open-source AI observability platform from Arize for tracing, evaluating and troubleshooting LLM applications and agents.

GitHub stars
12k
Last commit
today
Latest release
arize-phoenix-v20.19.0
Self-hosted
Yes
Hosted version
Available

Arize Phoenix is an open-source AI observability platform for experimentation, evaluation and troubleshooting. It helps teams building applications on large language models understand what their systems are doing at runtime, measure quality and debug problems before and after release.

Tracing relies on OpenTelemetry-based instrumentation to record an LLM application's runtime behavior. Evaluation uses LLMs to benchmark an application's performance with response and retrieval evals, and datasets can be versioned. The topics point to integrations with OpenAI, Anthropic, LangChain, LlamaIndex and smolagents, as well as work on agents and prompt engineering.

Phoenix is written in Python and published under a license listed in the repository. For managed production workflows, Arize also offers a separate product called Arize AX. The README is available in English and Simplified Chinese, and the project documentation is hosted on the Arize site.

Key features

  • OpenTelemetry-based tracing for LLM apps
  • LLM-assisted response and retrieval evals
  • Versioned datasets for experiments
  • Integrations with common LLM frameworks
  • Support for agent workflows

Pricing: Phoenix is open source; Arize offers a managed product, Arize AX, for production workflows.

LangWatch

LangWatch is an open-source platform for tracing, evaluating and testing LLM applications and AI agents, with prompt management and an AI gateway.

GitHub stars
4.9k
Last commit
yesterday
Latest release
langwatch-3.20.1
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
langwatch.aiLangWatch homepage screenshot

LangWatch is an Apache-2.0 platform for running AI in production. It lets teams trace, test, route and govern LLM calls across the company, covering both the agents they build themselves and the coding assistants their engineers use. The project is aimed at AI engineers and platform teams that need visibility and control over model usage.

Its feature areas include LLM operations such as observability, agent testing, evaluations and prompt management, plus tracking of coding agent sessions with cost per pull request and per team and privacy controls. An AI gateway offers a single endpoint compatible with the OpenAI and Anthropic APIs, with virtual keys, budgets and routing, and an AI governance layer sits on top. Repository topics mention DSPy, datasets, simulation testing and low-code evaluation.

LangWatch is written in TypeScript and can be used through LangWatch Cloud or self-hosted; the README says only Node.js is required to try it locally, with separate guidance for production deployments. It suits teams moving LLM prototypes into monitored, tested production systems.

Key features

  • LLM observability and tracing
  • Agent testing and simulation
  • Evaluations over datasets
  • Managing and testing prompts
  • AI gateway with virtual keys and budgets
  • Coding agent cost tracking per team

Pricing: Free Developer plan with 50k events a month. Growth costs €29 per core-seat per month plus €5 per 100k events beyond 200k; Enterprise is custom. Self-hosting is supported.

Laminar

Open-source observability platform for AI agents with OpenTelemetry-based tracing, evaluations, dashboards, and plain-English alerts on agent behavior.

GitHub stars
3.3k
Last commit
yesterday
Latest release
v0.2.5
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
laminar.shLaminar homepage screenshot

Laminar is an open-source observability platform built for AI agents. It records what agents do during each run so that teams can debug failures, track behavior, and measure quality. The project is released under the Apache-2.0 license and the backend is written largely in Rust, with a TypeScript front end.

Tracing is based on OpenTelemetry, and a single line of SDK setup can automatically trace libraries such as the Vercel AI SDK, Browser Use, Stagehand, LangChain, OpenAI, Anthropic, and Gemini. Signals let you describe a behavior, such as an agent getting stuck in a loop, in plain English and be notified in Slack when it occurs. Evals run through an SDK and CLI locally or in CI, with a UI for comparing results, and SQL queries and dashboards cover traces, metrics, and events.

Coding agents can reach the data through MCP and CLI access to investigate issues themselves. Laminar can be self-hosted locally with Docker Compose, or used as the managed platform at laminar.sh. It suits teams building production agents who need tracing, evaluation datasets, and annotation in one place.

Key features

  • OpenTelemetry-native tracing SDK
  • Plain-English signals with Slack alerts
  • Evals via SDK and CLI with comparison UI
  • SQL queries over traces, spans, and metrics
  • Custom dashboards for traces and events
  • MCP and CLI access for coding agents
  • Datasets and data annotation tools

Pricing: A free plan includes 1 GB of data. Starter costs $30 and Pro $150 per month with unlimited seats and per-GB overage; Enterprise is custom with an on-premise option.

TraceRoot

TraceRoot is an open-source observability layer for AI agents that turns production traces into detector findings, datasets and evaluations.

GitHub stars
790
Last commit
today
Latest release
v1.1.2
Self-hosted
Yes
Hosted version
Available
traceroot.aiTraceRoot homepage screenshot

TraceRoot describes itself as an open-source self-improving layer for AI agents. It captures traces from production agent systems and turns them into actionable feedback and evaluations, closing a loop with your coding agent so that problems found in production lead to improvements. The project is a Y Combinator S25 company.

Its capabilities include tracing of LLM calls, tool use and agent steps with OpenTelemetry-compatible Python and TypeScript SDKs, covering inputs, outputs, latency, tokens and cost. Detectors screen incoming traces for behaviors such as hallucinations, tool failures and logic errors, with configurable sampling and judge models. Datasets and evaluations let you version test cases, score runs and compare candidate versions, while dashboards and threshold alerts track quality, latency and cost.

A CLI exports traces and findings into your coding workflow, and an in-app AI assistant can explore traces with access to your source code and GitHub context, using a hosted model or your own key. TraceRoot offers a hosted cloud and a self-hosting guide, is written in TypeScript, and its license is listed as Other.

Key features

  • OpenTelemetry-compatible agent tracing
  • Detectors that flag failures in traces
  • Versioned datasets and evaluations
  • Dashboards and threshold alerts
  • CLI for exporting traces and findings
  • In-app AI assistant with code context

Arize AI alternatives: questions

What is the best open-source alternative to Arize AI?
Langfuse is the top-ranked open-source alternative to Arize AI on Enlisted: An open-source LLM engineering platform for tracing, evaluating and improving AI applications, with prompt management, datasets and a playground, self-hosted or cloud. Other strong options are MLflow, Opik, Arize Phoenix and LangWatch.
Are these Arize AI alternatives free?
All 7 are open source, so the code is free to use under its licence, and all of them can be self-hosted on your own server or computer. 7 also offer a paid or managed cloud version if you'd rather not host it yourself.
How is this list of Arize AI alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 7 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all