Langfuse
An open-source LLM engineering platform for tracing, evaluating and improving AI applications, with prompt management, datasets and a playground, self-hosted or cloud.
- GitHub stars
- 35k
- Last commit
- yesterday
- Latest release
- v4.50.0
- Self-hosted
- Yes
- Hosted version
- Available

Langfuse is an open-source LLM engineering platform that helps teams develop, monitor, evaluate and debug AI applications together. You instrument your app, and Langfuse records traces of LLM calls and surrounding logic such as retrieval, embeddings and agent actions, so complex runs and user sessions can be inspected and debugged.
Prompt management lets you centrally version and iterate on prompts, with caching on server and client so that changes do not add latency to your app. Evaluation features cover LLM-as-a-judge, code evaluators, user feedback, manual labeling and custom pipelines through the API and SDKs. Datasets provide test sets and benchmarks, there is an LLM playground for trying prompts, and integrations exist for OpenAI, LangChain and LlamaIndex. It is built on the ClickHouse database, and the Langfuse team has been part of ClickHouse since January 2026.
You can use Langfuse Cloud or self-host it, which the maintainers say takes minutes. The repository license is listed as 'Other' because it combines open-source and enterprise components, so check the terms. It suits teams shipping production LLM features who need visibility into quality, latency and cost.
Key features
- Tracing of LLM calls, retrieval and agent actions
- Prompt versioning with caching
- LLM-as-a-judge and custom evaluations
- Datasets for tests and benchmarks
- Interactive LLM playground
- Integrations with OpenAI, LangChain and LlamaIndex
Pricing: Core has a $29 monthly base that includes 100k units, then $8 per 100k units with volume discounts. Hobby and Pro tiers exist but their prices are not shown in the text; Enterprise is by sales.




