
Langfuse
Open source observability and evaluation for LLM applications
Overview
Langfuse is an open source platform for LLM engineering teams who need to see what's actually happening inside their AI agents once they leave the demo stage. It traces every call an application makes to a model, tracks prompt versions separately from code, and lets teams run evaluations using datasets, human review, or another LLM as a judge. It plugs into most major frameworks and languages through SDKs and OpenTelemetry, so adding it to an existing stack doesn't mean a rewrite.
It's aimed at developers and ML teams building production LLM products who need to debug why an agent gave a bad answer, track cost per user, or prove a prompt change actually improved output quality. The free Hobby tier covers small projects with 50,000 observations a month, and because the core product is open source, teams that want full control can self-host instead of paying at all.
Key Features
Tracing: Captures every LLM call, tool invocation, and retrieval step with session and user tracking.
Prompt Management: Version prompts separately from code, with one-click deploys and rollbacks.
Evaluation: Datasets, experiments, LLM-as-judge scoring, and human annotation queues.
Cost and Latency Monitoring: Filter traces by cost, latency, or custom metadata to find what's slow or expensive.
Framework Integrations: Native SDKs for Python, JavaScript, Java, and Go, plus OpenTelemetry compatibility.
Self-Hosting: Fully open source, so teams can run it on their own infrastructure for data control.
Pricing
Starting price
$29/mo (Core plan; a free Hobby tier also exists with 50k units/month)
Hobby: Free. 50k units/month, 30-day data access, 2 users.
Core: $29/month. 100k units/month, 90-day data access, unlimited users.
Pro: $199/month. 100k units/month, 3-year data access, advanced features.
Enterprise: $2,499/month. Custom limits, SLA, dedicated support.
Overage: $8 per 100k additional units, tiered down at higher volume.
Disclaimer: pricing may change, confirm on Langfuse's own pricing page before buying.
Pros
Open Source Core: Self-host for free with no vendor lock-in.
Deep Tracing: Full visibility into multi-step agent chains, not just single prompt-response pairs.
Framework Agnostic: Works with LangChain, OpenAI SDK, LiteLLM, and plain OpenTelemetry.
Built-In Evals: LLM-as-judge and human annotation in the same tool as tracing, no separate platform needed.
Cons
Developer-Only Tool: Requires SDK integration and code changes, not something a non-technical marketer can set up alone.
Pricing Jumps Fast: The gap from Core ($29) to Pro ($199) is steep once you need longer data retention.
Usage Overages Add Up: Extra observations beyond plan limits are billed per 100k units, which can surprise high-volume teams.
What Makes It Unique
Open Source With No Compromise: Unlike most LLM observability tools that are cloud-only, Langfuse's full feature set is available self-hosted for free, which matters a lot for teams with data residency rules.
Kay Score
7.8
/ 10
Tool Information
Pricing
$29/mo (Core plan; a free Hobby tier also exists with 50k units/month)
Category
Coding & Development
Platform
Web / iOS / Android
Last Updated
