Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
The best open source alternative to Arize Phoenix is Langfuse. If that doesn't suit you, we've compiled a ranked list of other open source Arize Phoenix alternatives to help you find a suitable replacement. Other interesting open source alternatives to Arize Phoenix are: Helicone, Latitude, OpenLIT , and Spanlens.
Arize Phoenix alternatives are mainly LLM Observability & Evaluation Tools. Browse these if you want a narrower list of alternatives or looking for a specific functionality of Arize Phoenix.
Langfuse provides tracing, evaluations, prompt management, and analytics to debug and improve LLM applications.

Langfuse is an open source LLM engineering platform designed to help teams build, debug, and improve AI-powered applications. With its comprehensive suite of tools, Langfuse empowers developers to gain deep insights into their LLM applications and optimize performance.
Key features of Langfuse include:
Tracing: Capture detailed production traces to quickly identify and resolve issues in your LLM applications. Visualize the entire request flow and pinpoint bottlenecks.
Evaluations: Collect user feedback, annotate data, and run custom evaluation functions to assess the quality and performance of your AI models.
Prompt Management: Collaboratively version and deploy prompts, with low-latency retrieval for production use. Streamline your prompt engineering workflow.
Analytics: Track key metrics like cost, latency, and quality to optimize your LLM application's performance and efficiency.
Playground: Test different prompts and models directly within the Langfuse UI, enabling rapid experimentation and iteration.
Datasets: Derive high-quality datasets from production data to fine-tune models and thoroughly test your LLM applications.
Langfuse integrates seamlessly with popular LLM frameworks and libraries, including LangChain, LlamaIndex, and OpenAI. It offers SDKs for Python and JavaScript/TypeScript, making it easy to incorporate into your existing workflow.
Built for teams of all sizes, Langfuse can be self-hosted or used as a cloud service. It's designed with enterprise-grade security in mind, offering SOC 2 Type II and ISO 27001 certifications for the cloud version.
By providing a comprehensive toolkit for LLM engineering, Langfuse helps teams build more reliable, efficient, and high-quality AI applications. Whether you're just starting with LLMs or scaling a complex AI system, Langfuse offers the observability and tools needed to succeed in the rapidly evolving field of AI engineering.
Open-source platform for logging, monitoring, and debugging LLM applications. Route, debug, and analyze AI apps with comprehensive observability tools.
Helicone is the open-source platform that helps developers build reliable AI applications through comprehensive observability. Trusted by the world's fastest-growing AI companies, it provides essential tools for routing, debugging, and analyzing LLM applications.
Key Features:
The platform offers a comprehensive dashboard for monitoring AI application performance, with detailed request tracking and user analytics. Developers can experiment with prompts, run evaluations, and manage datasets all within one unified interface.
Getting Started: No credit card required with a 7-day free trial. The platform is designed to help developers quickly identify issues, optimize performance, and ensure their AI applications run reliably at scale.
Open-source platform for monitoring AI agents: captures traces, surfaces failure patterns, alerts on issues, and helps you verify fixes with automated evals.

Latitude is an open-source monitoring platform built specifically for AI agents. It captures everything happening in production, including messages, tool calls, costs, and errors, then helps you understand what's actually going wrong and why. It's aimed at teams building AI agent platforms who need more than raw logs to debug production behavior.
The core idea is full-coverage observability. Latitude runs semantic search across 100% of your traces, no sampling, so you never miss a cohort of failing users. Combine that with exact text search and metadata filters to go from a broad hunch to a focused set of real examples fast.
Key capabilities:
Latitude is OpenTelemetry compatible, so you can point an existing OTEL pipeline at it without adopting a proprietary format. It also exposes an MCP server so coding agents can manage projects, traces, annotations, and datasets without touching the UI. Tools like Helicone and Arize Phoenix cover similar ground, but Latitude's automatic issue discovery and eval generation from production failures is a distinct angle.
It's SOC 2 Type II certified, GDPR compliant, and supports SSO with SAML 2.0, end-to-end encryption, data residency options, and audit logs.
Open-source observability platform for GenAI and LLM applications. Real-time monitoring, distributed tracing, prompt management, and AI model evaluation built on OpenTelemetry.

Monitor and optimize your LLM applications with comprehensive observability tools designed for production AI workloads. Built entirely on OpenTelemetry standards for seamless integration with existing infrastructure.
Key capabilities include:
Quick setup requires just a few lines of code with zero application changes. The platform supports automatic Kubernetes instrumentation through the OpenLIT Operator, making it perfect for containerized environments.
Privacy-first approach ensures your data never leaves your infrastructure, while the open-source nature eliminates vendor lock-in concerns. Compatible with all major LLM providers and frameworks including OpenAI, Anthropic, Google, AWS Bedrock, and popular vector databases.
Production-ready with minimal performance overhead, designed to scale with your AI applications from development to enterprise deployment.
Drop-in observability platform for OpenAI, Anthropic, and Gemini that logs every request, tracks costs, traces agent workflows, and flags anomalies and PII.

Spanlens is an MIT-licensed LLM observability platform that gives you full visibility into every request your app makes to OpenAI, Anthropic, or Gemini. It works as a drop-in replacement for the provider SDK, so you swap one import and start seeing data immediately. No agents to run, no infrastructure to wire up.
It's built for teams shipping LLM-powered products who need to answer real questions fast: why did the bill spike, which agent step is slow, did the new prompt actually improve quality?
Core capabilities:
Spanlens supports OpenAI, Anthropic, Google Gemini, Mistral, Azure, Bedrock, and Vertex, plus framework integrations for LangChain, LlamaIndex, Vercel AI SDK, and LangGraph. It also ingests OpenTelemetry spans over OTLP/HTTP.
For teams that can't send prompt data to a third party, it's fully self-hostable. Prompts and completions stay inside your own network. Data can be exported as JSON, CSV, or Parquet, or streamed to S3 or BigQuery.
If you're evaluating LangSmith alternatives, Spanlens covers similar ground with a flat monthly pricing model rather than per-seat fees, and a free self-hosted option that has no usage cap.
Managed Open Source software hosting in the EU: secure, compliant, fast.
Start using Open Source today