12 companies tracked in this category. Names and one-line descriptions are free to browse — pricing model, funding, founding year, headquarters, and cited sources are in the full dataset.
| Company | Description |
|---|---|
| Arize AI | Arize AI (Arize AX) is a unified AI observability and evaluation platform that lets teams trace, evaluate, and troubleshoot LLM applications and AI agents in development and production, alongside its open-source Phoenix project. |
| Braintrust | Braintrust provides an end-to-end evaluation, tracing, and observability platform for AI applications, letting teams inspect traces, run experiments, score outputs, and catch quality regressions in CI/CD. |
| Fiddler AI | Fiddler AI provides an AI control-plane platform delivering observability, evaluation, real-time guardrails, and governance for machine learning models, LLMs, and AI agents at enterprise scale. |
| Galileo | Galileo is an AI evaluation and observability platform that helps teams build datasets, run auto-tuned evaluations, and monitor reliability of AI applications in development and production; the company was acquired by Cisco in April 2026. |
| Helicone | Helicone is an open-source LLM observability and monitoring platform that helps developers track requests, costs, latency, and performance across multiple LLM providers. |
| Humanloop | Humanloop was an LLM evaluation, prompt management, and observability platform for enterprises; in mid-2025 its founders and much of its team joined Anthropic, and the platform's future was folded into Anthropic's roadmap. |
| Langfuse | Langfuse is an open-source LLM observability, evaluation, and prompt-management platform offering tracing, session tracking, and cost analytics for LLM applications, self-hostable or cloud-hosted; it was acquired by ClickHouse in January 2026. |
| LangSmith (LangChain) | LangSmith, built by LangChain, is a platform for debugging, testing, evaluating, and monitoring LLM applications and AI agents with tracing, evals, and deployment tooling. |
| Patronus AI | Patronus AI provides an LLM and AI agent evaluation platform, including automated evaluators, an agentic-trace failure-detection copilot (Percival), and digital-world simulation environments for stress-testing agents before production deployment. |
| Vals AI | Vals AI is an independent benchmarking platform that evaluates and ranks leading LLMs on real-world, industry-specific tasks (finance, law, healthcare, coding, and more) via its Vals Index and domain leaderboards. |
| Weights & Biases (Weave) | Weights & Biases offers Weave, an LLM observability and evaluation product (alongside its core MLOps/experiment-tracking platform) that traces, evaluates, and monitors generative AI applications in development and production; the company was acquired by CoreWeave in May 2025. |
| WhyLabs | WhyLabs built an AI/ML observability platform (and open-source tools whylogs and langkit) for monitoring data quality and LLM behavior in production; the company discontinued independent operations in 2025 after its founding team was acquired by Apple, and it open-sourced its full platform. |