← All categories

AI evaluation & observability

12 companies tracked in this category. Names and one-line descriptions are free to browse — pricing model, funding, founding year, headquarters, and cited sources are in the full dataset.

CompanyDescription
Arize AIArize AI (Arize AX) is a unified AI observability and evaluation platform that lets teams trace, evaluate, and troubleshoot LLM applications and AI agents in development and production, alongside its open-source Phoenix project.
BraintrustBraintrust provides an end-to-end evaluation, tracing, and observability platform for AI applications, letting teams inspect traces, run experiments, score outputs, and catch quality regressions in CI/CD.
Fiddler AIFiddler AI provides an AI control-plane platform delivering observability, evaluation, real-time guardrails, and governance for machine learning models, LLMs, and AI agents at enterprise scale.
GalileoGalileo is an AI evaluation and observability platform that helps teams build datasets, run auto-tuned evaluations, and monitor reliability of AI applications in development and production; the company was acquired by Cisco in April 2026.
HeliconeHelicone is an open-source LLM observability and monitoring platform that helps developers track requests, costs, latency, and performance across multiple LLM providers.
HumanloopHumanloop was an LLM evaluation, prompt management, and observability platform for enterprises; in mid-2025 its founders and much of its team joined Anthropic, and the platform's future was folded into Anthropic's roadmap.
LangfuseLangfuse is an open-source LLM observability, evaluation, and prompt-management platform offering tracing, session tracking, and cost analytics for LLM applications, self-hostable or cloud-hosted; it was acquired by ClickHouse in January 2026.
LangSmith (LangChain)LangSmith, built by LangChain, is a platform for debugging, testing, evaluating, and monitoring LLM applications and AI agents with tracing, evals, and deployment tooling.
Patronus AIPatronus AI provides an LLM and AI agent evaluation platform, including automated evaluators, an agentic-trace failure-detection copilot (Percival), and digital-world simulation environments for stress-testing agents before production deployment.
Vals AIVals AI is an independent benchmarking platform that evaluates and ranks leading LLMs on real-world, industry-specific tasks (finance, law, healthcare, coding, and more) via its Vals Index and domain leaderboards.
Weights & Biases (Weave)Weights & Biases offers Weave, an LLM observability and evaluation product (alongside its core MLOps/experiment-tracking platform) that traces, evaluates, and monitors generative AI applications in development and production; the company was acquired by CoreWeave in May 2025.
WhyLabsWhyLabs built an AI/ML observability platform (and open-source tools whylogs and langkit) for monitoring data quality and LLM behavior in production; the company discontinued independent operations in 2025 after its founding team was acquired by Apple, and it open-sourced its full platform.
Get the full dataset