- LangSmith Freemium ★4.86
AI agent observability platform — tracing, monitoring, and evals for any agent stack.
LLM Observability & Evals
- Okareo Free Trial ★4.84
Simulation testing for voice and text AI agents: synthetic users, audio-condition evals, CI/CD gating.
LLM Observability & Evals
- AgentOps Freemium ★4.84
Agent observability platform for OpenAI, CrewAI, Autogen, and 400+ LLMs. Visually track LLM calls, tools, multi-agent flows. Rewind and replay runs.
LLM Observability & Evals
- Langfuse Freemium ★4.84
Open-source LLM engineering platform — tracing, evals, prompt management, and metrics.
LLM Observability & Evals
- Coval Free Trial ★4.84
Simulation and observability platform that tests, monitors, and evaluates AI voice and chat agents.
LLM Observability & Evals
- Judgment Labs Paid ★4.83
Judgment Labs is a continuous-improvement stack for AI agents — monitoring, failure analysis, and pre-deploy testing.
LLM Observability & Evals
- Arthur Paid ★4.82
ML Observability platform ensuring transparent, compliant, and efficient AI operations.
LLM Observability & Evals
- Phoenix by Arize Free ★4.82
Open-source LLM and agent observability with tracing, evaluations, and experiment tracking.
LLM Observability & Evals
- Weights & Biases Paid ★4.82
Central dashboard for tracking hyperparameters, metrics, and ML workflows.
LLM Observability & Evals
- Braintrust Freemium ★4.81
AI evals and observability — turn production traces into evals and ship quality AI at scale.
LLM Observability & Evals
- Arize AI Paid ★4.80
ML observability platform for monitoring and fine-tuning machine learning models.
LLM Observability & Evals
- Patronus AI Paid ★4.78
Automated LLM and agent evaluation platform — detect hallucinations, bias, and performance regressions.
LLM Observability & Evals
- Helicone Paid ★4.75
Helicone: Open-source monitoring for generative AI applications.
LLM Observability & Evals
- Klu Freemium ★4.75
Generative AI app platform — design, deploy, and optimize LLM-enabled applications with collaborative prompt management, experiments, and evals.
LLM Observability & Evals
- Galileo Freemium ★4.75
Galileo is an LLM evaluation and observability platform that tests, monitors, and guardrails GenAI applications and agents at enterprise scale.
LLM Observability & Evals
- Fiddler AI Paid ★4.75
Fiddler AI: AI Observability platform for ML model monitoring and explainability.
LLM Observability & Evals
- Distributional Paid ★4.73
Distributional is an enterprise AI testing platform that statistically detects drift and regressions in agents and LLM applications before they hit product
LLM Observability & Evals
- HoneyHive Paid ★4.73
Optimization platform for GPT-4 applications, enhancing production LLM apps with observability, evaluation, and fine-tuning tools.
LLM Observability & Evals
- Athina AI Freemium ★4.72
Full-stack LLMOps platform — prototype pipelines, run evals, detect hallucinations and safety issues in LLM products. 50+ preset evals, YC-backed.
LLM Observability & Evals
- Deepchecks Paid ★4.64
Continuous ML validation platform for testing, CI/CD, and monitoring.
LLM Observability & Evals
- Kolena Paid ★4.62
AI quality platform for end-to-end testing, data curation, and model evaluation.
LLM Observability & Evals
- Freeplay Paid ★4.61
Freeplay is an LLM evaluation and observability platform that helps cross-functional teams test, monitor, and improve AI-powered products.
LLM Observability & Evals
- TruEra Paid ★4.59
ML monitoring, testing, and quality management solutions for AI.
LLM Observability & Evals
- Pay-i Paid ★4.49
Pay-i is the AI cost observability and governance platform tracking spend across OpenAI, Anthropic, Google, and self-hosted models. Khosla Ventures-backed.
LLM Observability & Evals
- Openlayer Freemium ★4.48
Openlayer is the LLM and ML observability platform for testing, monitoring, and improving AI models in production. YC alum; ~$5M seed.
LLM Observability & Evals