Best AI LLM Evaluation Tools in 2026

In short: Vellum is ranked #1 of 30 as of 3 October 2026, ahead of Weights & Biases and Evidently AI. The best-ranked option with a free plan is Weights & Biases. The lowest first paid tier on this page is Opik at $19/mo.

When you need to assess language-model outputs, prompts, or safety behavior, evaluation tools offer different methods and workflow features to compare. DeepEval, Opik, and Langfuse lead the ranked entries, followed by Braintrust and Confident AI. Look at evaluation methods, model support, and safety evaluations to see which assessment capabilities are listed. Prompt versioning can help distinguish how tools handle prompt changes, while API access and deployment give you additional points of comparison for fitting a tool into your workflow. Free-plan availability and paid-from pricing are also covered. Consider what you need to evaluate and which listed features matter most for your use.

30 AI LLM evaluation tools ranked on what their makers publish — plans and prices, free tiers, platforms and the facts on their own pages.

30ranked
14free plans on this page
$19/molowest paid tier
3 Oct 2026last checked
#1 Vellum Top pick · 7.4 Free plan · $30/mo #2 Weights & Biases Runner-up · 7.4 Free plan · $60/mo #3 Evidently AI Also great · 7.3 Free plan · $80/mo
  1. 1 7.4
    Free plan AndroidextensioniOSmacOSself-hostedWebWindows
    Free plan
    Yes
    Paid from
    30 /mo
    RecognisedDocumentedFree planPlatforms
    $30/mofirst paid tier About VellumVisit site
  2. Free planFree trial apiiOSLinuxmacOSself-hostedWebWindows
    Free plan
    Yes
    Paid from
    60 /mo
    RecognisedDocumentedFree planPlatforms
    $60/mofirst paid tier About Weights & BiasesVisit site
  3. Free planFree trial apiLinuxmacOSself-hostedWebWindows
    Free plan
    Yes
    RecognisedDocumentedFree planPlatforms
    $80/mofirst paid tier About Evidently AIVisit site
  4. 4 7.2
    Free plan apiLinuxmacOSself-hostedWebWindows
    Free plan
    Yes
    RecognisedDocumentedFree planPlatforms
  5. 5 7.1
    Free plan LinuxmacOSself-hostedWindows
    Free plan
    Yes
    Safety evaluations
    Yes
    RecognisedDocumentedFree planPlatforms
  6. 6 7.0
    Free plan apiLinuxself-hostedWeb
    Free plan
    Yes
    RecognisedDocumentedFree planPlatforms
  7. 7 7.0
    Free planFree trial apiLinuxself-hostedWeb
    Free plan
    Yes
    Paid from
    19 /mo
    RecognisedDocumentedFree planPlatforms
    $19/mofirst paid tier About OpikVisit site
  8. 8 7.0
    Free plan apiLinuxself-hostedWeb
    Free plan
    Yes
    Evaluation methods
    offline evaluation, online trace metrics, LLM-as-a-judge, custom metrics, single-turn testing, multi-turn testing, adversarial red-teaming
    Model support
    OpenAI, Anthropic, Google Gemini, Azure OpenAI, Mistral, Cohere, Groq, Together AI, Perplexity, Replicate, Ollama, vLLM, LiteLLM Proxy
    Safety evaluations
    Yes
    RecognisedDocumentedFree planPlatforms
  9. Free plan apiself-hostedWeb
    Free plan
    Yes
    Paid from
    200 /mo
    RecognisedDocumentedFree planPlatforms
    $200/mofirst paid tier About Confident AIVisit site
  10. 10 6.9
    Free plan apiself-hostedWeb
    Free plan
    Yes
    Paid from
    100 /mo
    RecognisedDocumentedFree planPlatforms
    $100/mofirst paid tier About GalileoVisit site
  11. 11 6.9
    Free plan apiself-hostedWeb
    Free plan
    Yes
    Paid from
    29 /mo
    RecognisedDocumentedFree planPlatforms
    $29/mofirst paid tier About LangfuseVisit site
  12. 12 6.9
    Free planFree trial apiself-hostedWeb
    Free plan
    Yes
    RecognisedDocumentedFree planPlatforms
    $29/mofirst paid tier About Maxim AIVisit site
  13. 13 6.8
    Free plan apiself-hostedWeb
    Free plan
    Yes
    Paid from
    249 /mo
    RecognisedDocumentedFree planPlatforms
    $249/mofirst paid tier About BraintrustVisit site
  14. Free plan apiLinuxself-hosted
    Evaluation methods
    Built-in benchmarks, exact match, fuzzy match, multiple-choice regex, answer-line, numeric match, code sandbox, LLM-as-judge, JSON-schema validation, regression comparison, quality gates
    Model support
    OpenAI-compatible model endpoints, NVIDIA API Catalog, vLLM, NIM, local vLLM, SGLang, TensorRT-LLM, vision-language models
    Safety evaluations
    Yes
    Deployment
    self-hosted
    RecognisedDocumentedFree planPlatforms
  15. 15 6.4
    apiself-hostedWeb
    Evaluation methods
    basic exact/match evaluations, model-graded evaluations, custom evaluation logic, academic benchmarks, meta-evaluations
    Model support
    OpenAI API models and custom CompletionFunction implementations
    Safety evaluations
    Yes
    Deployment
    hybrid
    RecognisedDocumentedFree planPlatforms
  16. 16 6.4
    Linux
    Free plan
    Yes
    Evaluation methods
    Deterministic checks; custom evaluators; LLM judges; G-Eval; performance checks; report evaluators; span-based evaluation; agentic trajectory evaluation
    Model support
    OpenAI; Anthropic; Gemini; xAI; Bedrock; Cerebras; Cohere; Groq; Hugging Face; Mistral; OpenRouter; and other listed Pydantic AI providers
    Safety evaluations
    Yes
    Deployment
    self-hosted
    RecognisedDocumentedFree planPlatforms
  17. 17 6.4
    apiself-hostedWeb
    Evaluation methods
    preconfigured checks; custom prompt evaluations; custom Python evaluations; model-graded evaluations; classification; chain-of-thought classification; regression testing; experiments
    Model support
    OpenAI; Azure; Claude; Mistral; Together AI; Anyscale; Ollama; Hugging Face; Replicate; custom endpoints
    Safety evaluations
    Yes
    Deployment
    hybrid
    RecognisedDocumentedFree planPlatforms
  18. 18 5.5
    Web
    Free plan
    Yes
    RecognisedDocumentedFree planPlatforms
  19. 19 5.5
    WebLinux
    Free plan
    Yes
    RecognisedDocumentedFree planPlatforms
  20. 20 5.5
    Evaluation methods
    objective; subjective; discriminative; generative; LLM-as-a-judge
    Model support
    Hugging Face models; API-based models; custom models; OpenAI; Anthropic; Gemini; Qwen; GLM; DeepSeek
    Safety evaluations
    Yes
    Deployment
    self-hosted
    RecognisedDocumentedFree planPlatforms
  21. 21 5.5
    Deployment
    self-hosted
    RecognisedDocumentedFree planPlatforms
  22. 22 5.4
    Web
    Deployment
    self-hosted
    RecognisedDocumentedFree planPlatforms
  23. 23 5.4
    Web
    Free plan
    Yes
    RecognisedDocumentedFree planPlatforms
  24. 24 5.4
    Free plan
    Yes
    Deployment
    self-hosted
    RecognisedDocumentedFree planPlatforms
  25. 25 5.2
    Linux
    Deployment
    self-hosted
    RecognisedDocumentedFree planPlatforms

Is your app on this list?

Numbered spots on this list can be sponsored. They are labelled, and the editorial order and scores never change for payment.

Questions about this list

Which AI LLM evaluation tool is ranked first on Everything Xiaomi?

Vellum is ranked #1 of 30 with a score of 7.4. Weights & Biases is second and Evidently AI third.

How many of these have a free plan?

14 of the 25 on this page publish a free plan on their own pricing pages.

Which is the cheapest paid option?

On this page, Opik has the lowest first paid tier we found: $19/mo.

How is this list ranked?

Ranked on what each maker publishes: documentation depth, a free tier and the platforms it runs on. Paid placements never change a rank.

More in AI Tools

All AI tools lists