Best LLM Evaluation Tools in 2026
Updated
In short: Promptfoo is ranked #1 of 29 as of 3 October 2026, ahead of DeepEval and Giskard. The best-ranked option with a free plan is DeepEval. The lowest first paid tier on this page is Maxim AI at $29/mo.
When assessing prompts and model outputs, LLM evaluation tools offer different approaches and workflow connections. Compare custom metrics, safety evaluations, and LLM-as-a-judge with human review workflows and prompt versioning. CI/CD integration and deployment options can help you assess fit with your development process; free-plan availability and paid-from pricing are also worth checking. The ranking opens with DeepEval, Galileo, and Maxim AI, with Braintrust and Confident AI among the early entries. Consider which evaluation methods your team needs and how you want review, testing, and prompt changes to fit into your workflow.
29 LLM evaluation tools ranked on what their makers publish — plans and prices, free tiers, platforms and the facts on their own pages.
#1 Promptfoo Top pick · 7.3 Free plan · Free
#2 DeepEval Runner-up · 7.2 Free plan · Free
#3 Giskard Also great · 7.1 Free plan · Free- Free plan apiLinuxmacOSself-hostedWebWindows
- Free plan
- Yes
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - Free plan LinuxmacOSself-hostedWindows
- Free plan
- Yes
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - Free plan apiLinuxself-hostedWeb
- Free plan
- Yes
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - Free plan apiself-hostedWeb
- Free plan
- Yes
- Paid from
- 249 /mo
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms$249/mofirst paid tier About BraintrustVisit site - Free plan apiself-hostedWeb
- Free plan
- Yes
- Paid from
- 200 /mo
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms$200/mofirst paid tier About Confident AIVisit site - Free plan apiself-hostedWeb
- Free plan
- Yes
- Paid from
- 100 /mo
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms$100/mofirst paid tier About GalileoVisit site - Free planFree trial apiself-hostedWeb
- Free plan
- Yes
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms$29/mofirst paid tier About Maxim AIVisit site - Free plan apiself-hostedWeb
- Free plan
- Yes
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms$150/mofirst paid tier About Parea AIVisit site - WindowsmacOSLinux
- Free plan
- Yes
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - WindowsmacOSLinux
- Free plan
- Yes
- Deployment options
- self-hosted
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - Web
- Deployment options
- self-hosted
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - WebmacOSLinux
- Free plan
- Yes
- Deployment options
- both
RecognisedDocumentedFree planPlatforms - Web
- Free plan
- Yes
- Deployment options
- cloud
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms -
- Deployment options
- self-hosted
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms -
- Free plan
- Yes
- Deployment options
- self-hosted
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms -
- Deployment options
- self-hosted
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
RecognisedDocumentedFree planPlatforms - Web
- Deployment options
- both
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - Linux
- Free plan
- Yes
- Deployment options
- self-hosted
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms -
- Free plan
- Yes
- Deployment options
- self-hosted
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
RecognisedDocumentedFree planPlatforms -
- Deployment options
- self-hosted
- Custom metrics
- Yes
- LLM-as-a-judge
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms - Linux
- Free plan
- Yes
- Deployment options
- self-hosted
RecognisedDocumentedFree planPlatforms - Linux
- Deployment options
- self-hosted
- LLM-as-a-judge
- Yes
RecognisedDocumentedFree planPlatforms - WebRecognisedDocumentedFree planPlatforms
- Web
- Deployment options
- both
- LLM-as-a-judge
- No
RecognisedDocumentedFree planPlatforms -
- Deployment options
- self-hosted
- Custom metrics
- Yes
- Safety evaluations
- Yes
RecognisedDocumentedFree planPlatforms
Is your app on this list?
Numbered spots on this list can be sponsored. They are labelled, and the editorial order and scores never change for payment.
Questions about this list
Which LLM evaluation tool is ranked first on Everything Xiaomi?
Promptfoo is ranked #1 of 29 with a score of 7.3. DeepEval is second and Giskard third.
How many of these have a free plan?
8 of the 25 on this page publish a free plan on their own pricing pages.
Which is the cheapest paid option?
On this page, Maxim AI has the lowest first paid tier we found: $29/mo.
How is this list ranked?
Ranked on what each maker publishes: documentation depth, a free tier and the platforms it runs on. Paid placements never change a rank.

































