Giskard vs Pydantic Evals
| Giskard | Pydantic Evals | |
|---|---|---|
| Free plan | Yes | |
| Paid from | Free | |
| Platforms | api, Linux, self-hosted, Web | Linux |
| Free plan | Yes | Yes |
| Deployment options | both | |
| Custom metrics | Yes | |
| LLM-as-a-judge | Yes | |
| Safety evaluations | Yes | Yes |
| Human review workflows | Yes | |
| CI/CD integration | Yes | |
| Evaluation methods | Deterministic checks; custom evaluators; LLM judges; G-Eval; performance checks; report evaluators; span-based evaluation; agentic trajectory evaluation | |
| Model support | OpenAI; Anthropic; Gemini; xAI; Bedrock; Cerebras; Cohere; Groq; Hugging Face; Mistral; OpenRouter; and other listed Pydantic AI providers | |
| Deployment | self-hosted | |
| Prompt versioning | Yes | |
| API access | Yes |
Listed together in Best AI LLM Evaluation Tools