DeepEval
A pytest-style evaluation framework for LLM output with ready-made metrics for hallucination, answer relevancy, faithfulness and task completion. Tests run locally and can be pushed to the Confident AI platform.
Services
- Pytest-style eval suites
- Prebuilt metrics
- RAG and agent metrics
- Red teaming
Highlights
Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.