Prompt Junky

DeepEval

Online · Nationwide

A pytest-style evaluation framework for LLM output with ready-made metrics for hallucination, answer relevancy, faithfulness and task completion. Tests run locally and can be pushed to the Confident AI platform.

Services

  • Pytest-style eval suites
  • Prebuilt metrics
  • RAG and agent metrics
  • Red teaming

Highlights

  • #python
  • #pytest
  • #metrics
  • #apache-2.0

Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.

Nearby and similar

More prompt tools

All prompt tools