Guide · 2026
Best Evaluation & Testing (2026)
The 10 most complete evaluation & testing listings on Prompt Junky, ranked by how much verified public information each publishes: description, services, pricing, and a live website. 97 are listed in this category. Order is not a quality rating, and no listing pays for its rank.
- 1
AdalFlow
A library for building LLM applications that also auto-tunes them, optimizing prompt text and few-shot demonstrations against a training set. It aims to make prompt tuning as routine as model training.
- Listed under Agent & Prompt Frameworks, Evaluation & Testing
- 4 services listed, including auto-optimized prompts, few-shot tuning, retriever and agent components
- Pricing published: Open source
- Serves the whole US
- 2
DeepEval
A pytest-style evaluation framework for LLM output with ready-made metrics for hallucination, answer relevancy, faithfulness and task completion. Tests run locally and can be pushed to the Confident AI platform.
- 4 services listed, including pytest-style eval suites, prebuilt metrics, rag and agent metrics
- Pricing published: Open source
- Serves the whole US
- 3
DSPy
A Python framework that replaces hand-written prompt strings with declarative modules and signatures, then compiles them against your data. Optimizers search for the wording and few-shot examples that score best on a metric you define.
- Listed under Agent & Prompt Frameworks, Evaluation & Testing
- 4 services listed, including declarative prompt modules, automatic prompt optimization, few-shot example search
- Pricing published: Open source
- Serves the whole US
- 4
garak
A vulnerability scanner for language models that probes a target with hundreds of attack prompts covering jailbreaks, prompt injection, data leakage and toxic output, then reports which probes succeeded.
- 4 services listed, including attack probe library, jailbreak detection, automated reporting
- Pricing published: Open source
- Serves the whole US
- 5
GEPA
An optimizer that improves prompts and other text artifacts through reflective evolution, reading execution traces and natural-language feedback to propose edits. It integrates with DSPy and other pipelines.
- Listed under Agent & Prompt Frameworks, Evaluation & Testing
- 4 services listed, including reflective prompt evolution, trace-driven feedback, pareto candidate selection
- Pricing published: Open source
- Serves the whole US
- 6
Haystack
deepset's orchestration framework for retrieval and agent pipelines built from composable components. Prompt builders are pipeline nodes, so prompt text is versioned with the rest of the pipeline definition.
- Listed under Agent & Prompt Frameworks, Evaluation & Testing
- 4 services listed, including composable pipelines, prompt builder components, rag orchestration
- Pricing published: Open source
- Serves the whole US
- 7
Inspect AI
An evaluation framework from the UK AI Security Institute with built-in components for prompt engineering, tool use, multi-turn dialogue and model-graded scoring, plus a log viewer for inspecting runs.
- 4 services listed, including eval task api, model-graded scoring, agent and tool evals
- Pricing published: Open source
- Serves the whole US
- 8
PromptBench
A Microsoft research framework for evaluating language models with a focus on prompt robustness, including adversarial prompt attacks and prompt engineering method comparisons. The repository is archived and kept for reference.
- Listed under Evaluation & Testing, Courses & Guides
- 4 services listed, including prompt robustness testing, adversarial prompt attacks, method comparison
- Pricing published: Open source
- Serves the whole US
- 9
Promptfoo
A command-line and library test runner for prompts, models and RAG pipelines. Test cases are declared in YAML, run side by side across providers, and the project also performs red-team scans for prompt injection and jailbreaks.
- Listed under Evaluation & Testing, Prompt Managers & Versioning
- 4 services listed, including side-by-side prompt testing, yaml test cases, red teaming scans
- Pricing published: Open source
- Serves the whole US
- 10
PromptLayer
A prompt management platform where non-engineers can edit prompts in a visual registry while engineers pull versions at runtime. It also logs requests and runs evaluation pipelines against prompt versions.
- Listed under Prompt Managers & Versioning, Evaluation & Testing
- 4 services listed, including visual prompt registry, version history, request logging
- Pricing published: Free plan; Pro $49/month
- Serves the whole US