Prompt Junky

Stanford HELM

Online · Nationwide

A holistic evaluation framework from Stanford CRFM that scores models across many scenarios and metrics including accuracy, calibration, robustness and bias, and publishes public leaderboards.

Services

  • Multi-metric evaluation
  • Scenario library
  • Public leaderboards
  • Reproducible runs

Highlights

  • #python
  • #benchmarks
  • #stanford
  • #research
  • #apache-2.0

Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.

Nearby and similar

More prompt tools

All prompt tools