LightEval
Hugging Face's evaluation toolkit for running benchmark tasks across several inference backends, with custom task and metric definitions and detailed per-sample output.
Services
- Multi-backend evaluation
- Custom tasks and metrics
- Per-sample logging
- Hub integration
Highlights
Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.