LMMs-Eval
An evaluation toolkit for multimodal models covering text, image, video and audio tasks with a unified interface, used to benchmark vision-language prompting.
Services
- Multimodal benchmarks
- Unified task interface
- Video and audio evals
- Reproducible configs
Highlights
Source: Official GitHub repository. Last checked 2026-09-18. Spot an error? Tell us.