EvolvingLMMs-Lab/lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
Appears on
Quick read
Latest capture 2026-08-03 03:12
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +74
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
No description.