EvolvingLMMs-Lab/lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Repository profile
Evaluating LLMs' abilities to generate structural output [TMLR2025]
Repository updates
Get generated TIGER-AI-Lab/StructEval development summaries by email, or follow the weekly and monthly RSS feeds.
Sign in to subscribe by email. RSS feeds are public.
Sign in to subscribeTracked growth, recent movement, and commit velocity from stored repository snapshots.
Latest capture 2026-07-28 10:52
1 capture since 2026-07-28
Stars from baseline 0
All tracked data
Frameworks, package managers, ecosystems, and dependency manifests found during catalog scans.
Scanned 2026-07-28 10:52
requirements.txt
python ecosystem,
8 dependencies
setup.py
python ecosystem,
13 dependencies
Searchable topics, generated tags, and stack labels that explain where this repository fits.
Agent instructions and tool configuration paths found in the repository tree.
Nearest indexed repositories by embedding similarity.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
LiveBench: A Challenging, Contamination-Free LLM Benchmark
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI