LiveBench/LiveBench
LiveBench: A Challenging, Contamination-Free LLM Benchmark
Humanity's Last Exam
Appears on
Quick read
Latest capture 2026-07-31 03:06
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +91
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
LiveBench: A Challenging, Contamination-Free LLM Benchmark
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
No description.
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends