THUDM/AgentBench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
[NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judge
Appears on
Quick read
Latest capture 2026-07-31 03:06
7 paths
Agent instructions and tool configuration found in this repository.
6 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +2
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
An agent benchmark with tasks in a simulated software company.
Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
LiveBench: A Challenging, Contamination-Free LLM Benchmark
Collection of evals for Inspect AI