THUDM/AgentBench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
[ICLR 2026] LLM/VLM gaming agents and model evaluation through games.
Quick read
Latest capture 2026-08-11 04:40
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
43 observed captures since 2026-06-02. Observed captures are shown by default.
Stars from first capture +31
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
A generalized information-seeking agent system with Large Language Models (LLMs).
An agent benchmark with tasks in a simulated software company.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
A lightweight framework for building LLM-based agents
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering