github Actively maintained

THUDM/AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

Quick read

Stars
3,634
Forks
272
Open issues
74
Commits
75

Activity and growth

Latest capture 2026-08-03 03:12

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+189
Stored snapshots
8

Classification

Technology stack

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2023-07-28
First commit
2023-10-18
Last pushed
2026-02-08
GitHub updated
2026-08-03
Last synced
2026-08-03 03:12
Stack scanned
2026-08-03 03:12
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

8 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +189

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

KwaiKEG/KwaiAgents

A generalized information-seeking agent system with Large Language Models (LLMs).

1,202 stars
Python 1 awesome list

lmgame-org/GamingAgent

[ICLR 2026] LLM/VLM gaming agents and model evaluation through games.

963 stars
Python 0 awesome lists

OSU-NLP-Group/Mind2Web-2

[NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judge

113 stars
Python 1 awesome list