THUDM/AgentBench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
ICML 2024: Improving Factuality and Reasoning in Language Models through Multiagent Debate
Appears on
Quick read
Latest capture 2026-08-01 03:07
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
7 observed captures since 2026-05-22. Observed captures are shown by default.
Stars from first capture +14
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
Code for Arxiv 2023: Improving Language Model Negociation with Self-Play and In-Context Learning from AI Feedback
[ICLR 2026] LLM/VLM gaming agents and model evaluation through games.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
data-to-paper: Backward-traceable AI-driven scientific research