TheAgentCompany/TheAgentCompany
An agent benchmark with tasks in a simulated software company.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Quick read
Latest capture 2026-08-03 03:12
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
8 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +189
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
An agent benchmark with tasks in a simulated software company.
A generalized information-seeking agent system with Large Language Models (LLMs).
[ICLR 2026] LLM/VLM gaming agents and model evaluation through games.
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
A benchmark for LLMs on complicated tasks in the terminal
[NeurIPS'25 D&B] Mind2Web-2 Benchmark: Evaluating Agentic Search with Agent-as-a-Judge