sierra-research/tau-bench
Code and Data for Tau-Bench
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Appears on
Quick read
Latest capture 2026-08-02 03:14
12 paths
Agent instructions and tool configuration found in this repository.
Agent instructions 7
Cursor 5
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +482
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Code and Data for Tau-Bench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
A benchmark for LLMs on complicated tasks in the terminal
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Collection of evals for Inspect AI
An agent benchmark with tasks in a simulated software company.