harbor-framework/terminal-bench
Measuring and evolving with the frontier of agent work
A benchmark for LLMs on complicated tasks in the terminal
Appears on
Quick read
Latest capture 2026-09-06 10:53
5 paths
Agent instructions and tool configuration found in this repository.
Claude Code
Gemini CLI 3
GitHub Copilot
1 observed capture since 2026-09-06. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Measuring and evolving with the frontier of agent work
No description.
An agent benchmark with tasks in a simulated software company.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models