ucsb-mlsec/terminal-bench-env
No description.
A benchmark for LLMs on complicated tasks in the terminal
Appears on
Quick read
Latest capture 2026-08-02 03:08
5 paths
Agent instructions and tool configuration found in this repository.
Claude Code
Gemini CLI 3
GitHub Copilot
7 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +245
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
No description.
An agent benchmark with tasks in a simulated software company.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models