github Actively maintained

harbor-framework/terminal-bench

A benchmark for LLMs on complicated tasks in the terminal

2 awesome lists

Quick read

Stars
2,509
Forks
564
Open issues
319
Commits
904

Activity and growth

Latest capture 2026-08-02 03:08

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+245
Stored snapshots
7

Classification

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2025-01-17
First commit
2025-01-17
Last pushed
2026-07-11
GitHub updated
2026-08-01
Last synced
2026-08-02 03:08
Stack scanned
2026-08-02 03:08
Archived
No

AI development signals

5 paths

Agent instructions and tool configuration found in this repository.

Growth history

Tracked growth

7 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +245

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

THUDM/AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

3,634 stars
Python 4 awesome lists

InternLM/WildClawBench

An in-the-wild benchmark for AI agents in the OpenClaw Environment.

498 stars
Python 1 awesome list

google/BIG-bench

Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models

3,249 stars
Python 2 awesome lists