github Actively maintained

OpenHands/benchmarks

Evaluation harness for OpenHands V1.

1 awesome list

Quick read

Stars
111
Forks
84
Open issues
65
Commits
432

Activity and growth

Latest capture 2026-08-02 03:13

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+26
Stored snapshots
6

Classification

Technology stack

Metadata

Language
Python
License
MIT
Default branch
main
Created
2025-09-02
First commit
2025-09-02
Last pushed
2026-07-31
GitHub updated
2026-07-31
Last synced
2026-08-02 03:13
Stack scanned
2026-08-02 03:13
Archived
No

AI development signals

1 path

Agent instructions and tool configuration found in this repository.

Agent instructions

View paths

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +26

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

scaleapi/SWE-bench_Pro-os

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

491 stars
Python 1 awesome list

InternLM/WildClawBench

An in-the-wild benchmark for AI agents in the OpenClaw Environment.

498 stars
Python 1 awesome list

modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

3,177 stars
Python 3 awesome lists

SWE-bench/SWE-bench

SWE-bench: Can Language Models Resolve Real-world Github Issues?

5,549 stars
Python 2 awesome lists