github Actively maintained

InternLM/WildClawBench

An in-the-wild benchmark for AI agents in the OpenClaw Environment.

1 awesome list

Quick read

Stars
498
Forks
56
Open issues
9
Commits
25

Activity and growth

Latest capture 2026-08-02 03:10

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+91
Stored snapshots
6

Classification

Metadata

Language
Python
License
MIT
Default branch
main
Created
2026-03-23
First commit
2026-05-11
Last pushed
2026-07-30
GitHub updated
2026-07-30
Last synced
2026-08-02 03:10
Stack scanned
2026-08-02 03:10
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +91

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

TIGER-AI-Lab/ClawBench

Open-source benchmark for browser AI agents on daily tasks.

538 stars
Python 1 awesome list

claw-eval/claw-eval

Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

739 stars
Python 2 awesome lists

pinchbench/skill

PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

1,301 stars
Python 1 awesome list

reacher-z/HarnessBench

Benchmark for comparing agent harnesses on everyday online tasks — fixes the base model, varies the harness. Sister project of ClawBench, same scoring pipeline.

46 stars
Python 1 awesome list