github Actively maintained

TIGER-AI-Lab/ClawBench

Open-source benchmark for browser AI agents on daily tasks.

1 awesome list

Quick read

Stars
538
Forks
30
Open issues
38
Commits
397

Activity and growth

Latest capture 2026-07-31 03:03

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+191
Stored snapshots
5

Classification

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2026-04-10
First commit
2026-04-10
Last pushed
2026-07-31
GitHub updated
2026-07-31
Last synced
2026-07-31 03:03
Stack scanned
2026-07-31 03:03
Archived
No

AI development signals

1 path

Agent instructions and tool configuration found in this repository.

Agent instructions

View paths

Growth history

Tracked growth

5 observed captures since 2026-05-30. Observed captures are shown by default.

Stars from first capture +191

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

InternLM/WildClawBench

An in-the-wild benchmark for AI agents in the OpenClaw Environment.

498 stars
Python 1 awesome list

reacher-z/HarnessBench

Benchmark for comparing agent harnesses on everyday online tasks — fixes the base model, varies the harness. Sister project of ClawBench, same scoring pipeline.

46 stars
Python 1 awesome list

pinchbench/skill

PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

1,301 stars
Python 1 awesome list

claw-eval/claw-eval

Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

739 stars
Python 2 awesome lists

vivekchand/clawmetry

See your agent think. Real-time observability for 14 AI agent runtimes - OpenClaw, NVIDIA NemoClaw, Claude Code, Codex & 8 more.

393 stars
Python 2 awesome lists