InternLM/WildClawBench
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
Open-source benchmark for browser AI agents on daily tasks.
Appears on
Quick read
Latest capture 2026-07-31 03:03
1 path
Agent instructions and tool configuration found in this repository.
Agent instructions
5 observed captures since 2026-05-30. Observed captures are shown by default.
Stars from first capture +191
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
Benchmark for comparing agent harnesses on everyday online tasks — fixes the base model, varies the harness. Sister project of ClawBench, same scoring pipeline.
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
A benchmark for LLMs on complicated tasks in the terminal
See your agent think. Real-time observability for 14 AI agent runtimes - OpenClaw, NVIDIA NemoClaw, Claude Code, Codex & 8 more.