InternLM/WildClawBench
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
Appears on
Quick read
Latest capture 2026-08-03 03:08
60 paths
Agent instructions and tool configuration found in this repository.
Agent workspace 60
36 more paths detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +105
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
BenchClaw — Multi-dimensional AI agent evaluation with 17-judge AI Tribunal, 10 scoring dimensions, radar charts, and deception detection. Benchmark any LLM agent.
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Open-source benchmark for browser AI agents on daily tasks.
An agent benchmark with tasks in a simulated software company.
A benchmark for LLMs on complicated tasks in the terminal