modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Collection of evals for Inspect AI
Appears on
Quick read
Latest capture 2026-08-03 03:03
52 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
Claude Code 48
Devin 3
28 more paths detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +100
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
An agent benchmark with tasks in a simulated software company.
SWE-bench: Can Language Models Resolve Real-world Github Issues?
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
The LLM Evaluation Framework
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents