modelscope/evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
RAG evaluation without the need for "golden answers"
Appears on
Quick read
Latest capture 2026-07-31 03:09
1 path
Agent instructions and tool configuration found in this repository.
Claude Code
6 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture +20
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
⚡FlashRAG: A Python Toolkit for Efficient RAG Research (WWW2025 Resource)
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
Supercharge Your LLM Application Evaluations 🚀