github Actively maintained

openai/evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

Quick read

Stars
19,090
Forks
3,042
Open issues
218
Commits
691

Activity and growth

Latest capture 2026-08-02 13:54

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+559
Stored snapshots
7

Classification

Technology stack

Metadata

Language
Python
License
NOASSERTION
Default branch
main
Created
2023-01-23
First commit
2023-03-14
Last pushed
2026-04-14
GitHub updated
2026-08-02
Last synced
2026-08-02 13:54
Stack scanned
2026-08-02 13:54
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

7 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +559

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

3,177 stars
Python 3 awesome lists

evidentlyai/evidently

Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

7,778 stars
Jupyter Notebook 5 awesome lists

huggingface/lighteval

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

2,501 stars
Python 3 awesome lists

hidai25/eval-view

Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic.

129 stars
Python 1 awesome list

evalplus/evalplus

Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024

1,803 stars
Python 2 awesome lists