github Actively maintained

stanford-crfm/helm

Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.

Quick read

Stars
2,872
Forks
407
Open issues
91
Commits
6,298

Activity and growth

Latest capture 2026-08-03 03:12

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+74
Stored snapshots
6

Classification

Technology stack

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2021-11-29
First commit
2021-11-15
Last pushed
2026-08-01
GitHub updated
2026-08-02
Last synced
2026-08-03 03:12
Stack scanned
2026-08-03 03:12
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +74

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

EvolvingLMMs-Lab/lmms-eval

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

4,343 stars
Python 1 awesome list

evalplus/evalplus

Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024

1,803 stars
Python 2 awesome lists

huggingface/lighteval

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

2,501 stars
Python 3 awesome lists

open-compass/VLMEvalKit

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

4,321 stars
Python 2 awesome lists