Skip to content

lm-evaluation-harness

Command-line and Python harness for running few-shot benchmark tasks against language models across many backends.

By EleutherAI · github.com/EleutherAI/lm-evaluation-harness

Package download activity

Each package is shown on its own; packages that can be installed together are never added up. Week 2026-W40 (2026-09-28 to 2026-10-04).

PyPIlm-eval · headline (ranked)

270,051 downloads in the week, +10.3% vs prior week · rank 14 on PyPI

2026-W29: 339,6192026-W30: 365,5092026-W31: 374,2972026-W32: 428,1612026-W33: 442,0502026-W34: 445,9982026-W35: 269,4752026-W36: 264,7392026-W37: 226,6542026-W38: 237,7812026-W39: 244,8042026-W40: 270,051
2026-W29 to 2026-W40, 12 weeks; min 226,654, max 445,998

Reproduce this number (ClickHouse SQL) · registry metadata (mapping checked 2026-10-07)

GitHub stars

EleutherAI/lm-evaluation-harness: 14,147 stars (rank 9). Stars are a point-in-time count of a repository's bookmarks, not people using the tool.

Series starts 2026-W40 (14,147).

Reproduce (GitHub API) · repository link source

Facts

Facts for lm-evaluation-harness
CategoriesEvaluation
StatusGenerally available
LicenceMITsource, checked 2026-10-07
Self-hostYessource, checked 2026-10-07
OpenTelemetryUnknownsource, checked 2026-10-07
Pricing modelfree
Entry priceFree; Open-source library; no paid plansource, checked 2026-10-07
Free tierYes
MeasurabilityMeasured (mapped SDK packages)

Facts last verified 2026-10-07 on first-party pages. Prices as published, taxes excluded; non-USD prices are not converted.

Compare

Spotted an outdated fact? The methodology page explains how facts are checked; corrections are applied at the next weekly build.