lm-evaluation-harness
Command-line and Python harness for running few-shot benchmark tasks against language models across many backends.
By EleutherAI · github.com/EleutherAI/lm-evaluation-harness
Package download activity
Each package is shown on its own; packages that can be installed together are never added up. Week 2026-W40 (2026-09-28 to 2026-10-04).
PyPIlm-eval · headline (ranked)
270,051 downloads in the week, +10.3% vs prior week · rank 14 on PyPI
Reproduce this number (ClickHouse SQL) · registry metadata (mapping checked 2026-10-07)
GitHub stars
EleutherAI/lm-evaluation-harness: 14,147 stars (rank 9). Stars are a point-in-time count of a repository's bookmarks, not people using the tool.
Series starts 2026-W40 (14,147).
Facts
| Categories | Evaluation | |
|---|---|---|
| Status | Generally available | |
| Licence | MIT | source, checked 2026-10-07 |
| Self-host | Yes | source, checked 2026-10-07 |
| OpenTelemetry | Unknown | source, checked 2026-10-07 |
| Pricing model | free | |
| Entry price | Free; Open-source library; no paid plan | source, checked 2026-10-07 |
| Free tier | Yes | |
| Measurability | Measured (mapped SDK packages) |
Facts last verified 2026-10-07 on first-party pages. Prices as published, taxes excluded; non-USD prices are not converted.
Compare
Spotted an outdated fact? The methodology page explains how facts are checked; corrections are applied at the next weekly build.