Observability & evaluation
LM Evaluation Harness
A broad, reusable benchmark task ecosystem with pluggable model backends and scoring components.
Overview
Research summary
LM Evaluation Harness provides a shared framework for measuring language-model performance across reusable benchmark tasks. It combines task definitions, prompts, metrics, model adapters, and result reporting in one Python package. Developers can run locally hosted models or supported APIs, configure tasks with YAML, and extend the system with additional backends and scoring components. The project also supports plugin registration so downstream packages can add evaluation functionality. It is primarily a model benchmarking framework, complementing application-specific tracing and production monitoring tools.
Repository summary
- Stars
- Unavailable
- Open issues
- Unavailable
- Last push
- Unavailable
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Version
- Unavailable
Recorded catalogue figures. View repository data and provenance →
Classification
Pricing & services
Paid services unknown
Whether the provider offers paid products or services has not been established.
Licence scope
Implementation
Recorded implementation details and interfaces for LM Evaluation Harness.
Implementation details
Python model-benchmarking framework and CLI
- Languages
- Python
- Repository type
- source
Recorded interfaces and capabilities
Licence scope
Repository
Repository snapshots, release information and recorded maintenance signals.
Repository snapshot
EleutherAI/lm-evaluation-harness ↗
- Stars
- Unavailable
- Open issues
- Unavailable
- Last push
- Unavailable
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Archived
- Not recorded
Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).
Maintenance and provenance
- Catalogue snapshot
- 2026-10-06
Documentation
Recorded references and research provenance for this entry.
Recorded sources 4
- https://github.com/EleutherAI/lm-evaluation-harness Project page · Linked repository · Research reference
- https://github.com/EleutherAI/lm-evaluation-harness/blob/main/README.md Research reference
- https://github.com/EleutherAI/lm-evaluation-harness/blob/main/LICENSE.md Research reference
- https://www.eleuther.ai/ Research reference
Research metadata
- Research date
- 2026-10-02
- Catalogue snapshot
- 2026-10-06