Observability & evaluation
OpenAI Evals
A public registry of evaluation templates combined with an extensible local test runner.
Overview
Research summary
OpenAI Evals supplies a Python framework and a registry of evaluation definitions for assessing models and systems built around them. Developers can use existing test templates, provide private datasets, or implement completion functions for more complex application flows. Evaluations can be run locally through the package and command-line tools. This entry covers the public repository, whose README also points to the separate hosted Evals service. Repository code is MIT licensed, while bundled datasets have their own terms, including attribution, share-alike, and noncommercial conditions.
Repository summary
- Stars
- Unavailable
- Open issues
- Unavailable
- Last push
- Unavailable
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Version
- Unavailable
Recorded catalogue figures. View repository data and provenance →
Classification
Pricing & services
Paid services unknown
Whether the provider offers paid products or services has not been established.
Licence scope
Other bundled datasets have their own attribution, share-alike, public-domain, and data-license terms. Consult LICENSE.md for the complete per-dataset mapping.
Implementation
Recorded implementation details and interfaces for OpenAI Evals.
Implementation details
Python evaluation runner, templates, and benchmark-data registry
- Languages
- Python
- Repository type
- source
Recorded interfaces and capabilities
Licence scope
Other bundled datasets have their own attribution, share-alike, public-domain, and data-license terms. Consult LICENSE.md for the complete per-dataset mapping.
Repository
Repository snapshots, release information and recorded maintenance signals.
Repository snapshot
- Stars
- Unavailable
- Open issues
- Unavailable
- Last push
- Unavailable
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Archived
- Not recorded
Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).
Maintenance and provenance
- Catalogue snapshot
- 2026-10-06
Documentation
Recorded references and research provenance for this entry.
Recorded sources 3
- https://github.com/openai/evals Project page · Linked repository · Research reference
- https://github.com/openai/evals/blob/main/README.md Research reference
- https://github.com/openai/evals/blob/main/LICENSE.md Research reference
Research metadata
- Research date
- 2026-10-02
- Catalogue snapshot
- 2026-10-06