← Browse tools

Observability & evaluation

OpenAI Evals

A public registry of evaluation templates combined with an extensible local test runner.

Overview

Research summary

OpenAI Evals supplies a Python framework and a registry of evaluation definitions for assessing models and systems built around them. Developers can use existing test templates, provide private datasets, or implement completion functions for more complex application flows. Evaluations can be run locally through the package and command-line tools. This entry covers the public repository, whose README also points to the separate hosted Evals service. Repository code is MIT licensed, while bundled datasets have their own terms, including attribution, share-alike, and noncommercial conditions.

Repository summary

Stars
Unavailable
Open issues
Unavailable
Last push
Unavailable
Commits, 90 days
Unavailable
Repository activity
Not scored
Version
Unavailable

Recorded catalogue figures. View repository data and provenance →

Classification

Pricing & services

Paid services unknown

Whether the provider offers paid products or services has not been established.

Licence scope

Framework code:MITSelected bundled datasets including PiC/phrase_similarity, alpaca-gpt4, and ToMi:CC-BY-NC-4.0

Other bundled datasets have their own attribution, share-alike, public-domain, and data-license terms. Consult LICENSE.md for the complete per-dataset mapping.

Implementation

Recorded implementation details and interfaces for OpenAI Evals.

Implementation details

Python evaluation runner, templates, and benchmark-data registry

Languages
Python
Repository type
source

Recorded interfaces and capabilities

Licence scope

Framework code:MITSelected bundled datasets including PiC/phrase_similarity, alpaca-gpt4, and ToMi:CC-BY-NC-4.0

Other bundled datasets have their own attribution, share-alike, public-domain, and data-license terms. Consult LICENSE.md for the complete per-dataset mapping.

Repository

Repository snapshots, release information and recorded maintenance signals.

Repository snapshot

openai/evals ↗

Stars
Unavailable
Open issues
Unavailable
Last push
Unavailable
Commits, 90 days
Unavailable
Repository activity
Not scored
Archived
Not recorded

Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).

Maintenance and provenance

Catalogue snapshot
2026-10-06

Documentation

Recorded references and research provenance for this entry.

Recorded sources 3

Research metadata

Research date
2026-10-02
Catalogue snapshot
2026-10-06