← Browse tools

Observability & evaluation

Lighteval

Backend-flexible benchmarking with detailed per-sample outputs for analyzing model differences.

Overview

Research summary

Lighteval is a Python toolkit for evaluating language models with reusable tasks and metrics. It can evaluate models loaded in memory or reached through supported serving backends and inference APIs. Runs retain detailed sample results so developers can inspect failures instead of relying only on an aggregate score. Users can define custom tasks, metrics, and model adapters, and the toolkit integrates with the Hugging Face ecosystem. Its current evaluation interface also supports Inspect AI as a backend alongside other execution paths.

Repository summary

Stars
Unavailable
Open issues
Unavailable
Last push
Unavailable
Commits, 90 days
Unavailable
Repository activity
Not scored
Version
Unavailable

Recorded catalogue figures. View repository data and provenance →

Classification

Pricing & services

Paid services unknown

Whether the provider offers paid products or services has not been established.

Licence scope

Implementation

Recorded implementation details and interfaces for Lighteval.

Implementation details

Python evaluation toolkit and command-line interfaces

Languages
Python
Repository type
source

Recorded interfaces and capabilities

Licence scope

Repository

Repository snapshots, release information and recorded maintenance signals.

Repository snapshot

huggingface/lighteval ↗

Stars
Unavailable
Open issues
Unavailable
Last push
Unavailable
Commits, 90 days
Unavailable
Repository activity
Not scored
Archived
Not recorded

Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).

Maintenance and provenance

Catalogue snapshot
2026-10-06

Documentation

Recorded references and research provenance for this entry.

Recorded sources 4

Research metadata

Research date
2026-10-02
Catalogue snapshot
2026-10-06