← Browse tools

Inference & serving

ExLlamaV3

A GPU-oriented language-model inference engine and EXL3 quantization toolkit for running supported models on local hardware.

1,422 stars

Overview

Research summary

ExLlamaV3 supplies the low-level inference and quantization layer for running language models efficiently on supported GPU hardware. The project includes EXL3 model quantization and runtime capabilities such as batching, parallel execution across devices, and options for managing memory and computation placement. It is an engine and toolkit rather than a complete chat product, coding agent, or hosted model service.

Applications can integrate the Python-facing implementation or use a separate serving project such as TabbyAPI to expose an API to clients. Practical model size, speed, and memory use depend on the selected checkpoint, quantization settings, workload, and hardware configuration, so a repository-level performance claim is not a universal guarantee. The code is MIT licensed; compatible model weights must be obtained and licensed separately.

Repository summary

Stars
1,422
Open issues
Unavailable
Last push
2026-09-15
Commits, 90 days
Unavailable
Repository activity
Not scored
Version
Unavailable

Recorded catalogue figures. View repository data and provenance →

Classification

Pricing & services

Paid services unknown

Whether the provider offers paid products or services has not been established.

Licence scope

Implementation

Recorded implementation details and interfaces for ExLlamaV3.

Implementation details

Python inference tooling with GPU kernels and EXL3 quantization

Languages
Python
Repository type
source

Recorded interfaces and capabilities

Licence scope

Repository

Repository snapshots, release information and recorded maintenance signals.

Repository snapshot

turboderp-org/exllamav3 ↗

Stars
1,422
Open issues
Unavailable
Last push
2026-09-15
Commits, 90 days
Unavailable
Repository activity
Not scored
Archived
Not recorded

Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).

Maintenance and provenance

Catalogue snapshot
2026-10-06
Stars source
Recorded fallback
Last push source
Recorded fallback

Documentation

Recorded references and research provenance for this entry.

Recorded sources 6

Research metadata

Research date
2026-09-16
Catalogue snapshot
2026-10-06

Community & social 1 channels

Communities