Inference & serving
ExLlamaV3
A GPU-oriented language-model inference engine and EXL3 quantization toolkit for running supported models on local hardware.
Overview
Research summary
ExLlamaV3 supplies the low-level inference and quantization layer for running language models efficiently on supported GPU hardware. The project includes EXL3 model quantization and runtime capabilities such as batching, parallel execution across devices, and options for managing memory and computation placement. It is an engine and toolkit rather than a complete chat product, coding agent, or hosted model service.
Applications can integrate the Python-facing implementation or use a separate serving project such as TabbyAPI to expose an API to clients. Practical model size, speed, and memory use depend on the selected checkpoint, quantization settings, workload, and hardware configuration, so a repository-level performance claim is not a universal guarantee. The code is MIT licensed; compatible model weights must be obtained and licensed separately.
Repository summary
- Stars
- 1,422
- Open issues
- Unavailable
- Last push
- 2026-09-15
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Version
- Unavailable
Recorded catalogue figures. View repository data and provenance →
Classification
Pricing & services
Paid services unknown
Whether the provider offers paid products or services has not been established.
Licence scope
Implementation
Recorded implementation details and interfaces for ExLlamaV3.
Implementation details
Python inference tooling with GPU kernels and EXL3 quantization
- Languages
- Python
- Repository type
- source
Recorded interfaces and capabilities
Licence scope
Repository
Repository snapshots, release information and recorded maintenance signals.
Repository snapshot
- Stars
- 1,422
- Open issues
- Unavailable
- Last push
- 2026-09-15
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Archived
- Not recorded
Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).
Maintenance and provenance
- Catalogue snapshot
- 2026-10-06
- Stars source
- Recorded fallback
- Last push source
- Recorded fallback
Documentation
Recorded references and research provenance for this entry.
Recorded sources 6
- https://github.com/turboderp-org/exllamav3 Project page · Linked repository · Research reference
- https://github.com/turboderp-org/exllamav3/blob/master/README.md Research reference
- https://github.com/turboderp-org/exllamav3/blob/master/LICENSE Research reference
- https://github.com/stars/angelcervera/lists/ai Research reference
- https://raw.githubusercontent.com/turboderp-org/exllamav3/HEAD/README.md Research reference
- https://discord.com/api/v10/invites/NSFwVuCjRq?with_counts=false&with_expiration=true Research reference
Research metadata
- Research date
- 2026-09-16
- Catalogue snapshot
- 2026-10-06
Community & social 1 channels
Communities
- Discorddiscord.gg/NSFwVuCjRqOfficial
Link details for Discord
- Access
- Public
- Last checked