← Browse tools

Inference & serving

TensorRT-LLM

NVIDIA-focused inference optimizations exposed through extensible Python and native runtimes.

Overview

Research summary

TensorRT-LLM is NVIDIA's inference framework for optimizing and serving large language models and supported generative workloads on NVIDIA GPUs. It combines specialized kernels, model implementations, runtime scheduling, and Python APIs that developers can configure or extend for their deployment. The repository documents supported hardware, model architectures, and serving examples.

Its current license is scoped: the main project is Apache-2.0, incorporated third-party code retains applicable notices, and the LTX-2 model subtree has separate community-license restrictions. This entry therefore does not describe the entire repository as uniformly Apache licensed.

Repository summary

Stars
Unavailable
Open issues
Unavailable
Last push
Unavailable
Commits, 90 days
Unavailable
Repository activity
Not scored
Version
Unavailable

Recorded catalogue figures. View repository data and provenance →

Classification

Pricing & services

Paid services unknown

Whether the provider offers paid products or services has not been established.

Licence scope

Main project code:Apache-2.0LTX-2 model subtree:LicenseRef-LTX-2-Community

The LTX-2 subtree has separate community-licence restrictions. Incorporated third-party code retains its own notices and terms; Apache-2.0 does not cover the entire repository uniformly.

Implementation

Recorded implementation details and interfaces for TensorRT-LLM.

Implementation details

Python and C++ inference runtime with CUDA kernels

Languages
Python
Repository type
source

Recorded interfaces and capabilities

Licence scope

Main project code:Apache-2.0LTX-2 model subtree:LicenseRef-LTX-2-Community

The LTX-2 subtree has separate community-licence restrictions. Incorporated third-party code retains its own notices and terms; Apache-2.0 does not cover the entire repository uniformly.

Repository

Repository snapshots, release information and recorded maintenance signals.

Repository snapshot

NVIDIA/TensorRT-LLM ↗

Stars
Unavailable
Open issues
Unavailable
Last push
Unavailable
Commits, 90 days
Unavailable
Repository activity
Not scored
Archived
Not recorded

Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).

Maintenance and provenance

Catalogue snapshot
2026-10-06

Documentation

Recorded references and research provenance for this entry.

Recorded sources 5

Research metadata

Research date
2026-10-02
Catalogue snapshot
2026-10-06

Community & social 1 channels

Communities