Inference & serving
TensorRT-LLM
NVIDIA-focused inference optimizations exposed through extensible Python and native runtimes.
Overview
Research summary
TensorRT-LLM is NVIDIA's inference framework for optimizing and serving large language models and supported generative workloads on NVIDIA GPUs. It combines specialized kernels, model implementations, runtime scheduling, and Python APIs that developers can configure or extend for their deployment. The repository documents supported hardware, model architectures, and serving examples.
Its current license is scoped: the main project is Apache-2.0, incorporated third-party code retains applicable notices, and the LTX-2 model subtree has separate community-license restrictions. This entry therefore does not describe the entire repository as uniformly Apache licensed.
Repository summary
- Stars
- Unavailable
- Open issues
- Unavailable
- Last push
- Unavailable
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Version
- Unavailable
Recorded catalogue figures. View repository data and provenance →
Classification
Pricing & services
Paid services unknown
Whether the provider offers paid products or services has not been established.
Licence scope
The LTX-2 subtree has separate community-licence restrictions. Incorporated third-party code retains its own notices and terms; Apache-2.0 does not cover the entire repository uniformly.
Implementation
Recorded implementation details and interfaces for TensorRT-LLM.
Implementation details
Python and C++ inference runtime with CUDA kernels
- Languages
- Python
- Repository type
- source
Recorded interfaces and capabilities
Licence scope
The LTX-2 subtree has separate community-licence restrictions. Incorporated third-party code retains its own notices and terms; Apache-2.0 does not cover the entire repository uniformly.
Repository
Repository snapshots, release information and recorded maintenance signals.
Repository snapshot
- Stars
- Unavailable
- Open issues
- Unavailable
- Last push
- Unavailable
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Archived
- Not recorded
Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).
Maintenance and provenance
- Catalogue snapshot
- 2026-10-06
Documentation
Recorded references and research provenance for this entry.
Recorded sources 5
- https://nvidia.github.io/TensorRT-LLM Project page · Documentation
- https://github.com/NVIDIA/TensorRT-LLM Linked repository · Research reference
- https://github.com/NVIDIA/TensorRT-LLM/blob/main/README.md Research reference
- https://api.github.com/repos/NVIDIA/TensorRT-LLM Research reference
- https://github.com/NVIDIA/TensorRT-LLM/blob/main/LICENSE Research reference
Research metadata
- Research date
- 2026-10-02
- Catalogue snapshot
- 2026-10-06
Community & social 1 channels
Communities
- GitHub Discussionsgithub.com/NVIDIA/TensorRT-LLM/discussionsOfficial
Link details for GitHub Discussions
- Last checked