github Actively maintained

NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Quick read

Stars
14,283
Forks
2,632
Open issues
1,593
Commits
8,311

Activity and growth

Latest capture 2026-08-03 03:07

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+558
Stored snapshots
6

Classification

Metadata

Language
Python
License
NOASSERTION
Default branch
main
Created
2023-08-16
First commit
2023-09-20
Last pushed
2026-08-03
GitHub updated
2026-08-03
Last synced
2026-08-03 03:07
Stack scanned
2026-08-03 03:07
Archived
No

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +558

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3,371 stars
Python 2 awesome lists

llm-d/llm-d

Achieve state of the art inference performance with modern accelerators on Kubernetes

3,959 stars
Shell 2 awesome lists

ikawrakow/ik_llama.cpp

llama.cpp fork with additional SOTA quants and improved performance

2,989 stars
C++ 1 awesome list

InternLM/lmdeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

7,985 stars
Python 3 awesome lists

tensorzero/tensorzero

TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.

11,728 stars
Rust 5 awesome lists

hiyouga/LlamaFactory

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

73,697 stars
Python 2 awesome lists