Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
A high-throughput and memory-efficient inference and serving engine for LLMs
The open-source AI voice studio. Clone, dictate, create.
SGLang is a high-performance serving framework for large language models and multimodal models.
Build and run Docker containers leveraging NVIDIA GPUs
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Open3D: A Modern Library for 3D Data Processing
NumPy & SciPy for GPU
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
NumPy aware dynamic Python compiler using LLVM
cuDF - GPU DataFrame Library
cuDF - GPU DataFrame Library
Containers for machine learning
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
FlashInfer: Kernel Library for LLM Serving
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
NVIDIA cuML: GPU-Accelerated Machine Learning
cuML - RAPIDS Machine Learning Library
Fast inference engine for Transformer models
A retargetable MLIR-based machine learning compiler and runtime toolkit.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.
Self-hosted, local only NVR and AI Computer Vision software. With features such as object detection, motion detection, face recognition and more, it gives you the power to keep an eye on your home, office or any other place you want to monitor.
Single C file, Realtime CPU/GPU Profiler with Remote Web Viewer
how to optimize some algorithm in cuda.
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
PyTorch native quantization for training and inference
LLM speculative inference server for heterogeneous hardware & consumer GPUs