Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
A high-throughput and memory-efficient inference and serving engine for LLMs
Port of OpenAI's Whisper model in C/C++
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Making large AI models cheaper, faster and more accessible
Cross-platform, customizable ML solutions for live and streaming media.
SGLang is a high-performance serving framework for large language models and multimodal models.
Faster Whisper transcription with CTranslate2
ncnn is a high-performance neural network inference framework optimized for the mobile platform
Machine Learning Engineering Open Book
Nano vLLM
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
Production ready toolkit to run AI locally
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Adversarial Robustness Toolbox (ART) - Python Library for Machine Learning Security - Evasion, Poisoning, Extraction, Inference - Red and Blue Teams
A programmable Mixture-of-Models router for heterogeneous LLM inference
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.
Fast inference engine for Transformer models
Achieve state of the art inference performance with modern accelerators on Kubernetes
Pre-trained Deep Learning models and demos (high quality and extremely fast)
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️
Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Open-source inference server and production cluster for all the models your agent needs.