Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
PyTorch native quantization for training and inference
Community maintained hardware plugin for vLLM on Ascend
Turn any computer or edge device into a command center for your computer vision projects.
Mastering Applied AI, One Concept at a Time
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
Manages Unified Access to Generative AI Services built on Envoy Gateway
Continual learning infra for self-improving agents
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
[deprecated] AI Gateway - core infrastructure stack for building production-ready AI Applications
[deprecated] AI Gateway - core infrastructure stack for building production-ready AI Applications
A DuckDB extension for in-database inference
Learn how LLMs work by building one in Zig -- from tensors to text generation.