Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
Awesome-LLM: a curated list of Large Language Model
GitHub stars and default-branch commits for Hannibal046/Awesome-LLM.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
A high-throughput and memory-efficient inference and serving engine for LLMs
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
Making large AI models cheaper, faster and more accessible
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and Llama) RAG and Agent app with langchain
DSPy: The framework for programming—not prompting—language models
SGLang is a high-performance serving framework for large language models and multimodal models.
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Ongoing research training transformer models at scale
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
A web interface for chatting with Alpaca through llama.cpp. Fully dockerized, with an easy to use API.
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
AdalFlow: The library to build & auto-optimize LLM applications.
The platform for LLM evaluations and AI agent testing
Infinity is a high-throughput, low-latency serving engine for text-embeddings, reranking models, clip, clap and colpali
A security scanner for your LLM agentic workflows