Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
A curated list of awesome LLM frameworks, libraries and software.
GitHub stars and default-branch commits for uhub/awesome-llm.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Unified framework for building enterprise RAG pipelines with small, specialized models
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
structured outputs for llms
Talk to any LLM with hands-free voice interaction, voice interruption, and Live2D taking face running locally across platforms
🌐 The open-source Agentic browser; alternative to ChatGPT Atlas, Perplexity Comet, Dia.
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows
Pocket Flow: Codebase to Tutorial
Go ahead and axolotl questions
Access large language models from the command-line
Retrieval and Retrieval-augmented LLMs
Expose your FastAPI endpoints as Model Context Protocol (MCP) tools, with Auth!
Private chat with local GPT with document, images, video, etc. 100% private, Apache 2.0. Supports oLLaMa, Mixtral, llama.cpp, and more. Demo: https://gpt.h2o.ai/ https://gpt-docs.h2o.ai/
Graph-Native Infrastructure for Context and Accountable AI Systems
Build high-quality LLM apps - from prototyping, testing to production deployment and monitoring.
Pocket Flow: 100-line LLM framework. Let Agents build Agents!
QLoRA: Efficient Finetuning of Quantized LLMs
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
A collection of GPT system prompts and various prompt injection/leaking knowledge.
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings
🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading
Build, Manage and Deploy AI/ML Systems
"AutoAgent: Fully-Automated and Zero-Code LLM Agent Framework"
High-speed Large Language Model Serving for Local Deployment
🐚 Python-powered shell. Full-featured, cross-platform and AI-friendly.
Cybersecurity AI (CAI), the framework for AI Security
ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source.
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
UFO³: Weaving the Digital Agent Galaxy
Running large language models on a single GPU for throughput-oriented scenarios.
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
Automate your mobile devices with natural language commands - an LLM agnostic mobile Agent 🤖
KAG is a logical form-guided reasoning and retrieval framework based on OpenSPG engine and LLMs. It is used to build logical reasoning and factual Q&A solutions for professional domain knowledge bases. It can effectively overcome the shortcomings of the traditional RAG vector similarity calculation model.
the LLM vulnerability scanner
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
PraisonAI 🦞 — Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute tasks. Deployed in 5 lines of code with built-in memory, RAG, and support for 100+ LLMs.
An Autonomous LLM Agent for Complex Task Solving
Build effective agents using Model Context Protocol and simple workflow patterns
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!
Accessible large language models via k-bit quantization for PyTorch.
Semantic cache for LLMs. Fully integrated with LangChain and llama_index.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
FinRobot: An Open-Source AI Agent Platform for Financial Applications using Large Language Models
Reverse engineered API of Microsoft's Bing Chat AI