Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
☁️ Build multimodal AI applications with cloud-native stack
YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)
Janus-Series: Unified Multimodal Understanding and Generation Models
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/
Production ready toolkit to run AI locally
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Open-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google
Solve Visual Understanding with Reinforced VLMs
A visual playground for agentic workflows: Iterate over your agents 10x faster
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
A Next-Generation Training Engine Built for Ultra-Large MoE Models
Align Anything: Training All-modality Model with Feedback
Plug in and Play Implementation of Tree of Thoughts: Deliberate Problem Solving with Large Language Models that Elevates Model Reasoning by atleast 70%
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai).
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.