Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Awesome List
Curated list of the best truly open-source AI projects, models, tools, and infrastructure. Daily updated.
GitHub stars and default-branch commits for alvinreal/awesome-opensource-ai.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Open-Source Frontier Voice AI
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
Making large AI models cheaper, faster and more accessible
A generative speech model for daily dialogue.
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
SOTA Open Source TTS
Generative Models by Stability AI
SoTA open-source TTS
An open-source RAG-based tool for chatting with your documents.
A simple screen parsing tool towards pure vision based GUI agent
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
FinGPT: Open-Source Financial Large Language Models! Revolutionize 🔥 We release the trained model on HuggingFace.
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
🦉 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
A TTS model capable of generating ultra-realistic dialogue in one pass.
🐫 CAMEL: The first and the best multi-agent framework. Finding the Scaling Law of Agents. https://www.camel-ai.org
Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
No description.
HunyuanVideo: A Systematic Framework For Large Video Generation Model
Go ahead and axolotl questions
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
Pocket Flow: 100-line LLM framework. Let Agents build Agents!
AI Observability & Evaluation
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
tiny vision language model
ModelScope: bring the notion of Model-as-a-Service to life.
[NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
PraisonAI 🦞 — Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute tasks. Deployed in 5 lines of code with built-in memory, RAG, and support for 100+ LLMs.
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks
ZenML 🙏: One AI Platform from Pipelines to Agents. https://zenml.io.
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
Align Anything: Training All-modality Model with Feedback
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
[AAAI 2025] EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditioning
Optimizing inference proxy for LLMs