Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
LlamaIndex is the leading document agent and OCR platform
The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
A roadmap connecting many of the most important concepts in machine learning, how to learn them and what tools to use to perform them.
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
The leader in Customer Data Infrastructure
Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.
data load tool (dlt) is an open source Python library that makes data loading easy 🛠️
Edit Banana: A framework for converting statistical formats into editable.
CKAN is an open-source DMS (data management system) for powering data hubs and data portals. CKAN makes it easy to publish, share and use data. It powers catalog.data.gov, open.canada.ca/data, data.humdata.org among many other sites.
Mimesis is a Python library for generating fake but realistic data in multiple languages and locales.
DeepAnalyze is the first agentic LLM for autonomous data science. 🎈你的AI数据分析师,自动分析大量数据,一键生成专业分析报告!
RAG (Retrieval Augmented Generation) Framework for building modular, open source applications for production by TrueFoundry
Preswald is a WASM packager for Python-based interactive data apps: bundle full complex data workflows, particularly visualizations, into single files, runnable completely in-browser, using Pyodide, DuckDB, Pandas, and Plotly, Matplotlib, etc. Build dashboards, reports, and notebooks that run offline, load fast, and share like a document.
A system for agentic LLM-powered data processing and ETL
Easiest and laziest way for building multi-agent LLMs applications.
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
Extract data from a wide range of Internet sources into a pandas DataFrame.
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.
Deepnote is a drop-in replacement for Jupyter with an AI-first design, sleek UI, new blocks, and native data integrations. Use Python, R, and SQL locally in your favorite IDE, then scale to Deepnote cloud for real-time collaboration, Deepnote agent, and deployable data apps. https://deepnote.com/
The fastest business intelligence tool for humans and agents.
QuantMind is an agent-native knowledge extraction and retrieval framework for quantitative finance.
Colour Science for Python
Jupyter extensions that help you write code faster: Context aware AI Chat, Autocomplete, and Spreadsheet