Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
The context API to search, scrape, and interact with the web at scale. 🔥
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Scrapy, a fast high-level web crawling & scraping framework for Python.
👾 Fast and simple video download library and CLI tool written in Go
Python scraper based on AI
Open source AI job application toolkit in Python: generate a resume and cover letter tailored to each job posting, and drive a stealth browser from any AI client over MCP.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Incredibly fast crawler designed for OSINT.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
A community-driven way to read and chat with AI bots - powered by chatGPT.
Get web data for AI agents and LLMs - fast, efficient, and reliable with Rust
Find web directories without bruteforce
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
❤️ Fredy - [F]ind [R]eal [E]state [D]amn Eas[y] - Fredy keeps searching for new apartments, houses, and flats in Europe on platforms like ImmoScout24, Immowelt, eBay Kleinanzeigen and instantly delivers the results to you via Slack, Telegram, Email, Discord or ntfy, so you can focus on the more important things in life ;)
Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.
Run a high-fidelity browser-based web archiving crawler in a single Docker container
Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG & semantic search. Apache-2.0.
An ergonomic, privacy-aware Rust HTTP Client
A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use.
The unix-way web crawler
Extract web archive data using Wayback Machine and Common Crawl
Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head
An automated website accessibility scanner and cli