github Actively maintained

s0rg/crawley

The unix-way web crawler

2 awesome lists

Quick read

Stars
341
Forks
18
Open issues
8
Commits
158

Activity and growth

Latest capture 2026-08-22 03:02

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+1
Stored snapshots
7

Classification

Metadata

Language
Go
License
MIT
Default branch
master
Created
2021-10-27
First commit
2021-10-27
Last pushed
2026-08-21
GitHub updated
2026-08-09
Last synced
2026-08-22 03:02
Stack scanned
2026-08-22 03:02
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

7 observed captures since 2026-05-22. Observed captures are shown by default.

Stars from first capture +1

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

apify/crawlee-python

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.

9,393 stars
Python 2 awesome lists

karust/gogetcrawl

Extract web archive data using Wayback Machine and Common Crawl

183 stars
Go 1 awesome list

AIMLPM/markcrawl

Fast Python web crawler for RAG and AI ingestion. Extracts clean Markdown from any site for LLMs and vector stores.

3 stars
Python 1 awesome list

Pyx-Corp/spectrawl

The unified web layer for AI agents. Search (8 engines), stealth browse, auth, and act on 24 platforms. One npm install, self-hosted.

28 stars
JavaScript 1 awesome list

N0taN3rd/Squidwarc

Squidwarc is a high fidelity, user scriptable, archival crawler that uses Chrome or Chromium with or without a head

178 stars
JavaScript 1 awesome list

ArchiveTeam/grab-site

The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns

1,602 stars
Python 1 awesome list