github Actively maintained

simplecto/sitemap_grabber

A python library to recursively crawl every sitemap.xml for a website. Also handles robots.txt and other well-knowns.

0 awesome lists

Quick read

Stars
1
Forks
0
Open issues
5
Commits
77

Activity and growth

Latest capture 2026-08-11 04:37

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
0
Stored snapshots
46

Classification

Metadata

Language
Python
License
MIT
Default branch
master
Created
2024-06-05
First commit
2024-06-05
Last pushed
2024-10-23
GitHub updated
2025-05-07
Last synced
2026-08-11 04:37
Stack scanned
2026-08-11 04:37
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

46 observed captures since 2026-06-02. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

apify/crawlee-python

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.

9,393 stars
Python 2 awesome lists

s0rg/crawley

The unix-way web crawler

341 stars
Go 2 awesome lists

scanapi/scanapi

Automated Integration Testing and Live Documentation for your API

1,579 stars
Python 1 awesome list

scrapy/scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

63,552 stars
Python 2 awesome lists