github Actively maintained

midwork-finds-jobs/duckdb-web-archive

DuckDB extension to fetch pages from Wayback Machine & Common Crawl

1 awesome list

Quick read

Stars
22
Forks
1
Open issues
2
Commits
75

Activity and growth

Latest capture 2026-07-31 03:05

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+2
Stored snapshots
6

Classification

Metadata

Language
C++
License
MIT
Default branch
main
Created
2025-11-21
First commit
2025-11-21
Last pushed
2026-06-27
GitHub updated
2026-07-28
Last synced
2026-07-31 03:05
Stack scanned
2026-07-31 03:05
Archived
No

AI development signals

1 path

Agent instructions and tool configuration found in this repository.

Claude Code

View paths

Growth history

Tracked growth

6 observed captures since 2026-05-23. Observed captures are shown by default.

Stars from first capture +2

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

Query-farm/httpserver

DuckDB HTTP API Server and Query Interface in a Community Extension

284 stars
C++ 1 awesome list

karust/gogetcrawl

Extract web archive data using Wayback Machine and Common Crawl

183 stars
Go 1 awesome list

duckdb/duckdb

DuckDB is an analytical in-process SQL database management system

39,892 stars
C++ 2 awesome lists

mehd-io/duckdb-extension-radar

This repo contains information about DuckDB extensions found on GitHub. Refreshed daily

114 stars
Python 1 awesome list

autumoswitzerland/Webduck

A self-hosted DuckDB-as-a-Service server with REST API and Web UI.

50 stars
Python 0 awesome lists