github Archived

ocropus-archive/DUP-ocropy

Python-based tools for document analysis and OCR

1 awesome list

Quick read

Stars
3,465
Forks
588
Open issues
83
Commits
998

Activity and growth

Latest capture 2026-09-01 03:04

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
-3
Stored snapshots
4

Metadata

Language
Jupyter Notebook
License
Apache-2.0
Default branch
master
Created
2014-10-28
First commit
2010-03-03
Last pushed
2021-05-22
GitHub updated
2026-08-10
Last synced
2026-09-01 03:04
Stack scanned
2026-09-01 03:04
Archived
Yes

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

4 observed captures since 2026-06-24. Observed captures are shown by default.

Stars from first capture -3

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

ocrmypdf/OCRmyPDF

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

34,614 stars
Python 1 awesome list

allenai/olmocr

Toolkit for linearizing PDFs for LLM datasets/training

19,454 stars
Python 3 awesome lists

JaidedAI/EasyOCR

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

29,930 stars
Python 1 awesome list

arcships/light-ocr

Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr

490 stars
C++ 1 awesome list

CatchTheTornado/text-extract-api

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

3,178 stars
Python 1 awesome list