ocrmypdf/OCRmyPDF
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Quick read
Latest capture 2026-08-11 04:41
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
43 observed captures since 2026-06-02. Observed captures are shown by default.
Stars from first capture -7
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
🔍 Better text detection by combining multiple OCR engines (EasyOCR, Tesseract, and Pororo) with 🧠 LLM.
Toolkit for linearizing PDFs for LLM datasets/training
Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
Get your documents ready for gen AI