ocrmypdf/OCRmyPDF
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Repository profile
Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr
Repository updates
Get generated arcships/light-ocr development summaries by email, or follow the weekly and monthly RSS feeds.
Sign in to subscribe by email. RSS feeds are public.
Sign in to subscribeTracked growth, recent movement, and commit velocity from stored repository snapshots.
Latest capture 2026-07-28 10:52
1 capture since 2026-07-28
Stars from baseline 0
All tracked data
Frameworks, package managers, ecosystems, and dependency manifests found during catalog scans.
Scanned 2026-07-28 10:52
CMakeLists.txt
c-cpp ecosystem,
1 dependency
package.json
javascript ecosystem,
2 dependencies
package-lock.json
javascript ecosystem,
100 dependencies
bindings/node/CMakeLists.txt
c-cpp ecosystem,
1 dependency
bindings/node/package.json
javascript ecosystem,
0 dependencies
packages/light-ocr-document/package.json
javascript ecosystem,
1 dependency
packages/light-ocr-medium/package.json
javascript ecosystem,
2 dependencies
packages/light-ocr-server/package.json
javascript ecosystem,
3 dependencies
Searchable topics, generated tags, and stack labels that explain where this repository fits.
Agent instructions and tool configuration paths found in the repository tree.
AI agent config detected
Nearest indexed repositories by embedding similarity.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
AI-powered StartUp Accelerator Engine built with Next.js, LangChain, PostgreSQL + pgvector. Upload, organize, and chat with documents. Includes predictive missing-document detection, role-based workflows, and page-level insight extraction.
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
🔍 Better text detection by combining multiple OCR engines (EasyOCR, Tesseract, and Pororo) with 🧠 LLM.
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.