arcships/light-ocr
Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Appears on
Quick read
Latest capture 2026-08-30 03:05
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
4 observed captures since 2026-06-24. Observed captures are shown by default.
Stars from first capture +655
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
A Python library for reading and writing PDF, powered by QPDF