raphael-seo/Versatile-OCR-Program
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Python-based tools for document analysis and OCR
Appears on
Quick read
Latest capture 2026-07-29 03:08
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
3 observed captures since 2026-06-24. Charts use measured snapshots only.
Stars from first capture -2
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Toolkit for linearizing PDFs for LLM datasets/training
Fast, offline OCR for Node.js & C++. PP-OCRv6 with Core ML / WebGPU hardware acceleration — recognize text in images with confidence scores & coordinates. npm: @arcships/light-ocr
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
An open source implementation of CLIP.