allenai/olmocr
Toolkit for linearizing PDFs for LLM datasets/training
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
Appears on
Quick read
Latest capture 2026-09-04 10:54
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
1 observed capture since 2026-09-04. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Toolkit for linearizing PDFs for LLM datasets/training
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
The official API server for Exllama. OAI compatible, lightweight, and fast.