allenai/olmocr
Toolkit for linearizing PDFs for LLM datasets/training
Convert PDF to markdown + JSON quickly with high accuracy
Appears on
Quick read
Latest capture 2026-08-03 04:44
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
42 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +2809
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Toolkit for linearizing PDFs for LLM datasets/training
Python tool for converting files and office documents to Markdown.
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
AI-native ontology engine: a Rust MCP server with tools for building, validating, querying, and reasoning over RDF/OWL ontologies. In-memory Oxigraph triple store, native OWL2-DL tableaux reasoner, SHACL validation, SPARQL, versioning. Single binary, no JVM.
☁️ The fastest HTML to markdown convertor on GitHub. Optimized for LLMs and supports streaming.
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI