Open highlighted repo slot
Put your repository first
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
GitHub projects from awesome lists
Search names, descriptions, topics, tags, and stacks, then tune results by ecosystem, freshness, health, and cross-list signal.
Open highlighted repo slot
Promote a GitHub repo at the top of Awesome repository list views for 7 days.
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
Render OG images and paged PDFs from JSX, HTML, and CSS. No headless browser. Runs on Node.js, Cloudflare Workers, browsers, and Rust.
Open Source Document Management System for Digital Archives (Scanned Documents)
A Python library for reading and writing PDF, powered by QPDF
A powerful Zotero AI and MCP plugin with ChatGPT, Gemini 3.7, Claude Fable 5, Claude Opus 5, DeepSeek V4, Grok, OpenRouter, Kimi k3, GLM 5.3, SiliconFlow, GPT-oss, Gemma 4, Qwen 3.8
Cross-platform desktop GUI app to clean image metadata
PDF exporter for HTML presentations
The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal
Document scanning app
Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
NOW MANAGED ON CODEBERG
A web interface to extract tabular data from PDFs
kramdown is a fast, pure Ruby Markdown superset converter, using a strict syntax definition and supporting several common extensions.
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
Versatile PDF creation and manipulation for Ruby
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
A Python Library for Generating PDFs and Images from HTML, powered by PlutoBook
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
Free open-source web software for signing PDF (alone or with others) and also organize pages, edit metadata and compress pdf
Control how your documents get signed
A Pure ruby library to merge PDF files, number pages and maybe more...
Papermerge DMS core backend, REST API server, and frontend UI
An open-source read-along document reader server with high-quality TTS options, synchronized highlighting, and audiobook export for EPUB, PDF, DOCX, TXT, and MD.
Mercury - Train your own custom GPT. Chat with any file, or website.
I, Librarian - open-source version of a PDF managing SaaS.
📚 Simplifying the way you read. Minimalistic Vim-like TUI document reader.
A tool for tagging files and archiving tasks.
Imgcompress is a self-hosted image processing toolbox that handles compression, format conversion, and AI background removal in a single web interface. It supports over 70 input formats (including PSD, HEIC, and RAW) and can output common formats or generate PDFs, all without sending files to external services.
Your bibliography on the command line
Tool for extracting pages from pdf as images and text as strings.