github Actively maintained

Dicklesworthstone/llm_aided_ocr

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

1 awesome list

Quick read

Stars
2,996
Forks
214
Open issues
0
Commits
64

Activity and growth

Latest capture 2026-09-04 10:55

Stars · last 7 days
0 0.0%
Commits · last 7 days
0 0.0%
Stars since tracking
0
Stored snapshots
1

Classification

Metadata

Language
Python
License
NOASSERTION
Default branch
main
Created
2023-07-26
First commit
2023-07-26
Last pushed
2026-08-03
GitHub updated
2026-09-04
Last synced
2026-09-04 10:55
Stack scanned
2026-09-04 10:55
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

1 observed capture since 2026-09-04. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

CatchTheTornado/text-extract-api

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

3,178 stars
Python 1 awesome list

ocrmypdf/OCRmyPDF

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

34,614 stars
Python 1 awesome list

Anil-matcha/AI-Youtube-Shorts-Generator

Open-source alternative to Opus Clip, Vidyo.ai, Klap & SubMagic. Turn long-form YouTube videos into viral 9:16 shorts using LLM highlight detection, Whisper transcription, and auto vertical cropping — free, no watermarks, no per-clip credits.

4,821 stars
Python 1 awesome list

allenai/olmocr

Toolkit for linearizing PDFs for LLM datasets/training

19,252 stars
Python 3 awesome lists