Soochi
All entries In Technology

Tool

Tesseract OCR

Open-source optical character recognition engine supporting over 100 languages, widely used to extract text from scanned documents and PDFs

tesseract-ocr.github.io

Tesseract is an open-source optical character recognition (OCR) engine originally developed at HP and maintained by Google and the open-source community. It uses LSTM-based neural network models to recognize text from scanned images, photos, and PDF files in more than 100 languages, including multiple Indian scripts (Devanagari, Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam, and Gurmukhi).

Tesseract is the foundational OCR engine powering numerous higher-level civic tech tools, including Paperless-ngx, Stirling PDF, and custom document pipelines used by investigative journalists and archival projects to digitize public records, court filings, and historical documents. It can be run from the command line or integrated programmatically via Python and C++ bindings.

Details

Kind
Tool
Run by
Collective
Topics
Technology
Where
Global, India
Licence
Open licence
Cost
Free
Licence terms
Apache-2.0

Elsewhere