Tesseract OCR
Open-source optical character recognition engine supporting over 100 languages, widely used to extract text from scanned documents and PDFs
tesseract-ocr.github.ioTesseract is an open-source optical character recognition (OCR) engine originally developed at HP and maintained by Google and the open-source community. It uses LSTM-based neural network models to recognize text from scanned images, photos, and PDF files in more than 100 languages, including multiple Indian scripts (Devanagari, Tamil, Telugu, Bengali, Gujarati, Kannada, Malayalam, and Gurmukhi).
Tesseract is the foundational OCR engine powering numerous higher-level civic tech tools, including Paperless-ngx, Stirling PDF, and custom document pipelines used by investigative journalists and archival projects to digitize public records, court filings, and historical documents. It can be run from the command line or integrated programmatically via Python and C++ bindings.
Details
- Kind
- Tool
- Run by
- Collective
- Topics
- Where
- Global, India
- Licence
- Open licence
- Cost
- Free
- Licence terms
- Apache-2.0