LiteParse
Fast, helpful, open-source document parser from LlamaIndex — extracts structured text from PDFs, Word, and other formats with high accuracy for LLM pipelines
github.comLiteParse is an open-source document parser built by the LlamaIndex team (run-llama), designed for the LLM era. Unlike traditional PDF parsers that struggle with tables, multi-column layouts, and complex formatting, LiteParse uses vision-language models to understand document structure and extract text with high fidelity. For civic technologists and researchers working with Indian government data — which is overwhelmingly published as scanned PDFs, multi-column tables, and poorly formatted spreadsheets — LiteParse represents a significant capability: the ability to reliably convert government documents into structured, machine-readable text at scale. The tool is MIT-licensed, installable via pip, and integrates naturally with LlamaIndex’s broader framework for building LLM applications over documents.
Details
- Kind
- Tool
- Run by
- Commercial
- Topics
- Where
- Global
- Licence
- Open licence
- Cost
- Free
- Licence terms
- MIT