LlamaIndex
LlamaIndex delivers the world’s most accurate agentic OCR and document-specific AI workflows, powering complete enterprise automation.
On HuggingFace and in our open-source repos, we are pushing forward document parsing and OCR with several projects:
- ParseBench -- A comprehensive dataset for evaluating document parsing pipelines and OCR models
- LiteParse -- Our lightweight, open-source document parser. Handles multiple formats, runs locally, and integrates with any OCR model
- Visual Document Retrieval -- An exploration into training models purely for document screenshot retrieval
LlamaParse
LlamaParse provides end-to-end document understanding with AI-powered parsing, extraction, and indexing. Transform complex layouts, tables, and handwriting into actionable insights with industry-leading accuracy.
- Parse -- Parse any document, any format, with SOTA accuracy
- Extraxt -- Extract structured data. using your own schemas, from any document
- Classify -- Classify documents into a subset of classes
- Split -- Split documents according to specific categories
- Index -- Index and retrieve your data