Open-source alternatives to ABBYY FineReader

ABBYY FineReader in the AI category. These are the open alternatives we recommend looking at.

Tesseract

Established OCR engine that recognises text in images, with support for over 100 languages.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

GitHub stars
76.8k
Deployment
Library

OCRmyPDF

Adds a searchable text layer to scanned PDF files using OCR.

MPL-2.0Weak copyleft

SaaS: YesClosed product: Yes, with conditions

GitHub stars
34.9k
Deployment
CLI

MarkItDown

Python tool from Microsoft that converts Office documents, PDFs and other files to Markdown.

MITPermissive

SaaS: YesClosed product: Yes

GitHub stars
188k
Deployment
Library

PaddleOCR

OCR toolkit that reads text and structure from images and PDF files, with support for many languages.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

GitHub stars
90.5k
Deployment
Library

Docling

Converts PDFs, Office files and images into structured text and tables that AI models can use.

MITPermissive

SaaS: YesClosed product: Yes

GitHub stars
68.3k
Deployment
Library

Marker

Converts PDF files to Markdown and JSON, including tables and equations.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

GitHub stars
40.2k
Deployment
Library

olmOCR

Toolkit that turns PDF pages into clean, readable text using a language model.

Apache-2.0PermissiveLow activity

SaaS: YesClosed product: Yes

GitHub stars
19.7k
Deployment
Library

Unstructured

Library that turns documents in many formats into structured data for language models and search.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

GitHub stars
15.5k
Deployment
Library

DeepSeek-OCR

Model from DeepSeek that reads text and layout in document images and compresses them for language models.

MITPermissive

SaaS: YesClosed product: Yes

Deployment
transformers

Florence-2-large

Vision model from Microsoft for captioning, object detection and OCR through text prompts.

MITPermissive

SaaS: YesClosed product: Yes

Deployment
transformers

PaddleOCR-VL

Compact model from PaddleOCR that parses documents with text, tables and formulas in many languages. Based on baidu/ERNIE-4.5-0.3B-Paddle.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

Deployment
PaddleOCR

dots.ocr

Model that parses document layout and reads text, tables and formulas in one step, in several languages.

MITPermissive

SaaS: YesClosed product: Yes

Deployment
dots_ocr

granite-docling-258M

Small model from IBM that turns document pages into structured text, for use with Docling.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

Deployment
transformers

olmOCR-7B-0225-preview

Model from the Allen Institute for AI that turns PDF pages into clean text. Based on Qwen/Qwen2-VL-7B-Instruct.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

Deployment
transformers

GOT-OCR-2.0-hf

OCR model from StepFun that reads text, tables, formulas and sheet music from images.

Apache-2.0Permissive

SaaS: YesClosed product: Yes

Deployment
transformers

This is guidance, not legal advice (inte juridisk rådgivning). Check with a lawyer before deciding.

Sign in for more

  • Full licence analysis per use case (self-hosting, modifying, hosted service, AI use)
  • Save products to your own lists
  • Systemkarta: map the tools you already use (coming soon)
  • Personal help and AI advice (coming soon)
Sign in for free