7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

OCRmyPDF

Open source

OCRmyPDF is a command-line tool that adds a searchable OCR text layer to scanned PDF files using the Tesseract engine.

Open-source alternative to

ocrmypdf.readthedocs.io
OCRmyPDF homepage screenshot
GitHub stars
35k
Last commit
3 days ago
Repository age
12 years
Version
v17.13.0
Licence
MPL-2.0
Self-hosted
Yes

About OCRmyPDF

OCRmyPDF adds an optical character recognition text layer to scanned PDF files so they can be searched and copied from. It turns an ordinary PDF into a searchable PDF/A file and places the recognized text under the page image so copy and paste works accurately. The tool is written in Python and released under the MPL-2.0 license.

It uses the Tesseract OCR engine, which recognizes more than a hundred languages. The tool keeps the original image resolution, inserts OCR data without disturbing other content where possible, and can optimize images so output files are often smaller than the input. Optional deskewing and image cleaning run before recognition, input and output files are validated, and work is spread across available CPU cores. Processing happens on your own machine, which keeps documents private, and it is available from PyPI and Homebrew.

Key features

  • Adds searchable text layer to scanned PDFs
  • Outputs PDF/A files
  • Tesseract OCR with 100+ language support
  • Optional deskew and image cleanup
  • Image optimization to shrink output size
  • Multi-core processing for large documents

Good fit for

  • →Making scanned document archives searchable
  • →Batch OCR on a server or workstation
Built with
Python
Tags
ocr
pdf
tesseract
document-processing
python
command-line
pdfa
scanning

OCRmyPDF: questions and answers

What is OCRmyPDF used for?
OCRmyPDF is a command-line tool that adds a searchable OCR text layer to scanned PDF files using the Tesseract engine. It is a good fit for making scanned document archives searchable and batch OCR on a server or workstation.
Is OCRmyPDF open source?
Yes. OCRmyPDF is open source under the MPL-2.0 licence. Its source code is on GitHub at ocrmypdf/OCRmyPDF and is written mainly in Python.
Is OCRmyPDF free?
Yes. OCRmyPDF is open source, so the software itself is free to use.
What is OCRmyPDF an alternative to?
OCRmyPDF is an open-source alternative to Adobe Acrobat. Other open-source alternatives to Adobe Acrobat include Stirling PDF, BentoPDF and PDFsam.
Is OCRmyPDF actively maintained?
Yes. The most recent commit to OCRmyPDF was on 29 September 2026, and the latest release is v17.13.0, published on 28 September 2026. The project has 35k stars on GitHub.

Open-source alternatives to OCRmyPDF

See all

SaaS alternatives to OCRmyPDF

See all