7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

Tesseract

Open source

An open-source OCR engine and command-line program that recognizes text in images in more than 100 languages, using a neural-network-based engine.

tesseract-ocr.github.io
Tesseract homepage screenshot
GitHub stars
77k
Last commit
4 days ago
Repository age
12 years
Version
5.5.3
Licence
Apache-2.0
Self-hosted
Yes

About Tesseract

Tesseract is an open-source optical character recognition engine. The package includes the libtesseract library and a command-line program called tesseract, and it is written in C++ under the Apache-2.0 license. It supports UTF-8 text and offers recognition for more than 100 languages.

Version 4 added a neural network engine based on LSTM that focuses on recognizing whole lines of text, while the older engine from version 3, which works by recognizing character patterns, remains available through a legacy mode with suitable training data. It reads image formats such as PNG, JPEG and TIFF and can output plain text, hOCR, PDF, invisible-text PDF, TSV, ALTO and PAGE.

The project does not include a graphical application, so GUIs come from third-party tools listed in its documentation, and the engine can be trained to recognize additional languages. It was originally developed at Hewlett-Packard laboratories. Image quality strongly affects results, so preprocessing often helps. It suits developers building document-digitization, search and data-entry workflows.

Key features

  • Command-line program and C++ library
  • LSTM neural network recognition engine
  • Legacy engine mode for older models
  • Recognizes more than 100 languages
  • Reads PNG, JPEG and TIFF images
  • Outputs text, hOCR, PDF, TSV, ALTO and PAGE
  • Trainable for additional languages

Good fit for

  • →Digitizing scanned documents
  • →Making scanned PDFs searchable
  • →Extracting text from screenshots and photos
Built with
C++
Tags
ocr
text-recognition
cpp
lstm
machine-learning
document-processing
command-line
pdf

Tesseract: questions and answers

What is Tesseract used for?
Tesseract is an open-source OCR engine and command-line program that recognizes text in images in more than 100 languages, using a neural-network-based engine. It is a good fit for digitizing scanned documents, making scanned PDFs searchable, and extracting text from screenshots and photos.
Is Tesseract open source?
Yes. Tesseract is open source under the Apache-2.0 licence. Its source code is on GitHub at tesseract-ocr/tesseract and is written mainly in C++.
Is Tesseract free?
Yes. Tesseract is open source, so the software itself is free to use.
What are some alternatives to Tesseract?
Similar open-source tools in the Developer Tools category include Windows Terminal, Godot and Termux. SaaS products in the same category include Qt, Visual Studio and Warp.
Is Tesseract actively maintained?
Yes. The most recent commit to Tesseract was on 28 September 2026, and the latest release is 5.5.3, published on 24 July 2026. The project has 77k stars on GitHub.

Open-source alternatives to Tesseract

See all

SaaS alternatives to Tesseract

See all