Skip to content

optical character recognition (OCR)

Optical character recognition (OCR) is the conversion of images of text, such as scanned pages or photographs, into machine-readable character data that software can search, edit, and feed into natural language processing pipelines.

A classical OCR pipeline first preprocesses the image by straightening any tilt, a step called deskewing, and removing noise. It then detects and segments text regions into lines and words, recognizes the characters in each segment, and applies dictionary or language-model correction to the raw output. The stepper below runs those four stages on a sample scan:

Interactive diagram — enable JavaScript to view.

Early engines matched glyphs against stored templates or hand-built features, while modern ones learn recognition from data using convolutional and recurrent networks. Tesseract, an open source engine built around a long short-term memory recognizer, covers over 100 languages and 35 scripts.

Accuracy depends heavily on input quality, and handwriting, unusual fonts, and low-resolution or skewed scans remain hard cases. Multimodal large language models, which take images as input alongside text, now read a page directly and return structured output that preserves tables, formulas, and reading order.

Image Processing With the Python Pillow Library

Tutorial

Image Processing With the Python Pillow Library

In this step-by-step tutorial, you'll learn how to use the Python Pillow library to deal with images and perform image processing. You'll also explore using NumPy for further processing, including to create animations.

intermediate data-science

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Martin Breuss • Updated Sept. 14, 2026