StudyKit browser utility

Image to Text (OCR)

Read the words out of a photo, a scan, or a screenshot and turn them into text you can copy, search, and edit.

1. Choose an image

Recognition runs in this tab using a copy of the engine served from this site. Your image is never uploaded.

Choose an image, or drop one here You can also paste an image with Ctrl or Command and V A photo, a scan, or a screenshot

The image you are reading will appear here, so you can check it is the right one.

Choose an image to begin. Nothing is uploaded.

2. Extracted text

The first run takes a moment. The recognition engine is about 16 MB and is downloaded from this site the first time you use it, then kept by your browser for later. Nothing is sent to any third party. Expect the least accurate results from very small text, heavy compression, handwriting, and photographs taken at an angle.

Getting the best result

Flat and straight

Photograph the page flat and square to the camera. Perspective distortion is the single biggest cause of wrong characters.

Big, dark text on light

Engines read dark text on a light background best. White text on black, low contrast, and very small captions are the hard cases.

Always check the numbers

Letters such as O and 0, l and 1, and 5 and S are routinely confused. Proofread anything that matters before you rely on it.

What OCR is doing

Text recognition matches shapes against known letterforms. It is not reading, and it has no idea what your document says, which is exactly why it is fast, private, and wrong in characteristic ways. A letter that is slightly blurred matches several candidates and the engine has to guess. It guesses correctly most of the time, which is what makes the result dangerous: a plausible sentence with one wrong digit in it looks exactly as trustworthy as the right one.

It is good at

  • Text that was generated on a screen and then captured, because those pixels are exact.
  • Flat, straight, front-on pages with dark text on a light background.
  • Common, ordinary fonts at a reasonable size.

It is bad at

  • Handwriting, which it will produce confident nonsense for.
  • Photographs of text, where glare, lens distortion and the camera’s own sharpening blur the letter edges.
  • Text at a shallow angle, and text that is very small on screen even in a large file.
  • Low-contrast print, and coloured text on a coloured background.
  • Tables and multi-column layouts, where the reading order is often scrambled even when the words are right.

Crop closer than feels necessary

Resolution matters far more than the total number of pixels in the picture. A tight, straight crop of one paragraph beats a full page shot every time, because the engine only has to resolve the glyphs you kept. If you can choose what to scan, choose the smallest area that still contains the text.

Where the engine comes from

The recognition engine is roughly 16 MB of JavaScript and WebAssembly, fetched from this site the first time you use the tool and cached by your browser afterwards, which is why the first scan is slower than the ones after it. It runs entirely in your browser. Your image is not uploaded, and no third party is sent a copy of what you scanned.

Use it as a first pass

Recognise the text, then read it. For anything you are going to send to someone else, publish, or rely on, check every number, date, address and name against the original.