PDF StudioFree in-browser PDF editor

How to extract text from a scanned PDF (OCR)

Recognise text in scanned PDFs in your browser with OCR in over 20 languages, including Lao, Thai, Chinese and Japanese.

A scanned document is a picture of a page, so copying text from it normally does nothing. OCR (optical character recognition) reads the picture and produces text. This runs in your browser — the pages are never uploaded.

Steps

  1. Open the PDF in the editor.
  2. Choose Extract Text.
  3. Pages that look scanned are detected automatically and flagged. Pick your document's language, then run OCR on them.
  4. If a page has text you can select but it comes out as nonsense, use Force OCR — that happens with older PDFs whose fonts decode incorrectly.
  5. Copy the result, or download it as plain text or Markdown.

Getting better results

Open the free PDF editor