How to extract text from a scanned PDF (OCR)
Recognise text in scanned PDFs in your browser with OCR in over 20 languages, including Lao, Thai, Chinese and Japanese.
A scanned document is a picture of a page, so copying text from it normally does nothing. OCR (optical character recognition) reads the picture and produces text. This runs in your browser — the pages are never uploaded.
Steps
- Open the PDF in the editor.
- Choose Extract Text.
- Pages that look scanned are detected automatically and flagged. Pick your document's language, then run OCR on them.
- If a page has text you can select but it comes out as nonsense, use Force OCR — that happens with older PDFs whose fonts decode incorrectly.
- Copy the result, or download it as plain text or Markdown.
Getting better results
- Pick the right language before running — it matters more than anything else.
- Straight, high-contrast scans read far better than photographs taken at an angle.
- OCR is slower than reading a normal text layer. A long document takes a while, and the first run downloads the language data for that language.
- Expect to proofread. OCR is very good, not perfect — especially on tables, handwriting and low-resolution scans.