OCR a scanned PDF

Turn scans and photos into searchable, copyable text, entirely on your device.

Runs in this tab.
Runs on your device

How to oCR a scanned PDF

  1. Add the scan

    Drop a scanned PDF or a photo of a page.

  2. Choose the language

    English is bundled and works offline. Other languages fetch only a small language file the first time; your document is never sent.

  3. Download

    Get a searchable PDF (your original pages with an invisible text layer) and a plain .txt of everything recognised.

Frequently asked questions

Does this run in my browser?
Yes. Tesseract (the same open-source OCR engine behind many desktop tools) runs as WebAssembly in a web worker on your device. Expect a few seconds per page.
Which languages are supported?
English plus Spanish, French, German, Italian, Portuguese, Dutch, Polish, Turkish, Russian, Arabic, Chinese (Simplified) and Japanese. The invisible text layer in the searchable PDF uses a standard Latin font, so non-Latin scripts are complete in the .txt but only partially present in the PDF layer.
Is the scan kept?
No. It is processed in memory in this tab and discarded when you leave.
How accurate is it?
Clean, upright scans at 200 DPI or better read well. Handwriting, skew, low resolution and unusual fonts reduce accuracy; check the per-page text before relying on it.

More pdf tools