How to Turn a Scanned PDF Into Searchable, Editable Text (OCR)
A scanned PDF looks like a document, but to a computer it is a stack of photos of pages. You can't search it, copy a sentence out of it or paste its numbers into a spreadsheet, and a screen reader has nothing to read aloud. OCR (optical character recognition) fixes that by finding the letters in the picture and turning them into real text. Here is how to tell whether your PDF needs it, how to get an accurate result, and what to do with the text afterwards.
Does your PDF need OCR?
Open the PDF and try to select a line of text, or press Ctrl+F (⌘F on a Mac) and search for a word you can see on the page. If nothing highlights, the page is an image. Other clues: the file came from a scanner or a phone scanning app, the text looks slightly tilted or grainy when you zoom in, or a text extractor such as PDF to Text finds nothing.
This matters beyond convenience. The W3C's accessibility guidance on OCR for scanned PDFs points out that with image-only pages, assistive technologies "cannot read or extract the words; users cannot select, edit, resize, or reflow text." If you publish documents for other people, a text layer is part of making them usable.
What OCR produces: text and a searchable PDF
OCR gives you two useful things. The first is plain text you can copy or save. The second is a searchable PDF: the same scanned pages, with an invisible layer of recognised text laid exactly over the words. The pages look unchanged, but you can now search, select and copy, and screen readers can read them.
Recognition is good on clean, printed pages and never perfect. Expect to check names, numbers and anything in small print. Printed text works far better than handwriting: the Tesseract project's own FAQ says you can try handwriting "but it won't work very well, as Tesseract is designed for printed text." Our tools use Tesseract.js, a WebAssembly version of Tesseract that runs in the browser.
Get a clean scan first
Most OCR errors start with the scan, not the software. The Tesseract documentation on improving quality gives the main points:
- Resolution: "Tesseract works best on images which have a DPI of at least 300 dpi." An A4 page scanned at 300 DPI is 2480 × 3508 pixels; a US Letter page is 2550 × 3300.
- Straight pages: line detection "reduces significantly if a page is too skewed". Put the page square on the scanner glass.
- Even background: the page is turned into black and white internally, and that step struggles when the background is of uneven darkness, such as a shadow across a phone photo.
- Little noise: specks, stains and heavy grain lower accuracy.
If you are photographing pages with a phone, lay each one flat, light it evenly without your phone's shadow, hold the phone parallel to the paper and fill the frame with the page. A rescan at 300 DPI usually does more for accuracy than any setting.
How to OCR a PDF with our tools
- Open PDF OCR and click Browse to choose the PDF. The page tells you how many pages it has, and whether the first page already has real text. If the PDF asks for one, type it in Password.
- Choose the Language of the text. There are 22 to pick from, including English, Arabic, Chinese, French, German, Hindi, Japanese, Russian, Spanish and Urdu. Leave Also English (for mixed text) ticked if English words appear alongside another language.
- Pick an Accuracy. Standard reads each page at about 180 DPI; High reads at about 290 DPI, which helps with small print but is slower. For an A4 page, that is roughly 1488 × 2105 pixels against 2380 × 3368.
- Optionally type Pages (optional), such as
1-3, 5, and keep Skip pages that already have real text ticked so mixed documents aren't read twice. - Click Make searchable. The first time, the reading engine (about 4 MB) and the language data download once; after that it takes a few seconds per page. The PDF itself is never uploaded.
- Check the results: Pages read, Words found and Average confidence. Below about 60% the page warns you that the scan may be blurry, small or in another language.
- Click Download searchable PDF, Download text (.txt) or Copy text.
For photos, screenshots and single scanned pages saved as pictures, use Image to Text instead. Drop up to 20 pictures or paste a screenshot with Ctrl+V (⌘V), choose the language, and click Read text. Tick Also make a searchable PDF to get the pictures back as one searchable PDF.
From searchable PDF to editable text
Once the PDF has a text layer, our other PDF tools can read it:
- Plain text: open the searchable PDF in PDF to Text and click Get text. Join lines into paragraphs turns the line-by-line OCR output into flowing text.
- A Word document: open it in PDF to Word, choose the layout Text only, in reading order (easiest to edit) and click Convert to Word. Pictures aren't copied, so you get the recognised words, not the scanned page.
Proofread before you rely on it. Common slips are 0 and O, 1, l and I, "rn" read as "m", 5 and S, and a comma read as a full stop. Check every amount, date, account number and name against the original. Tables usually come out as lines of text, so a column of figures may need rearranging.
Limits worth knowing
- Protected PDFs: if a PDF has a password or restrictions, the searchable copy is rebuilt from pictures of the pages. If you know the password, remove the protection first to keep the original pages as they are.
- Rotated pages: a page stored with a rotation setting is replaced by a picture of itself, with the text on top.
- Size: the searchable PDF can come out bigger than the scan. If it is too big to email, see how to compress a PDF, and use Light mode: Strong mode saves each page as a picture, which throws the text layer away again.
- Time: a long document takes a while. Test a few pages first with Pages (optional) to check the language and accuracy settings.
Quick checklist
- Can you already select the text? Then you don't need OCR.
- Scan at 300 DPI, straight, with an even background.
- Pick the right language, and tick Also English for mixed text.
- Try High accuracy for small print or low confidence.
- Proofread numbers, names and dates against the original.
- Keep the original scan alongside the searchable copy.
Sources
Spotted a mistake or something out of date? Tell us and we'll fix it.