How to tell whether a PDF is searchable or just a scanned image
A searchable PDF and a scanned one can look identical on screen, and the difference only shows up when you try to find, copy or upload the text.
The short answer
To tell if a PDF is searchable, press Ctrl+F (Cmd+F on a Mac) and search for a word you can see on the page. If it is found, the PDF has a text layer. If not, try selecting a line of text with your cursor. If neither works, the page is an image with no text behind it and needs OCR before it can be searched.
Those two checks settle most files. A third, extracting the text, catches the awkward cases where the text layer exists but is incomplete or wrong.
Check one: search for a word you can see
Pick a distinctive word from the middle of a paragraph, not a heading. Headings are sometimes real text on an otherwise scanned page, added by whoever assembled the file, and searching for one can give a false positive.
Search for it. A match that highlights the word on the page means the characters exist in the file as text. No match means one of three things: the page is an image, the text layer is garbled, or the word is split by a hyphen or unusual spacing. Try a second word before concluding anything.
Do this on more than one page. Merged documents frequently mix born-digital pages with scanned ones, and a file can be searchable on page 1 and blind from page 4 onwards.
Check two: try to select a line
Switch to the text-selection cursor and drag across a line. On a searchable page, the highlight follows individual words and stops at the end of the line. On an image-only page, one of two things happens: nothing highlights at all, or the whole page turns blue as a single object, because the only thing to select is the picture.
Paste what you selected into a plain text editor. Correct words confirm a working text layer. Symbols or nonsense mean a text layer exists but its encoding is broken, which behaves like an unsearchable file for practical purposes.
Check three: extract the text
Run the file through a text extractor and skim the output. This is the most honest check, because it shows exactly what software sees: what a search index, a screen reader or an AI tool would receive.
Empty or nearly empty output means image-only pages. Output that is mostly right but missing whole blocks means a partial text layer, often from OCR that skipped a table or a faint section. Output full of wrong characters means the text layer exists but is unreliable.
World of PDF's Extract Text shows the result in a preview pane before you save anything, which makes this a quick look rather than a download. Its OCR tool also reports, as soon as you add a file, how many pages have no text layer, which is the per-page answer for a long document.
Scanned, OCR'd and born-digital
A PDF is not simply searchable or not. There are three kinds, and knowing which one you have explains its behaviour.
- Born-digital. Exported from a word processor or other software, with real text drawn by fonts. Fully searchable, and text stays sharp at any zoom.
- Scanned, image-only. Each page is a photograph of paper. It looks like text and contains none. Search, selection and extraction all return nothing.
- Scanned with OCR. The same photograph, with an invisible layer of recognised words placed on top. It passes the search and selection checks, but the text is only as accurate as the recognition, and selection highlights may sit slightly off the visible words.
Telling a scan with OCR from a real document
The zoom test separates them. Zoom to 400 percent or more on a line of text. Born-digital text stays perfectly crisp, because it is drawn from font outlines at whatever size you ask for. Scanned text, OCR'd or not, goes soft and pixelated, because it is an image with a fixed resolution.
That matters because OCR text can be wrong in ways born-digital text never is: a 0 read as an O, a 1 as an l, a number in a table attached to the wrong cell. If a document is a scan with OCR, search will work, but do not trust it to find every occurrence, and proofread anything you copy out of it. How recognition works, and where it fails, is covered in the guide to OCR.
If the checks show an image-only file, running OCR adds the missing text layer without changing how the page looks. Pages that already have text can be left alone, which avoids doubling up the words on mixed documents.
Frequently asked questions
What is the difference between a searchable PDF and a scanned PDF?
A searchable PDF contains text as characters the software can read. A scanned PDF contains only images of pages, so it looks like text but cannot be searched or copied until OCR adds a text layer.
Why does Ctrl+F not find words in my PDF?
Usually because the pages are scanned images with no text layer. It can also mean the text layer has broken encoding, or that the word is hyphenated or spaced unusually in the file.
How can I make a scanned PDF searchable?
Run it through OCR, which recognises the words in the page images and adds them as an invisible text layer. The pages look the same but become searchable and selectable.
Can a PDF be partly searchable?
Yes, and it is common. Merged files often mix born-digital and scanned pages, so check several pages rather than just the first.
Is an OCR'd PDF the same as a born-digital one?
No. It is searchable, but the text is a recognition guess layered over an image. Zoom in: born-digital text stays sharp, scanned text pixelates.