Why text copied from a PDF comes out garbled
The page looks perfect and the paste is nonsense, because what you see and what the file knows about its own text are two different things.
The short answer
Text copied from a PDF comes out garbled because a PDF stores glyphs at positions, not words in order. When the font has no map from its glyphs back to real characters, you get symbols or shifted letters. When the map is fine, spacing and reading order still have to be guessed from geometry, and the guess can fail.
Those are two different problems with different fixes, and a ten-second test tells you which one you have.
The ten-second test
Paste the copied text into a plain text editor and look at the individual letters, ignoring spacing and order for a moment.
If the letters themselves are wrong, such as boxes, symbols, accented characters where none belong, or words where every letter is consistently replaced by another, the problem is encoding. The file does not know which characters its glyphs represent, and no amount of clever extraction will fix that from the existing text.
If the letters are right but the words run together, break mid-sentence, or arrive in the wrong order, the problem is layout. The characters are correct and only the structure was lost, which is far more fixable.
Encoding: missing ToUnicode maps and ligatures
A PDF font does not have to store text as Unicode. It stores glyph codes, numbers that pick shapes out of the embedded font. To turn those codes back into characters, the font needs a lookup table called a ToUnicode map. Most modern software writes one. Some does not, especially when it embeds a subset of a font containing only the glyphs actually used.
Subsetting often renumbers glyphs in the order they first appear, so code 1 might be the letter T and code 2 the letter h. Without a ToUnicode map, a viewer that copies the codes as if they were standard character codes produces a consistent substitution: every e becomes the same wrong symbol throughout the document. That uniform wrongness is the signature of this cause.
Ligatures are a smaller version of the same problem. Typesetting software often replaces fi, fl, ff, ffi and ffl with single combined glyphs. If the ToUnicode map is missing an entry for the combined glyph, copying drops it, and efficient becomes ecient and office becomes oce. If the entry exists, some viewers paste the single ligature character, which looks right but fails a search for the plain letters.
Layout: fake spaces, tab stops and columns
Even with perfect encoding, a PDF has no concept of a space between words, a line, or a paragraph. It says: draw these glyphs starting here. Everything else is inference, and the common failures are predictable.
- Missing spaces. Many writers position each word separately and never emit a space character. The extractor has to decide whether a gap is wide enough to be a word break. Tight justification fools it, and words run together.
- Extra spaces. Letter-spaced headings are drawn one glyph at a time with gaps between them, so T I T L E comes out with a space between every letter.
- Tab stops. A job title on the left and a location aligned on the right sit on one baseline with a large gap between them. PDF.js, the engine World of PDF uses and the one built into Firefox, reports that gap as a single space character whose width is the entire distance, sometimes hundreds of points, so naive copying glues the two together as if they were one phrase.
- Columns. Selecting across a two-column page reads straight along each baseline, so the paste alternates between a line from the left column and a line from the right.
- Hyphens and furniture. Words split at line ends keep their hyphens, and running headers, footers and page numbers land in the middle of sentences.
Fixing layout problems
Layout problems are fixed by extracting deliberately rather than through a selection rectangle. World of PDF's Extract Text groups characters into lines by vertical position and treats any gap wider than a multiple of the font size, including those wide single-space items at tab stops, as a break between separate runs rather than a word space. It looks for a consistent vertical gutter before splitting a page into columns, so a centred heading does not trigger a false split.
It also joins lines into paragraphs where the spacing suggests continuation, rejoins words hyphenated across a line break, and drops lines that repeat at the same position near the top or bottom of most pages, which is how running headers and page numbers are recognised. The result is still inference, so check the reading order on anything with an unusual layout.
When OCR is the fix
Encoding problems cannot be extracted away, because the information is not in the file. The fix is to ignore the broken text layer and read the page as an image, which is what OCR does. Recognition looks at the glyph shapes, not the codes behind them, so it produces correct characters from a page whose own text layer is gibberish.
There is a catch. World of PDF's OCR tool skips pages that already have a text layer by default, to avoid duplicating text. You can untick that, but the garbled layer stays underneath, and copying would then return both. For a clean result, turn the pages into images first, for example by converting to JPG and back to PDF, then run OCR on that. The OCR tool currently recognises English only.
Text drawn as vector outlines, common in logos and some exported designs, has no text layer at all and needs the same OCR route.
Frequently asked questions
Why do I get weird symbols when I copy text from a PDF?
The font in the PDF is missing a ToUnicode map, so the viewer cannot work out which real characters its glyphs represent. OCR, which reads the shapes rather than the codes, is the reliable fix.
Why are letters missing when I copy from a PDF?
Usually ligatures. Pairs like fi and fl are drawn as one combined glyph, and if the file does not map that glyph back to its letters, copying drops it.
Why do words run together when I paste PDF text?
The PDF positions words without storing space characters, and the extractor misjudged a narrow gap. Tightly justified text is the usual culprit.
Why does copied text from a two-column PDF come out mixed up?
Selection reads straight across the page, alternating between columns. A text extractor that detects the column gutter reads each column in turn instead.
Can OCR fix a PDF with a broken text layer?
Yes, because it ignores the existing text and reads the visible shapes. Remove the broken layer first by turning the pages into images, or the old garbled text stays in the file alongside the new layer.