Skip to content
WorldofPDFs

How a PDF can still contain what you deleted from it

Saving a PDF often adds to the file rather than replacing it. The page you see is the latest version; the earlier ones may be sitting right behind it.

The short answer

An incremental save writes a PDF's changes to the end of the file and leaves everything before them untouched. The viewer shows the newest version, but the old objects — deleted text, replaced images, previous metadata, removed pages — are still in the bytes and can be recovered. Only a full rewrite, which rebuilds the file from its current state, drops them.

This is a feature of the format, not a bug in any one editor, and it is why "I deleted it before sending" is not always true.

How a PDF is put together

A PDF is a sequence of numbered objects — pages, content streams, fonts, images, dictionaries — followed by a cross-reference table that records the byte offset of each object, and a trailer that says where the table is and which object is the document's root. A reader starts at the end of the file, finds the table, and uses it to jump to whatever it needs.

That design makes a cheap kind of update possible. To change a page, a writer does not have to touch the existing bytes. It appends the new or modified objects, then a new cross-reference section listing only those objects, then a new trailer pointing back to the previous section. Readers follow the chain from the newest section backwards, and for any object listed in several sections the newest entry wins.

Why deleted content survives

Because the old objects are never removed, only superseded, a file that has been edited and saved this way is a stack of revisions.

  • Edited text. Changing a sentence typically writes a new content stream for the page. The old stream, with the original sentence, stays where it was.
  • Deleted pages. Removing a page updates the page tree so the page is no longer listed. The page object and everything it drew are still in the file, just unreferenced.
  • Replaced images. Swapping a photo appends the new one; the old one is still embedded.
  • Changed metadata. Rewriting the author or title appends a new information dictionary, leaving the previous one, with the old name, behind it.
  • Cosmetic redaction. Drawing a black box and saving appends the box. Even a proper redaction applied by an editor can leave the original page content in an earlier revision if the file is saved incrementally rather than rewritten.

Getting the old version back

Recovery is not specialist work. Every incremental save ends with its own end-of-file marker, so cutting the file just after an earlier marker produces a valid PDF of the document as it stood at that save. Short of that, extracting every string and stream from the raw bytes will turn up text that no longer appears on any page.

There have been public cases of court filings and official reports where redacted or deleted passages were recovered this way. The pattern is always the same: the change was made, the document looked right, and the file was saved by appending rather than rewriting.

The same structure is also useful. Digital signatures depend on it: a signature covers the exact bytes up to that point, and later changes are appended so the signed revision stays intact and verifiable. Spotting multiple revisions is one of the more reliable signs that a file has been edited after it was created.

What a full rewrite removes, and what it might not

A full rewrite serialises the document's current state as a fresh file with one cross-reference table. Superseded versions of an object are gone, because only the latest version of each is written. That alone removes most of what incremental saving leaves behind.

It does not necessarily remove objects that are no longer referenced at all. Some writers, including the widely used pdf-lib library that the tools on this site are built on, write out every object they parsed, referenced or not. A page deleted in an earlier revision can survive such a re-save as an orphan nobody points to. Invisible in every viewer, still present in the bytes.

So there are really three levels of clean, and it is worth knowing which one a tool gives you.

  • Re-save in place. Every World of PDF tool writes a complete new file rather than appending, so revision history is collapsed. Tools that modify the existing document, such as Compress PDF's image re-encoding presets, stop there: superseded versions are gone, orphans may not be.
  • Rebuild from pages. Extract Pages copies the pages you choose into a brand-new document, carrying only what those pages actually reference. Earlier revisions, orphans and the original metadata dictionary are left behind. Extracting every page is a simple way to get a structurally clean copy.
  • Rebuild from pixels. Redact PDF renders every page to an image and builds a new document from those images. Nothing from the original file survives except what is visible, which is why it is the right level for anything that must not be recoverable.

Practical rules

Before sending an edited document outside your organisation, assume it carries its history unless you know it was rewritten. Producing the final copy through a tool that rebuilds the file — export to a fresh PDF from the source document, or rebuild it from its pages — is the habit that makes this a non-issue.

Do not rewrite a signed document you need to stay valid. A full rewrite changes the byte ranges the signature covers, and the signature will no longer verify. Sign last, after every other change.

And for redaction specifically, use a method that removes content rather than one that hides it, then make sure the saved file is a rebuild. Either half on its own can leave the text recoverable.

Frequently asked questions

What is an incremental save in a PDF?

Saving changes by appending new objects and a new cross-reference section to the end of the file, instead of rewriting it. It is fast and keeps earlier bytes intact, which also means earlier versions of the content remain in the file.

Can deleted text in a PDF be recovered?

Often, if the file was saved incrementally. The old content stream is still in the file, and truncating at an earlier end-of-file marker can restore the previous version outright.

How do I remove old versions from a PDF?

Rewrite the whole file rather than appending to it. Rebuilding it from its pages, or rendering it to images for anything sensitive, also drops objects nothing refers to any more.

Does compressing a PDF remove its edit history?

A compressor that writes a complete new file collapses the revisions, so superseded versions are dropped. It may still keep objects nothing refers to, so it is not a substitute for a proper rebuild.

Will rewriting a PDF break its digital signature?

Yes. A signature covers specific bytes, and a full rewrite changes them. Make every other change first and sign last.

Tools mentioned in this guide