Can redacted text be recovered from a PDF?
Whether a redaction holds depends entirely on whether the text was removed or merely covered, and from the outside the two look identical.
The short answer
Redacted text can be recovered whenever the redaction only covered it. A black rectangle drawn over a PDF is usually a new shape painted on top, and the original text stays in the file underneath, ready to be selected, copied or extracted. Text that was genuinely removed from the page's content cannot be recovered, because it is no longer there.
The difficulty is that both kinds look exactly the same on screen. The only way to know which one you have is to test it.
How black boxes fail
This is not a hypothetical weakness. Court filings have been published where the black bars over names and figures were drawn shapes, and anyone who selected the blacked-out area and pasted it into a text editor got the full original wording. Government reports have gone out the same way. In most cases nobody broke anything; they simply copied and pasted.
The cause is how a PDF page is built. Each page has a content stream, a list of drawing instructions read top to bottom: set this font, show these characters at these coordinates, fill this rectangle in black. A viewer paints them in order, so a rectangle drawn after the text covers it visually. But the instruction that shows the characters is still in the list. Copy, search and text extraction read that instruction directly and never look at what was painted over it.
Most general-purpose PDF editors add to a content stream rather than rewrite it, because adding is simple and rewriting is hard. That is why a highlight tool set to black, a filled shape or a black comment box all fail the same way.
What true removal does to the file
Real redaction has to change the content stream so the characters are not in it any more. There are two ways to get there.
The surgical way edits the instructions themselves: find the text-showing operators that fall inside the marked area and delete or replace those characters, leaving the rest of the page as live text. It is precise and keeps the page searchable, but it is genuinely difficult to do correctly, because a single show-text instruction can span the boundary of a box and fonts can encode characters in ways that make the boundary hard to find.
The robust way rebuilds the page. Render it to an image, paint the marked areas black in the pixels, and replace the page with that image. Nothing from the original drawing instructions survives, so there is nothing to extract. The cost is that the whole page becomes a picture and the unredacted text on it is no longer selectable either.
World of PDF's Redact PDF and Auto-Redact PII both take the robust route. Every page is rendered at 200 DPI, the marked rectangles are filled in the pixels, and a new document is assembled from those page images. The redacted output contains no text layer at all.
The places text hides besides the page
Even a correct edit to the visible page can leave the same words elsewhere in the file. A thorough redaction accounts for all of these.
- Metadata. The title, author and subject fields in the document properties often repeat a case name, a client or the person who wrote the file.
- Earlier revisions. PDFs can be saved incrementally, with changes appended to the end of the file and the original bytes left in place. A redaction saved this way can leave the unredacted version sitting earlier in the same file.
- Hidden layers. Optional content groups can be switched off in the viewer but still carry text. A layer you cannot see is a layer someone else can turn on.
- Annotations and form fields. Comments, sticky notes and filled form values live outside the page content and are untouched by anything drawn on the page.
- Bookmarks and alt text. Outline titles and image descriptions are stored separately and frequently quote the headings or names being redacted.
- The shape of the box. A redaction box the exact width of a name gives away its length, and in a short list of candidates that can be enough.
Rebuilding versus editing
Rebuilding the document from rendered images handles most of that list as a side effect. A new document assembled from page images carries no earlier revisions, no annotations, no form fields, no bookmarks and no hidden layers, because none of them are part of the rendered picture. A layer switched off at render time is simply not drawn.
Metadata is the one worth checking separately, since document properties are easy to overlook and some tools copy them across deliberately. Run the finished file through a metadata scanner before sending it.
The box-width leak survives every method, because it is a property of what you chose to cover. Where it matters, draw boxes a uniform size rather than tight to each word.
How to test your own redaction
Do this on the output file, the one you are about to send, not on the working copy. It takes two minutes.
- Search for a redacted word. Open the file and press Ctrl+F (Cmd+F on a Mac) for a name or number you covered. Any hit means the text is still there.
- Select across the box. Drag a selection over the black area, copy and paste into a plain text editor. Pasted text means the redaction is cosmetic.
- Extract all the text. Run the whole file through a text extractor and search the output. This catches text in white, in tiny type or under a box you missed.
- Check the properties. Look at the title, author and subject fields, or use a metadata scanner, for names repeated outside the page.
- Open it in a second viewer. A browser's built-in viewer and a desktop reader will sometimes disagree about layers and annotations, which is itself a warning sign.
Frequently asked questions
Can you remove a black box from a redacted PDF?
If the box is a shape drawn on top of the page, yes: it can be deleted in many editors, or the text under it simply copied out. If the page was rebuilt with the area painted out, there is nothing beneath the box to reveal.
Does flattening a PDF make redaction permanent?
No. Flattening merges annotations into the page, but the text under a black box stays in the content stream and remains selectable. Redaction has to remove the text, not just fix the box in place.
Can redacted text be recovered from a scanned PDF?
Only if the scan has a text layer from OCR and the redaction did not remove it. A pure image with the area blacked out in the pixels has nothing left to recover.
Why is the rest of my redacted page no longer searchable?
The tool rebuilt the page as an image so that nothing under the boxes survives. Run OCR on the redacted file afterwards if you need the remaining text searchable again.
Is it safe to redact a PDF with an online tool?
Only if it runs in your browser. A server-based redaction tool has to receive the unredacted document first, which defeats much of the purpose.