Why a Black Box Is Not Redaction

Open a PDF that someone has redacted. Find a black bar. Drag your cursor across it as though you were selecting ordinary text, copy, and paste the result into a text editor.

If words appear, the document was never redacted. It was decorated.

This is not an exotic attack. It requires no tools, no expertise and about thirty seconds. It has exposed the names of confidential informants, the terms of sealed settlements, and the internal deliberations of governments. And it keeps happening, because the mental model almost everyone brings to a PDF is wrong in a specific and consequential way.

A PDF is a list of instructions, not a picture

When you look at a page on screen, you see ink on paper. That is not what the file contains.

A PDF is closer to a set of stage directions. It holds a list of drawing operations: place this font at this size, move to these coordinates, paint these characters. Render them in order and a page appears. The page you perceive is the result of the instructions, not the thing being stored.

Now consider what happens when a tool lets you “redact” by drawing a rectangle. It appends one more instruction to the list: paint a filled black box at these coordinates. That instruction runs last, so it covers what came before.

The text instruction is still there. Untouched, in its original position, in the original font, spelling the original words. You have not removed anything. You have drawn over it — and because the text object is still a text object, every PDF reader in the world will still happily select it, copy it, search it and read it aloud.

The same applies to the other shortcuts people reach for. Highlighting text in black changes a colour value; the characters remain. Placing an opaque image over a region adds an image; the characters remain. Setting text to white-on-white changes a fill colour; the characters remain, and turn up the moment anyone selects the page.

The failures are not obscure

In January 2019, lawyers for Paul Manafort filed a court document with sections blacked out. The blackouts were the drawn-rectangle kind. Reporters selected the text underneath and read allegations that Manafort had discussed a Ukraine peace plan with Konstantin Kilimnik and shared 2016 polling data with him — material the filing had been intended to keep sealed. (BBC)

That case is famous because of who it involved, not because it was unusual. Court records, freedom-of-information responses and corporate disclosures are full of the same mistake. It is made by organisations with legal departments, compliance officers and document-management budgets — because the tool showed them a black box, the black box looked convincing, and nothing in the interface suggested the words were still in the file.

The part almost nobody mentions

Here is where it gets genuinely uncomfortable, and where most articles on this subject stop too early.

Suppose you do it properly. You use a tool that actually deletes the text objects rather than covering them. Copy-paste returns nothing. The words are gone from the file. Are you safe?

Not necessarily.

Researchers at the University of Illinois Urbana-Champaign examined this question across eleven popular redaction tools and published the results as Story Beyond the Eye: Glyph Positions Break PDF Text Redaction (Bland, Iyer and Levchenko, 2022). Two of the eleven tools failed outright, leaving the text copy-pastable. That was the least interesting finding.

The important one is that the space where the text used to be still describes the text that was there.

Proportional fonts do not give every character the same width. An “i” is narrow, a “W” is wide. When a line of text is typeset, the position of every character after the redacted portion depends on the exact widths of the characters that were removed. The paper describes recovering what it calls sub-pixel-sized horizontal shifts in the surrounding characters — measuring, in effect, the shape of the hole.

From that hole you can work backwards. The researchers found that PDFs produced by Microsoft Word leaked up to 15.9 bits of information about a redacted surname — enough to identify it correctly 38% of the time from a list of common surnames, and enough in many cases to narrow thousands of candidates down to one.

Applying their tooling to real corpora, they found more than 6,000 redactions in US court documents where the text was still directly copy-pastable, and hundreds more where the text had been properly removed but remained recoverable from the layout alone.

Rasterising the page — flattening it to an image — is the intuitive fix, and it is better. But the paper found that even tools claiming to rasterise still leaked positional information, because the flattening happened after the layout had already been computed around the missing text.

What real redaction requires

Three things, in order, and skipping any one of them undoes the others:

  1. Remove the content, not the view of it. The text objects, the image data, the annotation, the attachment — deleted from the file, not covered.
  2. Destroy the layout evidence. Re-flow or re-typeset the remaining text, or rasterise the affected region at a resolution that does not preserve sub-pixel positioning, so the gap no longer measures the thing that was in it.
  3. Strip everything that is not the page. PDFs carry document metadata, revision history, embedded thumbnails, attached files, form-field values and comment threads. A page can be immaculately redacted while the author name, the original filename and an earlier draft sit in the file’s metadata. This is the same class of problem as metadata in photographs, and it is missed just as often.

For anything genuinely sensitive — legal filings, medical records, anything where exposure has consequences — the honest advice is to use dedicated redaction software from a vendor who will state in writing what their tool removes, and then verify the output yourself with the copy-paste test before it leaves your building.

Why we did not build a redaction tool

We build browser-based file tools. Our PDF toolkit merges, splits, extracts, rotates, reorders, numbers and watermarks — all of it in your browser, with nothing uploaded anywhere.

A redaction feature would have been one of the easiest things on that list to build. Drawing a filled rectangle onto a PDF page is a handful of lines. It would have looked exactly like the redaction tools people already use, and it would have been just as useless.

We decided against it, deliberately, and the reasoning is worth stating plainly: shipping a black-rectangle tool under the word “redact” would not be a feature with a caveat. It would be a confidentiality risk wearing the label of a safeguard. Someone would use it on a document that mattered, trust the black bar the way people trusted it in the Manafort filing, and be worse off than if we had offered nothing at all — because our tool would have supplied the false confidence.

Doing it properly, to the standard the research above demands, is not a weekend feature in a browser tab. It is a serious piece of software with a serious testing burden. We would rather tell you that than sell you a rectangle.

The rest of the toolkit does what it says. Where we cannot do something safely, we would rather say so than pretend.


Check your own documents. Take a PDF you have already sent to someone with sections blacked out. Select the text under a bar and paste it somewhere. It takes thirty seconds, and it is better to find out now.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top