What Actually Makes a PDF Huge

A ten-page text document should be a few hundred kilobytes. When it arrives at forty megabytes, something specific has happened, and it is nearly always one of four things.

Working out which one matters, because the fixes are completely different and most of the advice online is a generic instruction to “compress the PDF” that either does nothing or destroys the document.

1. It is not a document, it is photographs of a document

This is the overwhelming most common cause.

Scan ten pages, or photograph them with a phone, and you do not have ten pages of text. You have ten large images that happen to depict text. A phone camera produces perhaps 3–4 MB per shot. Ten of those, wrapped in a PDF container, is a 35 MB file containing no text at all — you cannot select it, search it, or copy from it.

How to tell: open the PDF and try to select a line of text. If your cursor draws a rectangle instead of highlighting words, it is images.

What actually helps: reduce the images before they become a PDF, not after. A page of black text on white paper does not need a 12-megapixel photograph. Downscale to something around 1,500–2,000 pixels on the long edge and the text stays perfectly legible while the file collapses. Our image compressor and images-to-PDF tools do this pair in sequence, in your browser.

What does not help: running the finished PDF through a “compressor”. By then the images are already embedded and already compressed once; squeezing them again mostly costs quality.

2. Embedded fonts

A PDF that is guaranteed to look identical everywhere has to carry its fonts, because it cannot assume the reader’s machine has them.

Done well, this is cheap: a subset embeds only the glyphs actually used, and costs tens of kilobytes. Done badly, the producing application embeds complete font files — and a single large CJK or heavily-featured OpenType font can run to several megabytes on its own. A document using six weights of two families, unsubsetted, is carrying a typography library.

How to tell: document properties in most readers lists the fonts and whether each is embedded and subsetted.

What helps: re-export from the source document with font subsetting enabled. This is a setting in whatever produced the file, not something a downstream tool can fix well.

3. Everything that was ever in it

PDF supports incremental saving. Edit a file, save, edit again, save again, and rather than rewriting the document a writer can append the changes to the end, leaving the previous version intact underneath.

Do that thirty times and the file contains thirty versions. You see the newest. The file carries them all.

This is a size problem and, more seriously, a disclosure problem — earlier drafts, deleted paragraphs and removed images can all still be present in a file that displays none of them. It is the same family of mistake as drawing a black box over text and calling it redaction: what the page shows and what the file contains are different questions.

What helps: “Save As” rather than “Save” in the producing application, or any export that rewrites the file from scratch rather than appending to it.

4. The same thing, many times

Images and fonts in a PDF are objects that pages reference. A logo on every page of a fifty-page document should be one object referenced fifty times.

Producers that build PDFs by concatenation — assembling from separately generated pages, or merging many files — frequently embed a fresh copy for every occurrence. Fifty copies of a 400 KB logo is 20 MB of one picture.

How to tell: a file whose size scales suspiciously linearly with page count, when the pages are visually similar, is usually doing this.

What we deliberately do not offer

Our PDF toolkit merges, splits, extracts, rotates, reorders, numbers and watermarks. It does not offer PDF compression, and that omission is deliberate rather than an oversight.

Doing it properly means decompressing every embedded image stream, re-encoding each at a sensible resolution and quality, deduplicating repeated objects, subsetting fonts, and rewriting the object graph. That is a serious piece of software. Shipping something that renamed the file and shaved two per cent off it would be dishonest, and shipping something that silently degraded every image would be worse.

Where we cannot do a thing well in a browser, we would rather say so — the same reason we did not build a redaction tool.

The diagnosis, in order

  1. Select some text. Cannot? It is scans. Fix the images, not the PDF.
  2. Check document properties for fonts. Large, unsubsetted entries? Re-export with subsetting.
  3. Compare the page count to the size. A few hundred KB per page for text is normal. Several MB per page means images.
  4. Ask how the file was made. Many rounds of editing, or assembled from many files? Re-export from source rather than patching downstream.

Almost every genuinely enormous PDF is scans. If you fix nothing else, fix the resolution of your scans and you will fix ninety per cent of this problem permanently.

Reducing a scanned document properly

Since scans cause most of the problem, here is the sequence that actually works, in order.

1. Decide the resolution you need. For text you intend to read on screen, 150–200 DPI is ample. For text you intend to run through OCR later, 300 DPI is the usual recommendation. Above that you are storing paper texture.

2. Reduce the images before they become pages. A phone photograph at 4,000 pixels wide becomes a perfectly legible A4 page at around 1,700. Do this with an image compressor on the source images.

3. Convert to greyscale if the document is monochrome. Black text on white paper does not need colour information, and dropping it removes two of the three channels.

4. Assemble. Images to PDF will build the document, matching page orientation to each image.

Done in that order, a 40 MB scan routinely becomes 2–3 MB with no visible difference at reading size. Done in the opposite order — assemble first, compress the PDF afterwards — you get a fraction of the benefit and a worse-looking document, because you are re-compressing images that have already been compressed once.

What “PDF compressors” actually do

Since the advice is everywhere, it is worth knowing what these tools do when they work.

The honest ones decompress each embedded image stream, downsample it to a target resolution, re-encode it at a chosen quality, and rewrite the file. That is genuinely useful and it is genuinely lossy — you are throwing away image data, and the setting that controls how much is often not exposed.

Some also subset fonts, remove unreferenced objects, and drop the accumulated revision history, which is where a surprising share of the reduction sometimes comes from. If a “compressor” shrinks a text-only PDF dramatically, it almost certainly found old revisions rather than doing anything clever with the pages.

The dishonest ones re-save the file with slightly different stream compression, achieve two per cent, and present a progress bar. There is no way to tell which you have without checking the output, so check: compare file sizes, then zoom into a scanned page at 200% and look at the text edges.


Merge, split, extract, rotate, number and watermark PDFs in your browser with the PDF toolkit. Nothing is uploaded.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top