What Merge, Split and Rotate Actually Do to a PDF

Most people picture a PDF as a stack of pages, the way a paper document is a stack of sheets. Under that picture, merging two files means putting one pile on top of another, and rotating a page means turning a sheet around.

The reality is different in ways that explain several otherwise baffling behaviours — why merging two 2 MB files sometimes gives you 4 MB and sometimes 7 MB, why rotating a page is instantaneous regardless of what is on it, and why extracting one page from a large document can produce a file almost as large as the original.

A PDF is an object graph

A PDF file is a collection of numbered objects. Objects are dictionaries, streams, arrays, numbers, names. They reference each other by number.

A page is an object. It is mostly a dictionary saying: my dimensions are this, my content stream is object 47, the fonts and images I use are objects 12, 13 and 31, my rotation is 0.

The pages are tied together by a page tree — a structure whose leaves are page objects, in order. And a cross-reference table at the end of the file records the byte offset of every object, so a reader can jump straight to object 47 without scanning the whole file.

Nothing in a page contains its own resources. It points at them. That single fact explains everything below.

Rotation is one number

/Rotate is a key in the page dictionary. Its value is 0, 90, 180 or 270.

Rotating a page means writing a different number into that key. The content stream is untouched. The images are untouched. Nothing is re-rendered, nothing is re-encoded, no pixels move.

This is why rotating a 300-page scanned document is instant while rotating a single image in a photo editor takes a moment: the PDF is not rotating anything, it is leaving an instruction for the reader to rotate it at display time.

When we tested our toolkit’s rotation, the check was exactly this — after rotating page 2 of a five-page file, page 2 carries /Rotate 90 and pages 1, 3, 4 and 5 carry nothing. That is the whole operation.

Reordering is a list

Page order lives in the page tree’s array of references. Reordering is rearranging that array.

Again, nothing moves. Object 47 stays where it is in the file; it is simply referenced earlier or later. Reversing a five-page document is five swaps in a list, and the file size does not change at all.

Our test for this asked for 5-1 and checked that the resulting pages came back in exact reverse order — a cheap test that catches the off-by-one errors this kind of index manipulation invites.

Extraction is a subgraph, and this is why it disappoints

Extract pages 2–3 from a fifty-page document and you might reasonably expect a file about 4% of the original size.

You often get much more, because the extracted pages carry their dependencies. If those two pages use a font, the font comes with them — and if the font was embedded unsubsetted, that might be several megabytes. If they contain a full-page image, that image comes too.

Meanwhile the forty-eight pages you dropped were mostly text, contributing very little. You removed the cheap pages and kept an expensive one.

This is not a flaw in the extraction. It is the object graph doing exactly what it should: a page without its resources would not render.

Merging is where size actually goes wrong

To merge, you renumber every object in the second document so its numbers do not collide with the first, append them, splice the page trees together and write a new cross-reference table.

Straightforward — and it is why merged files bloat.

Both documents embed Helvetica. After merging, the file contains two copies of Helvetica, because they were separate objects with separate numbers and nothing compared them. Both had the same company logo on every page: two copies. Merge six departmental reports built from the same template and you have six copies of every shared resource.

Deduplicating properly means comparing object contents across documents, deciding when two streams are genuinely identical, and rewriting every reference to the survivors. It is real work, and most simple merge tools — ours included — do not attempt it.

The honest guidance: if a merged file is unexpectedly large, this is usually why, and the fix is upstream. Merge from sources built on the same template, or re-export the finished document from a producer that will rewrite it cleanly. See what makes PDFs huge for the other three causes.

Page numbers and watermarks are new content

These two are different in kind from the operations above. They do not rearrange existing objects — they add drawing instructions to a page’s content stream, or add a new stream that renders on top.

Which makes them the only operations in the list that genuinely modify a page’s appearance, and the only ones with a bug surface worth mentioning.

Ours had one. The watermark was sized by a crude character-count formula and came out roughly 1.7× too wide, starting at x = −21 with the final letter clipped off the page — “CONFIDENTIAL” extracted from the finished file as “CONFIDENTIA”. It is now sized from the longest line that actually fits inside the page at 35°, capped so a short word does not become a monster, and centred on the page along its own baseline.

The detail worth taking from that: the bug was found by extracting the text back out of the output and comparing it to what went in. A visual check on a preview would have shown a large diagonal word and looked fine.

Why none of this needs a server

Every operation above is structural. Parsing an object graph, editing a dictionary value, rearranging an array, renumbering objects and writing a new cross-reference table — none of it requires decoding an image or rendering a page.

That is why our PDF toolkit runs entirely in your browser with nothing uploaded. It is also why the things we left out — real compression, OCR, PDF-to-Word — are exactly the operations that do require decoding and re-rendering everything.

The structure is cheap. The pixels are expensive.

Two structures worth knowing about

Linearisation. A PDF can be arranged so a reader can display page one before the whole file has arrived, with a special structure at the front and the cross-reference information duplicated. This is what “fast web view” means in export dialogues.

Most editing operations break it. Merge two linearised PDFs and the result is not linearised, because the careful ordering no longer holds. Nothing is damaged and nothing is lost — the file simply loads all at once rather than progressively. For a document served over the web at any size, re-linearising after editing is worth doing; for a file that will be downloaded and opened locally it makes no difference at all.

Encryption and permissions. PDF supports two distinct things that get conflated. A user password encrypts the content — without it, the file cannot be read. An owner password sets permission flags: no printing, no copying, no editing.

The flags are advisory. They are stored in the document and honoured by well-behaved readers as a matter of convention, and a reader that chooses to ignore them can. A PDF that opens without a password but claims to forbid copying is asking politely, not enforcing anything, and any structural operation on the file can drop those flags entirely.

This matters when someone sends you a “protected” document and expects the protection to survive being merged into a larger one. It will not, and it was never as strong as its name suggested.

Why this shapes what a tool can honestly offer

Everything in this article divides cleanly into two categories, and that division explains the whole shape of our toolkit.

Structural operations — merge, split, extract, rotate, reorder, page numbers, watermarks — manipulate the object graph. They are fast, they are lossless, and they need nothing but a parser. They run comfortably in a browser tab on a document of any size.

Content operations — compression, OCR, conversion to an editable format — require decoding every image and rendering every page. They are slow, they are lossy, and they need substantially more than a parser.

Our toolkit does all of the first category and none of the second, and the line between those lists is not arbitrary. It is exactly where a browser stops being the right place to do the work.


Merge, split, extract, rotate, reorder, number and watermark PDFs in your browser with the PDF toolkit.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top