PNG and WebP Carry Metadata Too

Search for photo metadata and you will find a thousand articles about JPEG. Almost all of them stop there, which leaves a gap, because the images most people handle daily — screenshots, exported graphics, images saved from the web — are frequently not JPEGs at all.

PNG and WebP both carry metadata. They store it in their own structures, which means a tool built only for JPEG will find nothing in them and, worse, may report them as clean.

PNG: a sequence of named chunks

A PNG file is a signature followed by a series of chunks. Each chunk has a length, a four-character type, its data, and a checksum. The image data lives in IDAT chunks; everything else lives alongside it.

The ones that carry information about you:

  • eXIf — a full EXIF block, the same structure a JPEG carries, including GPS and camera fields. Standardised relatively recently, and increasingly common in files produced by cameras and editors that export PNG.
  • tEXt — uncompressed key-value text. Commonly holds Software, Author, Comment, Description, Copyright.
  • iTXt — the international version, UTF-8, optionally compressed. This is where XMP usually lives when a PNG carries it, and XMP can contain a full edit history.
  • zTXt — compressed text, same purpose as tEXt. Being compressed, it is invisible to anyone grepping the file for readable strings, which is a reason not to rely on that as a check.
  • tIME — last modification time.

Screenshots are the interesting case. Many screenshot tools write a tEXt or iTXt chunk naming the software, and some write the window title or the source application. That is rarely catastrophic, but it does mean a screenshot can quietly disclose which tool, and sometimes which application, it came from.

The good news: PNG is losslessly compressed, so removing chunks is unambiguously safe. You drop the ones you do not want, fix the sequence, and every pixel comes back exactly as it went in. There is no quality trade-off to weigh.

WebP: a RIFF container with a trap

WebP is built on RIFF — the same container idea as WAV audio. A file is a RIFF header, a WEBP marker, and then a series of chunks.

Metadata lives in two of them:

  • EXIF — a standard EXIF block.
  • XMP — note the trailing space. RIFF chunk identifiers are exactly four bytes, and XMP is three characters, so the standard pads it. Code that looks for XMP without the space will not find it.

And then the part that catches people out.

An extended WebP file begins with a VP8X chunk, and inside it is a flags byte. That byte advertises which optional features the file uses — one bit for an ICC profile, one for alpha, one for EXIF, one for XMP, one for animation.

Delete the EXIF chunk and leave that byte alone, and you have produced a file that announces it contains EXIF metadata which is not there. Lenient decoders shrug. Strict ones go looking for a chunk that does not exist, and some of those will complain or fail outright.

Our tool patches the VP8X flags to clear the EXIF, XMP and ICC bits whenever the corresponding chunks are removed. It is a small thing — one byte — and it is the difference between a valid file and one that merely works in most viewers.

If you strip WebP metadata with something and the resulting file behaves oddly in a particular application, this is the first thing to suspect.

What this means for the files you actually handle

Screenshots. Usually PNG. Frequently carry a software name. Check before publishing anything from a work machine.

Images exported from design tools. PNG or WebP, and among the most likely to carry XMP with an edit history, an author name and sometimes a full local file path — which reveals a username and directory structure.

Images saved from websites. Increasingly WebP. Carry whatever the site’s pipeline left in them.

Photographs exported as PNG. People do this believing PNG is “clean” because it is lossless. Lossless refers to the pixels. The eXIf chunk comes along regardless, GPS and all.

That last one is the most common misconception in this whole area. Converting a JPEG to PNG does not remove its metadata. It usually preserves it, in a different container, in a much larger file.

Checking and stripping

Our EXIF tool reads all three formats — JPEG via the FFE1 APP1 segment, PNG via eXIf, WebP via its EXIF chunk — with a parser written from scratch rather than borrowed, and it flags XMP, IPTC and comment segments separately from EXIF, because “we removed the metadata” is a claim that needs to say which metadata.

Colour profiles are the deliberate exception. An ICC profile describes how the numbers in the file map to real colours; removing it does not make an image anonymous, it makes it wrong. It also says nothing about you. It stays by default, with the option to remove it if you have a reason.

The general lesson

Format guides tend to discuss compression and transparency and stop. Every container format has somewhere to put text, and somewhere to put text is somewhere information about you accumulates — usually written by software you did not choose, recording things you did not decide to disclose.

If you take one habit from this: check the file you are about to share, whatever its extension. Not the file you started with, not the file your editor shows you — the one that is about to leave.

HEIC and AVIF, briefly

Two newer formats are worth a mention because they are increasingly what phones and websites actually produce.

HEIC is the default photo format on recent iPhones. It is a container built on the same family of structures as modern video files, and it carries full EXIF along with everything else discussed here. Converting a HEIC to JPEG for sharing typically preserves the metadata rather than removing it.

AVIF is the newer web format, built on the AV1 video codec, and likewise carries EXIF and XMP in its container.

The general rule holds for both: a modern image container has somewhere to put metadata, and something in the pipeline will use it.

How to look for yourself

You do not need a tool to confirm a file carries text-bearing chunks — though you do need one to read them properly.

Open any PNG in a hex viewer and the four-character chunk names are visible as readable text near the start and end of the file: IHDR, then whatever ancillary chunks are present, then IDAT, then IEND. Spotting tEXt or eXIf in that list tells you something is there.

The limit of that approach is zTXt and compressed iTXt, which are exactly the chunks a casual inspection misses, because their contents are deflate-compressed and will not appear as readable strings. This is worth knowing if you have ever satisfied yourself that a file was clean by searching it for your own name.

Proper parsing means walking the chunk sequence, reading each length prefix, and decompressing where required — which is what a real tool does and what a text search cannot.

The same caution applies to WebP. A search for Exif will miss a file whose metadata is stored in an XMP chunk, and neither search tells you anything about the VP8X flags byte.


Read and strip metadata from JPEG, PNG and WebP in your browser with the EXIF viewer and remover.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top