Printer and scanner marks that metadata removal can't touch
Every other guide on this site treats a photo as a file with a header wrapped around a picture, and treats cleaning as the act of removing that header. This page follows an image across a boundary where the header stops existing altogether. Paper has no EXIF block. Nothing on a printed sheet records a camera model or a set of coordinates. That sounds like the end of the problem, and it is the beginning of a different one — because printers write identifiers of their own onto the page, and a scanner turns everything on that page into ordinary pixels. Once information becomes picture, a metadata tool does not reach it. That is not a flaw in the tool; it is the boundary of what the word "metadata" means.
Ready to clean a photo? MetadataWipe processes JPEG and PNG files locally — no account and no server transfer.
Open MetadataWipe toolThe round trip is ordinary. A tenant prints photographs of a leak and hands them over at a meeting. A claimant prints and posts hard copies to an insurer. Someone photocopies a letter before returning the original. A small business scans signed paperwork back into a folder. A person with a sensitive document prints it precisely because paper feels like it escapes the digital trail. In each case an image or document crosses from file to page and, often, back to file again, and at every crossing the set of identifiers that travels with it changes completely.
Three layers, and only one of them is metadata
It helps to name the layers separately, because people routinely clean one and assume they have cleaned all three.
- The tag layer. The structured fields inside an image file: EXIF, XMP, IPTC, PNG text chunks, color profiles, embedded thumbnails. When you print, this layer simply does not make the journey. When you scan, a brand-new tag layer is written by the scanner and its software — commonly device and software names, versions, a scan date and time, and often a color profile. This is the layer a metadata tool removes.
- The picture layer. Everything that is part of the image as an image. In a scan, this is a photograph of a piece of paper, so it contains every mark on that paper: the printed content, handwriting, staple holes, a fold shadow, a coffee ring, the edge of the platen — and any marks the printer itself laid down. No metadata tool touches this layer, because none of it is a field.
- The context layer. Everything outside the bytes: the filename the scanning software generated, the folder it landed in, the file's timestamps on disk, and — for shared office equipment — the job records a multifunction device or print server may retain. What gets logged and for how long depends on the product and its configuration, so it is worth asking rather than assuming, but it is entirely outside the reach of any browser tool.
Our page on what metadata a scanned document has covers the first layer in detail — what scanners write and why GPS is usually absent. This page is mostly about the second, which is the one people do not expect.
What color laser printers put on every page
The Electronic Frontier Foundation has researched printer tracking since 2004 under what it calls the Machine Identification Code Technology Project. Its published account is that many color laser printers and copiers place a faint pattern of yellow dots on every page they produce, forming a coded pattern that can be used to identify specific details about the output. For at least one printer family EFF was able to decode the pattern, and it describes the dots as allowing a document's origin and its date of printing to be ascertained — in effect, a printer serial number and a timestamp, laid down in a color and at a size that is difficult to notice under ordinary light.
Three things about EFF's own framing matter more than the headline, and they are the parts most summaries drop.
First, EFF explicitly warns against reading its model list as a clearance. The list is no longer being updated, and EFF's standing note is that it appears likely that all recent commercial color laser printers print some kind of forensic tracking code — not necessarily using yellow dots, whether or not those codes are visible to the eye, whether or not the model appears on the list, and including models the list records as producing no dots. A "no" entry only ever meant that EFF could not see yellow dots on that model's output; it was never a finding that no forensic marking was present.
Second, EFF describes a later generation of the technique that does not work by adding anything. Documents it obtained indicated a subsequent approach that slightly rearranges dots the printer was going to print anyway. That has a direct consequence for anyone hoping to check their own equipment by looking: a visual inspection that finds no stray yellow dots does not establish that a page is unmarked, because the newer approach leaves nothing extra to find.
Third, EFF's account of the origin is that the technology results from arrangements between governments and the printer industry going back more than a decade, with the stated original motivation being currency counterfeiting — and that nothing in the technology limits its use to that purpose, with no law restricting the tracing of non-currency documents. Some manufacturers acknowledge that a tracking mechanism exists but offer few details.
For printing technologies other than color laser, EFF's position is narrower: as far as it knows, printers other than color laser and similar technologies do not deliberately encode their serial numbers in their output. Read that as a statement about what has been established, not as a clean bill of health for home inkjets. Forensic examination of printed output by other means remains possible.
Why a scan carries the marks into the file
This is the join that makes the topic belong on a metadata site at all. A scanner does not read a page; it photographs it. Whatever is on the paper becomes pixel values in a new image, and from that point on it is content, indistinguishable in kind from the text or the photograph you actually meant to capture.
The clearest public illustration is a 2017 episode EFF wrote about, in which journalists and researchers noticed that a scanned document published by a news outlet contained tiny yellow dots produced by a Xerox DocuColor printer, and one researcher decoded them using a tool EFF had built. EFF was careful about what that did and did not show, and so should anyone citing it: the arrest affidavit in the associated case did not mention the tracking dots at all and referred only to other sources of information, so it is quite possible the dots played no role in that investigation. What the episode does demonstrate cleanly is the mechanism — that marks laid down by a printer can survive into a distributed digital scan and be read out of it by a third party, long after the person who printed the page has stopped thinking about the printer.
Whether faint marks survive any particular scan is genuinely uncertain and depends on the scanner's resolution, how it handles color, and what compression is applied afterwards. That uncertainty cuts in the direction of caution rather than comfort: not knowing whether the marks made it into your file is not the same as knowing they did not.
It is worth being explicit about what MetadataWipe does here, because the honest answer is instructive. Cleaning a JPEG or PNG in this tool builds a clean copy by redrawing the image onto a canvas at its original width and height and exporting that as a new file, which discards the tag layer and carries the picture across unaltered — no resizing, no cropping, no filtering. That is exactly the right behaviour for a metadata remover, and it is exactly why it cannot help with layer two. This page is a description of a limitation, not a workaround for it, and it should not be read as advice on defeating forensic marking. Anyone whose safety depends on the untraceability of a printed document needs specialist advice, not a browser tool.
A working checklist for the paper round trip
- Ask whether paper needs to be in the chain at all. A file that is cleaned and sent digitally never acquires printer marks. Printing is sometimes required and sometimes just habit; only the first is worth the added layer.
- Treat every conversion as a new file. Cleaning the original before you print does nothing for the scan you make afterwards. The scan is a different file with a different, freshly written header, and it needs its own pass.
- Keep two questions apart. "Is this file clean?" and "Is this page clean?" have different answers and different tools. A metadata check only ever answers the first.
- Assume the printer marks the page. Following EFF's guidance, treat any recent commercial color laser printer or copier as emitting some form of tracking code, and do not rely on a visual check to rule it out.
- Remember whose printer it was. Scanning a page someone else printed imports their device's marks into your file. Scanning something you printed at work imports your employer's device into whatever you send.
- Watch photocopies specifically. A copier is a scanner joined to a printer, so a copy can carry both whatever the original page already bore and whatever the copying machine adds on its own account.
- Look at the shared-device trail. Scan-to-email on an office machine sends from an account and leaves a mail record; multifunction devices and print servers can hold job logs. None of that lives in the file, and none of it is affected by cleaning the file.
- Check the frame before the header. Scans routinely capture more than intended — a second document underneath, a margin note, a reference number in a corner. Crop or re-scan where that matters; those are picture problems with picture solutions.
- Then clean the scan. Once the page and the frame are settled, strip the scan's tag layer before sharing it, the same as you would any other image.
Mistakes and misconceptions
"Paper has no metadata, so a printout is anonymous." The first half is right and the second does not follow. The header ends at the printer; identification does not, because the printer participates in the output.
"I stripped the metadata from the scan, so the scan is clean." You closed the layer a stripper reaches. The picture is untouched by definition, and that is where printer marks live.
"My model isn't on EFF's list, so it's fine." This is the specific misreading EFF's own warning is written to prevent: the list is stale, a "no" only meant no visible dots, and the standing advice is to assume some form of marking regardless.
"I looked closely and there are no yellow dots." Visual inspection is not a test. Beyond the difficulty of spotting pale yellow on white, EFF describes a later technique that rearranges dots the page already contains rather than adding any.
"Printing and re-scanning is a good way to strip metadata." It does end the original tag layer, which is why the idea circulates. But it substitutes a new device's identifiers for the old ones, adds whatever the paper picked up, and costs real image quality through the optical round trip. If removing the tag layer is the goal, removing the tag layer is the cheaper and more complete way to do it — see stripping metadata from a photo before printing for the file-handoff side of the same workflow.
"Home and office printing are the same problem." They are not. An office device usually means an account, a queue, a log, and frequently a color laser engine. A personal printer removes the account and the log from the picture but not the marking question.
Related guides
See also:
Frequently asked questions
Does printing a photo remove its metadata?
It ends the tag layer, because paper has no header. There is no EXIF block, no XMP packet and no PNG chunk on a sheet of paper, so the capture time, the camera model and the GPS coordinates that were in the file do not travel onto the page. What people then infer from that — that the printed output is therefore anonymous — does not follow. Printing moves the question from the file to the page, and the page has identifiers of its own: what is visible in the image, whatever the printer itself adds to every sheet, and the physical characteristics of the paper and the print. If the page is later scanned, the new file starts with a fresh header written by the scanner and a picture that contains everything that was on the paper. Printing is a change of layer, not a cleaning step.
Do all printers add tracking marks to every page?
The most careful public position comes from the Electronic Frontier Foundation, which has researched this since 2004 and maintained a list of color laser models that did or did not show yellow tracking dots. EFF now states that the list is no longer updated and that it appears likely all recent commercial color laser printers print some kind of forensic tracking code — not necessarily using yellow dots, whether or not the codes are visible to the eye, whether or not the model is on their list, and including models their list recorded as producing no dots. They also note that a 'no' in that list only ever meant they could not see dots, not that no forensic marking was present. For other printing technologies, EFF says that as far as they know printers other than color laser and similar do not deliberately encode their serial numbers in output — which is a statement about what is known, not a guarantee of cleanliness. The practical reading: assume a color laser printer or copier marks every page, and do not treat any printer as verified clean.
Can MetadataWipe remove printer tracking dots from a scan?
No, and no metadata tool can, because those marks are not metadata. MetadataWipe works on JPEG and PNG files, one at a time, entirely in your browser. It runs a quick heuristic scan of roughly the first half-megabyte of the file to flag whether EXIF-like markers, GPS-like markers or PNG metadata chunks appear to be present, then produces a clean copy by redrawing the image onto a canvas at its original width and height and exporting that as a new JPEG or PNG. That export deliberately carries the picture across unchanged and discards the tag layer around it. Anything that exists as part of the image — a printer's marks captured by a scanner, a reflection, a visible document number, handwriting in a margin — is picture, so it survives by design. The tool is the right instrument for the header of a scanned file and the wrong instrument for the content of the page.
If I scan a document and strip the metadata, is it untraceable?
No. Stripping the scan's metadata closes one layer out of several. The picture itself still shows whatever was on the paper, including any marks the printer placed there. The file still sits in a filename that scanning software usually generates from a device or job name. The route the scan travelled can carry more than the file does — a scan-to-email feature on a shared office machine sends from an account, and multifunction devices and print servers can retain job records with user identities and timestamps, though what is logged and for how long varies by product and configuration. And the delivery path you choose afterwards has its own record. Treating a metadata strip as an anonymity guarantee is the mistake this page exists to prevent; it is a hygiene step that closes a specific, real leak, not a way to make a document untraceable.
Remove EXIF data, GPS location, and common photo metadata in your browser — one file at a time, before you share it.
Try MetadataWipe free