Handling metadata in photos your users upload
Almost every guide to photo metadata — including almost every page on this site — is written for the person holding the file. You took the photo, you know roughly where you were, and you are deciding what to send. If you build something that accepts photo uploads, none of that describes your position. You are the pipeline. The files are other people's, taken with devices you have never seen, by people who in most cases have no idea what is inside them, arriving at a volume where nobody will ever open a single header. Whatever you decide happens automatically, to everyone, including the people who would have chosen differently. That is a different problem, and it fails in both directions: you can leak what you never looked at, and you can break images by cleaning them carelessly.
Inspecting a real upload while you design the pipeline? MetadataWipe scans and rebuilds one JPEG or PNG at a time in your browser — no account, and the file stays on this device.
Open MetadataWipe toolThis is for anyone whose software receives images from other people: a marketplace with listing photos, a support desk that takes screenshots, a claims or intake form, a community site with avatars and posts, a booking platform where hosts upload rooms, an internal tool where staff attach site photos. It is not a guide to cleaning a photo before you post it somewhere — the rest of the site covers that. The decisions here are about defaults, storage and derived copies, and the usual sender-side advice does not transfer.
Why being the pipeline is a different problem
You are choosing on behalf of people who did not choose. A sender weighs their own risk against their own convenience. You are setting one behaviour for a population that includes someone photographing a room they are subletting, someone uploading a screenshot of a bank statement to a support ticket, and someone whose home coordinates are in the first and last photo of every batch they ever send you. Your default is the decision for all of them, because effectively none of them will change it.
Volume removes review. On a single photo, a human can open the header and look. At a thousand uploads a day, nobody does, ever. Anything that depends on someone noticing a bad value will not be noticed. The only things that hold at that scale are the operations your code performs unconditionally on every file.
You create copies, and a strip touches one of them. The upload becomes a stored original, a set of resized variants, one or more thumbnails, perhaps a preview or a PDF rendering, a backup, an object-storage version, a cached object at a CDN edge, and sometimes a text layer: an auto-generated caption, OCR output, a moderation label, or EXIF values you copied into a database column. Cleaning the file you serve does nothing about the others. The question is never "do we strip" but "at which point, and which artefacts are downstream of it".
You can break images. A sender who mangles one photo notices immediately. A pipeline that mangles a class of photos ships the bug to everybody and finds out from support tickets. Orientation is the classic case, and it is covered below, but colour, sort order and quality all have the same shape: a strip that is slightly too blunt is a visible product regression.
It is one-way for your users. Once an image has been served from a public URL it may have been fetched, cached, mirrored and scraped. Replacing it later with a cleaned version does not retrieve the copies. For your users, your ingest decision is effectively permanent, which is why it belongs at ingest rather than in a later cleanup.
What actually arrives with an upload
Three layers arrive at once, and they are not interchangeable.
Inside the file bytes. The EXIF, XMP and IPTC payload: GPS coordinates, altitude and heading; capture timestamps with time-zone offsets; camera make and model; body and lens serial numbers; the software that last wrote the file; author, artist and copyright strings; free-text description and comment fields that may contain anything a previous application or a previous person put there; and an embedded preview thumbnail that in some files was generated before the last edit, so it can show a frame the visible image no longer shows. XMP can also carry named person regions from face tagging, which means a name attached to a face without anything in the pixels changing.
Around the file, in the request. The multipart filename your client sent, which is routinely something like IMG_4417.JPG but is just as often lease-signed-final.jpg or a name containing a person, a client or a place. The declared content type, which is a claim and not a fact. And the request itself: source address, user agent, your own receipt timestamp, the authenticated account. That last group is metadata you are creating, not metadata you received, and no amount of header stripping affects it — it lives in your logs.
Not in the file at all. Filesystem dates and any provenance an operating system attached on the uploader's machine do not travel inside the bytes of a normal upload, so they are not yours to strip and not yours to worry about. What does travel, and surprises people, is the paired-file case: motion photos, live photos and raw formats can arrive as more than one object, and a container format such as a PDF can arrive with a fully intact image embedded inside it.
Decide where in the pipeline the strip happens
There are three defensible places, with real trade-offs.
In the user's browser, before the bytes are sent. The strongest version of the privacy story, because the sensitive header never exists on your infrastructure in the first place — not in a temp file, not in a request log, not in a backup, not in an incident. It is also the only option that protects users from you. The limitation is that it is unenforceable: an old app build, a direct API call, or a client someone else wrote will send whatever it wants. Treat it as a valuable default path, not as a control.
On ingest, before the original is persisted. The strongest guarantee you can actually enforce, because every route into your storage passes through it. The catch is that it is irreversible, so if a feature six months from now needs capture times you no longer have them. The fix is ordering: extract the handful of values you have decided you need into your own fields first, at a granularity you choose, and only then clear the header.
On serve, or only when generating derivatives. Convenient, because the pristine original stays available for internal needs. It is also the option that quietly keeps the thing you were worried about: the original is now something you store, replicate, back up, restore and can leak, and every code path that hands out the original — a download button, a signed link, an admin export, an email attachment, a direct object-storage URL — bypasses the strip by definition.
If you want a default: extract first, strip on ingest, generate every derivative from the stripped master, and either do not keep the original or keep it in a separate store with a written reason and a retention period. The reason matters more than it sounds, because "we might need it" is how an unreviewed archive of other people's headers accumulates for years.
An ingest checklist
- Write down which header facts your product genuinely needs. For a large majority of products the honest answer is none. Some legitimately need capture time, and nearly all need orientation. If the list is undefined, it resolves to "keep everything", which is the thing you are trying to avoid.
- Extract those values into your own fields, at your own granularity. A date to the day rather than the second. A device as a model name, never a serial. Coordinates dropped, or coarsened to whatever precision the feature actually needs. Storing them in your database rather than in the file means you control who can read them and you can delete them.
- Normalise orientation into the pixels rather than relying on a tag that you are about to delete. This is the step teams skip, and it is why stripped photos appear sideways.
- Make a deliberate decision about the colour profile. An embedded ICC profile is display instruction rather than identity, so deleting it leaks nothing and can shift how the image looks. Either convert to a standard space on ingest or keep the profile; do not simply delete it and hope.
- Rename on ingest, to a server-generated identifier. Never reflect the uploaded filename into a public URL, a page title, an alt attribute or an email. Filenames carry names, dates, places and client references, and no header strip touches them.
- Strip, then build every derivative from the stripped master. Variants generated from the raw upload reintroduce whatever the upload contained, and this is a common ordering bug in pipelines that otherwise do the right thing.
- Audit the text layer you generate. Auto-captions, OCR text, moderation labels, image-search embeddings and any EXIF you copied into a column can all be more specific than the header you removed — and unlike the header, something in your product probably displays them.
- Check what your logs and error reports retain. An error tracker, request log or crash report that captures the raw request body is a copy of the original you forgot you had, usually stored somewhere with wider access than your image bucket.
- Verify by measurement, from outside. Fetch an image from its public URL the way a stranger would and inspect the bytes you received. Do it on files you did not pick: a portrait phone photo, a screenshot, a PNG, something from a scanner, something a user reported a problem with. A setting name, a plugin description or a library's reputation is not evidence.
- Describe the behaviour to users accurately, and no more. If you strip on ingest, say so. Do not say files are not stored if they are, and do not call an image anonymous because its header is empty.
Two neighbouring problems have more depth elsewhere on this site. Once your pipeline is serving images publicly, what the served files still carry is the subject of what metadata the images on your own website still carry, which is the right reference for step 9. And the orientation and colour failures in steps 3 and 4 are explained in more detail in does stripping EXIF affect photo orientation or colour profile.
Things that break, and the one that quietly does not
Orientation. A phone held sideways often stores sideways pixels plus a tag saying which way is up. Browsers honour that tag: MDN documents the initial value of the CSS image-orientation property as from-image, meaning EXIF orientation is applied unless you opt out, and the imageOrientation option of createImageBitmap defaults to from-image as well. Delete the tag without rotating the data and the photo flips on its side for everyone. Bake the rotation in.
Colour. Drop an embedded profile and a later consumer assumes some default space instead, which can show up as a visible shift in saturation on wide-gamut material. Users report this as "your site washed out my photo".
Chronology. If any part of your product sorts, groups or labels images by when they were taken, and you strip before extracting, every photo in the system becomes "uploaded today". This is irreversible for files already processed, which is the strongest argument for the extract-then-strip ordering.
Fidelity. A strip implemented as a full decode and re-encode is a quality change as well as a cleaning. For most products that is irrelevant; for anything where the image is evidence, measurement or a print source, prefer an approach that edits the container and leaves the compressed image data alone, and record which one you used.
And the quiet failure: nothing breaks, no error is logged, and the strip never happened. A pipeline that silently no-ops on a format it does not recognise, or that was configured to preserve metadata by a setting nobody remembers, produces exactly the same clean logs as one that works. Absence of an error is not evidence of a strip, which is why step 9 is a measurement rather than a code review.
Common mistakes and misconceptions
"Our image library strips metadata by default." Maybe, for some outputs, in the version you have, with your configuration, and only until someone changes a flag or a dependency bumps. Whether a payload survives depends on the library, its version, its options, and whether a build step, plugin or CDN transformation sits in front of it — none of which is visible from the page. This is a thing to measure on your own served output, not a thing to assume from a library's reputation.
"We resize everything, so it is gone." Even where that holds, it holds for the resized copy, as a side effect rather than a guarantee, and it says nothing about the original you also kept.
"We only ever serve the thumbnail." Enumerate the routes before believing this. Download-original buttons, signed links, admin and support tooling, CSV or ZIP exports, email attachments, webhook payloads and publicly readable object-storage paths are all ways the original leaves, and most of them were added by someone who was not thinking about headers.
"We strip it in the browser before upload." Good as a default path, not a control. Your API accepts what it accepts, regardless of which client called it.
"We do not keep the original." Verify that against your backups, your bucket's versioning setting, your soft-delete window and your temp directory. "Deleted from the database" and "no longer in existence" are different claims.
"Stripping metadata anonymises the upload." It does not come close. Faces, documents, screens, licence plates, house numbers and recognisable scenery stay in the pixels. Your logs still hold an address and a timestamp. Your database still holds the account that uploaded it. Header hygiene is one narrow control in a much larger picture.
"Our privacy policy already says we remove metadata." Then the policy is a commitment your pipeline has to keep, and an unverified claim is worse than no claim. Test the statement against a file you fetched from your own public URL before you publish the sentence.
Where this site's tool fits, and where it does not
Plainly: MetadataWipe is not pipeline software and cannot be made into it. Its source accepts image/jpeg and image/png only — HEIC, raw, PDF and video are rejected outright — and it handles one file at a time. There is no batch mode, no folder handling, no API and no scripting hook, so there is nothing for a server to call. It decodes the image, draws it onto a canvas at its original pixel dimensions and exports a new file from the canvas, which for a JPEG means a fresh lossy encode at a fixed quality written into the code. The clean copy arrives with -metadatawipe added to the name, leaving your original untouched, and all of it happens inside the page on your own device with nothing sent anywhere.
Its built-in check is worth understanding before you lean on it, since this page is about not trusting unverified things. The scan is a heuristic over the beginning of the file rather than a full parse: for a JPEG it looks within the first 256 KB for an APP1 segment carrying the Exif identifier and for the literal text GPS; for a PNG it walks the chunk types looking for eXIf, tEXt, iTXt, zTXt, tIME and iCCP. It reports presence, never values, so a file whose only payload sits beyond that window or in a form it does not test can legitimately come back as clean. For pipeline verification you want a full extraction tool that prints every field it can find.
What it is genuinely good for while you are building: inspecting a few real uploads to see what your actual users' files contain before you design around assumptions; producing a known-clean JPEG or PNG to commit as a test fixture; and checking a single served image after you have downloaded it from your own public URL, as a quick sanity pass between the real measurements.
Related guides
See also:
- What metadata the images on your own website still carry
- Does stripping EXIF affect photo orientation or color profile?
Frequently asked questions
Should I strip metadata from user uploads in the browser or on the server?
On the server, and treat anything the client does as a courtesy rather than a control. A browser-side strip before the bytes are sent is genuinely valuable because the sensitive header never reaches your infrastructure at all, so it never appears in a temp directory, a log, a backup or an incident. But you cannot enforce it. Your own client is one of several ways files arrive: a modified page, an old cached bundle, a mobile app on a version you shipped two years ago, a direct call to your API with a token, an integration someone built against your endpoint. Any of those can send whatever bytes they like, and a privacy promise that only holds when the client cooperates is not a privacy promise. The practical arrangement is both: strip in the browser so that the common path never transmits the header, and strip again on ingest so that the guarantee does not depend on what the client did. Do not treat the server strip as redundant just because the client usually handles it.
Will stripping EXIF on ingest break image orientation?
It can, and this is the single most common visible bug from a server-side strip. Many cameras store a photo's upright presentation as an Orientation tag rather than rotating the stored pixels, so the pixel grid is sideways and a tag tells the viewer to turn it. Browsers do act on that tag: MDN documents the initial value of the CSS image-orientation property as from-image, meaning EXIF orientation is applied by default, and the imageOrientation option of createImageBitmap also defaults to from-image. So if you delete the tag without rotating the pixels, a photo that looked fine before your strip will render on its side afterwards. The fix is to normalise orientation into the pixel data as part of the same operation, then strip, then verify on a real portrait photo taken by a phone held sideways rather than on a test image you made in an editor. Consumers outside the browser vary more, which is another reason to bake the rotation in rather than rely on a tag at all.
We resize every upload. Does that already remove the metadata?
Possibly, for the resized copy only, and not in a way you should rely on without measuring. Discarding the header is a frequent side effect of decoding an image and encoding a new one at a different size, but it is a side effect of a particular library at a particular version with a particular configuration, and some pipelines deliberately carry parts of the payload across to preserve copyright or colour information. A build step, a plugin or a CDN transformation in front of your code can change the answer again, and none of it is visible from the page. More importantly, the resized copy is usually not the only artefact: the original upload often still exists in object storage, in a backup, in a versioned bucket, behind a download-original route or in an admin export. Measure it the only way that settles it, which is to fetch an image from its public URL as a stranger would and inspect the file you actually received.
Can I use MetadataWipe as part of my upload pipeline?
No. It is a single-file browser tool with no batch mode, no folder handling, no API and no scripting hook, so there is nothing to call from a server and nothing to run over a directory. Its source accepts image/jpeg and image/png only, handles one file at a time, decodes the image, draws it to a canvas at its original dimensions and exports a new file from the canvas, which for a JPEG is a fresh lossy encode at a fixed quality. A production pipeline needs a library or command line tool you can pin, configure and re-run identically. Where it is useful to someone building that pipeline is as an inspection and fixture tool: open a representative real upload and see what the built-in scan flags, produce a known-clean JPEG or PNG to use as a test fixture, or download a served image from your own public URL and check what your pipeline actually emitted. Everything it does happens in the page on your own device, with nothing sent anywhere.
Checking what a real upload actually contains? JPEG or PNG, scanned and rebuilt in your browser, nothing sent anywhere.
Try MetadataWipe free