Publishing an image dataset: what to strip and what to document

Nearly every page on this site — and nearly every guide to photo metadata anywhere — assumes a photo is on its way to an audience that might look at it, and that the metadata in it is purely a liability. Releasing a collection of images as data breaks both assumptions. The thing you publish is the set plus its documentation, the people who use it are strangers running code rather than viewers scrolling, and some of the metadata is not a liability at all: it is the reason the collection is worth having. The question stops being whether to strip and becomes which facts belong in the file, which belong in a record you write yourself, and which belong nowhere.

Sampling a few files before you plan the release? MetadataWipe rebuilds one JPEG or PNG at a time in your browser — no account, and the file stays on this device.

Open MetadataWipe tool

This covers image collections published for reuse: a research dataset attached to a paper, a citizen-science or wildlife-survey archive, a training or benchmark set, a municipal open-data release, a digitised collection going online under a reuse licence. It is not about sharing holiday photos with a group, and it is not another entry in the long list of guides to cleaning a photo before you post it somewhere. The failure modes are different enough that the usual advice actively misleads.

Why a published dataset is a different problem

Completeness is a stated virtue, so the usual instinct inverts. Everywhere else, less metadata is better. In a dataset, a reuser who cannot tell when an image was captured, by what device, or under what settings may not be able to use it for anything serious — and will have no way to detect that two apparently independent images came from the same camera five seconds apart. Stripping everything and saying nothing is not a cautious choice; it is a choice to publish something nobody can evaluate.

The exposure usually belongs to other people. The photographer is rarely the person at risk. It is the people in the frame, the landowner whose driveway is a survey point, the volunteer whose home coordinates sit in the first and last file of every batch they contributed, or — in ecological work — the animal whose den the coordinates lead to. None of them are in the room when you decide what to include, and none of them will see the release.

Publication is one-way. A dataset that is worth publishing gets mirrored, archived, forked, repackaged inside other datasets, and cached by things that do not take takedown requests. A photo posted to an account can at least be deleted from that account. A release cannot be recalled, which means the review has to happen before, not in response to a complaint.

The biggest leak is usually not in the images. Collections ship with the apparatus that makes them usable: a filename scheme, a folder hierarchy, a table of annotations, a README, occasionally a notebook with output still in the cells or a version history nobody pruned. All of that is metadata, written deliberately by a human to be read, and it is routinely more precise than anything a camera wrote. Header hygiene that stops at the image files can leave an exactly located, exactly dated, exactly attributed dataset with clean EXIF.

Three places a fact can live

The organising idea for the whole job is that every fact about an image can sit in one of three places, and the work is deciding which.

In the image header. EXIF, XMP and IPTC blocks travel with the file into every copy anyone ever makes. They are invisible in normal use, which is why they go unreviewed. Their granularity was chosen by the camera, not by you — a timestamp to the second, coordinates to several decimal places. And most tooling treats them as a block, so you get little or no say over which parts stay.

In a companion record you write. A manifest, a CSV, a JSON descriptor, a README table. This is the only one of the three where you control granularity, where the content is visible to you and to reviewers before release, and where a reuser will actually look. Capture time can be published to the day. Coordinates can be coarsened to a grid cell or replaced with a site code whose key you keep. A device can appear as a model without a serial.

Nowhere. Either never collected, or collected for your own use and deliberately excluded from the release — with the exclusion itself written down, so a reuser knows the limitation rather than guessing.

Stated as a rule: migrate the provenance you need out of the headers into documentation you control, then clear the headers. Not "preserve the headers just in case", which is how undocumented, unreviewed, unchosen detail ends up in a permanent public artefact.

This is not a novel idea in the data-publishing world, and it is worth knowing that the conventions already exist. The paper Datasheets for Datasets by Timnit Gebru and colleagues, posted in 2018 and later published in Communications of the ACM in December 2021, argues that the machine learning community has no standardised process for documenting datasets and proposes that every dataset ship with a datasheet covering its motivation, composition, collection process and recommended uses — by analogy with the datasheet that accompanies an electronic component. Container conventions exist too: the Data Package standard describes itself as a simple container format for a coherent collection of data, with a descriptor file carrying properties such as licences, contributors, sources, version and creation date.

The ecology world has a worked example that shows the model and its trap at the same time. Camera Trap Data Package, a community-developed exchange format for camera-trap data, is a Frictionless Data Package made of a datapackage.json describing the package and project, a deployments.csv listing camera placements, a media.csv listing the media files, and an observations.csv listing what was observed in them. That is exactly the right architecture: the provenance lives in reviewable tables rather than in image headers. It is also a reminder that the architecture does not make the decision for you. If deployments.csv carries exact placements for a sensitive species, stripping GPS tags out of ten thousand JPEGs has protected nothing — you moved the coordinates into a file that is easier to read.

Which fields are load-bearing and which are residue

It helps to sort the EXIF block into two piles before you decide anything. The field names below are as documented in the ExifTool tag reference, which is a practical place to check what a name actually means rather than guessing from it.

Often load-bearing for reuse: DateTimeOriginal and the time-zone offset fields such as OffsetTimeOriginal; Make and Model; the exposure parameters ExposureTime, FNumber, ISO and FocalLength; LensModel; pixel dimensions; and the colour profile, which is load-bearing in a different sense — drop it and the numbers in the file no longer mean what they meant.

Almost always residue: the serial fields, which ExifTool documents as SerialNumber — called BodySerialNumber by the EXIF specification — plus LensSerialNumber, ImageUniqueID, and OwnerName, which the specification calls CameraOwnerName. Then Software, HostComputer, Artist, Copyright, ImageDescription and UserComment, the last two of which are free text and may contain anything a previous application or a previous person put there. And the whole GPS block: GPSLatitude, GPSLongitude, GPSAltitude, GPSTimeStamp, GPSDateStamp, GPSImgDirection, GPSHPositioningError.

The serial fields deserve emphasis, because in a dataset they do something they do not do in a single photo. A body serial is a stable identifier for one physical device. Publish it across a few thousand images and you have published a join key: anyone can now group your dataset by device, link those groups to any other dataset or public gallery where the same camera's output appears, and attribute a contribution to a person who was never named anywhere in your files. That is a cross-dataset linkage you created by not thinking about a field you never looked at.

A release checklist

  1. Write down what the dataset is for. Every later decision is a trade-off against a purpose. If the purpose is undefined, the trade-offs get resolved by inertia, which means everything stays.
  2. Inventory what is actually in the files, by sampling. Do not assume uniformity. Contributions arrive from different devices, apps, exports and eras, so the fields present in one batch tell you very little about another. Sample across contributors, dates and sources.
  3. Establish your basis for publishing images of people and places — consent, permission, licence, institutional approval. What is required varies by jurisdiction, institution and subject matter, and this page is not legal advice; the point is that it belongs before the technical work, because the answer can change what you are willing to include at all.
  4. Draft the documentation before you touch a file, and decide granularity in it explicitly: time to the day or the second, coordinates exact or coarsened or coded, devices by model or by nothing. Writing it first forces the decisions to be deliberate.
  5. Extract the fields you are keeping into a table. Read them from the originals, into the manifest, at the granularity you chose.
  6. Build a fresh release folder. Copy in only what the release contains, rather than pruning your working directory. Pruning leaves whatever you forgot; building adds only what you chose. This single habit prevents most accidental inclusions.
  7. Rename to a neutral scheme and keep the mapping back to the original names in a private crosswalk that is not part of the release. Original filenames carry dates, places, client names, event names and camera sequence patterns, and no header strip touches them.
  8. Strip the image files with a batch tool you can run over a directory and re-run identically later.
  9. Audit the companion layer with the same seriousness. Annotation tables, coordinate columns, contributor columns, free-text notes, the README, notebook output, and any version history you are shipping alongside. This is where the specific detail lives.
  10. Spot-check from the release folder, not from your working copy, on randomly chosen files rather than the first one. Then freeze and version it, and record in the documentation what was removed or coarsened — so that a reuser treats the absence as a known limitation rather than as evidence of anything.

Two of those steps have more depth elsewhere on this site. The aggregate problem — the way a group of images gives away things no single image does — is the subject of why a set of photos leaks more than one photo alone, and it applies directly to step 9, because the risk you are auditing for is a pattern across rows rather than a bad value in one. And if the release is going onto infrastructure you run, the pipeline that serves the files has its own behaviour worth knowing, which is covered in what metadata the images on your own website still carry.

Where this site's tool fits, and where it does not

Being direct about this, because the honest answer is mostly "not here". MetadataWipe takes one JPEG or PNG at a time, decodes it in your browser, draws it onto a canvas at its original pixel dimensions, and exports a new file from the canvas. Nothing is sent anywhere; the clean copy is generated from data the page already holds and arrives with -metadatawipe added to the name, leaving the original untouched.

Three limits make it the wrong instrument for a release. There is no batch mode, no folder handling and no scripting hook, so a few thousand files is simply not the shape of work it does — and a release needs a repeatable command anyway, because you will publish a version two. The strip works by re-encoding, which for a JPEG is a fresh lossy encode at a fixed quality setting written into the code, so pixel values change slightly; for a dataset whose value is measurement rather than appearance, that is a transformation of the data and not merely a cleaning of it. And it is all-or-nothing: there is no field-level control, so you cannot use it to keep DateTimeOriginal while dropping the GPS block, which is exactly the operation a dataset release wants.

Its built-in check has limits worth stating too, since this page is about not trusting unreviewed things. The scan reads only the beginning of the file — the first 512 KB — and reports presence rather than values: for a JPEG it looks for the APP1 segment carrying the Exif identifier and for the literal text GPS in the first part of the file, and for a PNG it walks the chunk types looking for eXIf, tEXt, iTXt, zTXt, tIME and iCCP. It never displays a field value, and it is not an audit: a file whose only payload is XMP can legitimately come back as "No obvious metadata detected". For a dataset you want a real extraction tool that prints every field it can find.

What it is good for is the sampling step. Open a handful of representative files, see what gets flagged, produce a header-free version of one of them, and compare — a quick way to build intuition about your own material before you commit to a pipeline, and a reasonable way to clean the two or three figure images that go in the README itself.

Common mistakes and misconceptions

"We stripped the EXIF, so the dataset is anonymous." The annotation table, the filenames and the folder names are the dataset's real metadata, and a person wrote them to be understood.

"Keeping everything is the scientifically responsible choice." Completeness is not reproducibility. A field nobody documented is not documentation — it is an unlabelled number that a reuser has to guess the meaning of, shipped alongside things you would not have published on purpose.

"We can strip it later if someone objects." Mirrors, archives and derivative datasets do not revise themselves. Treat the release as irreversible, because it is.

"The locations are fine, the sites are public." Public to stand in is not the same as published to several decimal places with a timestamp. Precise coordinates for a nest, a den, a dig site, a shelter or someone's property are a different class of fact, and a season of them is a schedule.

"The images came from volunteers, so their metadata is their business." You are the publisher. Contributors did not choose the granularity of your release and in most cases do not know what their files contain.

"Re-encoding to clean the files is harmless." For a dataset used to measure something, a lossy re-encode alters the measurements. Prefer a tool that edits the container and leaves the compressed image data alone, and say in the documentation which you used.

"Nobody is going to read the headers of forty thousand files." Reading them is one command over a directory, and plenty of software does it automatically on ingest without anyone deciding to look.

Related guides

See also:

Frequently asked questions

Should I strip metadata from the images in a dataset at all?

In almost every case yes, but only after you have moved the facts a reuser genuinely needs into documentation you write and control. Those are two halves of one job and doing either alone fails. Stripping without documenting throws away capture times, device details and exposure settings that may be the reason the collection is worth anything, and leaves a reuser unable to tell whether two images came from the same camera on the same morning. Documenting without stripping ships serial numbers, exact coordinates, software fingerprints and free-text comment fields that you never reviewed, at a granularity nobody chose, inside every copy of every file forever. The headers are the wrong container for provenance because you cannot edit them selectively at scale, nobody reads them before reuse, and they carry whatever the camera felt like writing alongside the parts you wanted.

How do I keep capture dates and camera settings without publishing GPS and serial numbers?

Read the fields you want out of the originals into a table, then strip the files completely. Extracting to a table and then clearing the header is more reliable than trying to delete a precise list of fields and leave the rest, because a selective delete leaves you responsible for every field you did not think to name, including proprietary maker blocks and any free-text field a previous application wrote. Working from a table also lets you set granularity deliberately: a capture time can be published to the day rather than the second, coordinates can be coarsened to a grid cell or replaced with a site code, and a camera can appear as a model name without its body serial. Keep the full-precision version in a private working copy that is not part of the release.

Does stripping metadata make an image dataset anonymous?

No, and in a dataset that gap is wider than usual. A collection is published with the things that make a collection usable: filenames, a folder structure, an annotation table, a README, sometimes a notebook with cell output still in it. Those are metadata too, and they are often more specific than anything in the headers, because a human wrote them to be read. Faces, house numbers, licence plates, screens, documents and distinctive scenery stay in the pixels after any header strip. And a set has aggregate properties that no single file has, so repetition across hundreds of images can establish a place, a schedule or a route that no individual image gives away. Removing headers is one step in a de-identification process, not the process.

Can I use MetadataWipe for a dataset of a few thousand images?

No. This tool handles one JPEG or PNG at a time in your browser, with no queue, no folder walking and no scripting hook, so a release of any real size needs a local command line tool you can run over a directory and re-run identically when you publish version two. There are two more reasons it is the wrong instrument here. It rebuilds the image by drawing it to a canvas and exporting a new file, which for a JPEG means a fresh lossy encode at a fixed quality setting, so the pixel values change slightly and that matters for a dataset whose value is measurement rather than appearance. And it cannot touch anything that is not a single JPEG or PNG, so your annotation tables, filenames and README are entirely outside its reach. It is still genuinely useful for the sampling step: pull a handful of representative files, see what the scanner flags, and get a feel for what a header free copy of your material looks like.

Checking a representative file before you build the pipeline? JPEG or PNG, rebuilt in your browser, nothing sent anywhere.

Try MetadataWipe free