Does stripping metadata remove an image's alt text?

Nearly every page on this site treats metadata as something to get rid of. There is one field where that instinct runs into a second, equally legitimate requirement: the written description that lets a blind or low-vision reader know what is in a picture. Since 2021 there has been a standardised place to put that description inside the image file, and a metadata strip deletes it along with everything else. So two groups of people pull in opposite directions on the same bytes — the privacy-minded want the block gone, the accessibility-minded want part of it kept — and the usual advice on this site answers only one of them. The resolution is not a compromise. It is a fact about where accessible text is actually read from, and once you know it the two goals stop competing.

Need the tag layer cleared on a finished JPEG or PNG? MetadataWipe scans and rebuilds one image at a time in your browser — no account, and the file stays on this device.

Open MetadataWipe tool

This question deserves its own page because it is the first one on this site where stripping metadata can plausibly be accused of making something worse for a real person. Everywhere else the cost of over-cleaning is cosmetic: a lost credit line, a lost colour profile, a caption you have to retype. Here the thing at stake is whether a reader who cannot see the image gets any information about it at all — and that is a legal requirement in a good many places, not a nicety. It is worth getting the mechanics exactly right rather than guessing.

Two different things are both called alt text

The HTML attribute. This is the one that does the work. An <img> element in a page carries an alt attribute, and the browser uses the markup to compute the accessible name it exposes to assistive technology. The W3C Web Content Accessibility Guidelines put the requirement at the very top of the list: Success Criterion 1.1.1, Non-text Content, a Level A criterion, states that all non-text content that is presented to the user has a text alternative that serves the equivalent purpose, with a short list of exceptions. That text alternative is markup. It is written, stored and served by your site.

The field inside the file. The IPTC Photo Metadata Standard added a property called Alt Text (Accessibility) in version 2021.1, published in November 2021, alongside a companion called Extended Description (Accessibility). The standard defines Alt Text (Accessibility) as a brief textual description of the purpose and meaning of an image that can be accessed by assistive technology or displayed when the image is disabled in the browser, and says it should not exceed 250 characters. Extended Description (Accessibility) is the longer form: a more detailed description that elaborates on the alt text, with no character limit, not required when the short version is sufficient. Both are stored in the XMP block, under the names Iptc4xmpCore:AltTextAccessibility and Iptc4xmpCore:ExtDescrAccessibility, as language-alternative values so the same picture can carry descriptions in several languages. The standard is at version 2025.1 at the time of writing.

And the older fields that are not alt text. A photo may also carry a Description or Caption — the legacy IPTC field numbered 2:120, and its XMP and Exif equivalents — plus a Headline. IPTC is explicit that these are different jobs: the description is the who, what, when, where, why and how of the image, and the headline is a brief synopsis, while the accessibility properties describe the purpose and meaning of the picture for someone who cannot see it. A caption written for a photo library is usually a poor alt text, and stock libraries have been exporting captions for decades, so plenty of images contain a description that is not an accessible description.

The important point is the asymmetry. The attribute is read by browsers. The embedded field is read by software that processes files. No browser today opens the XMP packet in a JPEG and feeds that text to a screen reader, so an image published with no alt attribute is not quietly rescued by a description buried in the file. Treat the embedded value as a transport format — a way for a description to travel with the picture from the photographer to the picture desk to the publishing system — rather than as an accessibility feature of the image itself.

What a strip actually does, in both directions

Because the accessibility properties live in XMP, any all-or-nothing cleaner removes them. That includes this site's tool, which decodes a JPEG or PNG, draws it to a canvas at its original pixel dimensions and exports a new file from that canvas: the output is written from pixels, so the whole tag layer — EXIF, IPTC, XMP, accessibility fields included — simply is not carried across. There is no option to keep one property and drop the rest, because nothing is being edited; a new file is being written.

If your alt attributes already exist in your pages, that is the end of the story and nothing is lost. The failure case is narrower and more specific: a pipeline that was going to read the description out of the file.

WordPress is a good worked example, because the code is public. Its helper wp_get_image_alttext(), in wp-admin/includes/image.php and documented as added in version 7.0.0, reads the file, finds the XMP packet, and runs an XPath query for Iptc4xmpCore:AltTextAccessibility — looking first for a language entry matching the site locale, then a partial match such as en for an en_US site, then the x-default entry. The upload handlers in wp-admin/includes/media.php take whatever comes back and save it as the _wp_attachment_image_alt post meta, which is the Alt Text field in the media library. The same metadata reader pulls the legacy IPTC caption into the attachment excerpt and the headline or title field into the attachment title. Strip the file first, and every one of those arrives blank — which looks like a broken feature and is really just an ordering mistake. Exact behaviour will depend on your version and on the server having the extensions the reader needs, so check your own stack rather than assuming.

The mirror-image mistake is assuming your CMS cleans up after you. In the same codebase, the ImageMagick-backed image editor's metadata-stripping routine deliberately protects several profiles from removal, and the code comments name them: icc and icm for colour, iptc for copyright data, exif for orientation data and xmp for rights usage data. Everything else goes; those stay. So a platform can be stripping metadata, in good faith and by design, and still be serving the blocks that hold location, camera serial and edit history. What the pipeline does to your own files is covered in more depth in what metadata the images on your own website still carry.

The sequence that satisfies both requirements

  1. Decide which layer owns the description. For a website, the answer is the site: a field in your CMS, your DAM or your templates. A description that exists only inside a file is one re-export away from gone.
  2. Read the file before you clear it. Look at what is actually there rather than what you assume. The accessibility properties appear in a full metadata reader under names like XMP-iptcCore:AltTextAccessibility and XMP-iptcCore:ExtDescrAccessibility; the older text is under Description, Caption-Abstract or Headline.
  3. Copy out what is load-bearing. Accessibility text, any credit or licence line you are contractually required to carry, and anything a downstream system needs. Paste it somewhere durable before you touch the image.
  4. Let the import run first if you are relying on one. If your CMS populates alt text from the file, that read has to happen while the field is still in the file. Clean afterwards, or clean first and type the text in by hand — either works, but pick one deliberately.
  5. Then clear the tag layer on the copy you are going to publish, and keep your untouched original somewhere separate.
  6. Write the alt attribute from the text you saved, edited for the page it is on. The same photo legitimately gets different alt text in a news story and in a product listing, because the purpose differs.
  7. Put the long version where people can reach it. If the picture needs more than a sentence — a chart, a diagram, a complex scene — a visible caption or adjacent text serves everyone, not only screen reader users.
  8. Mark decorative images as decorative. An empty alt="" on a purely ornamental image is the correct answer, and is better than a description of a background texture.
  9. Verify on the served file, not the uploaded one. Fetch the exact image URL your page uses, including the resized variant the CMS generated, and inspect that. The contents of the blocks that survive a resize are the subject of does removing EXIF also remove XMP and IPTC metadata.
  10. Check the rendered page with a keyboard and a screen reader if you can, or at minimum view source and confirm each image has an alt attribute that is present, deliberate and not the filename.

Common mistakes and misconceptions

"Embedded alt text makes the image accessible wherever it goes." It makes the text available to software that chooses to read it. A browser rendering an <img> with no alt attribute does not consult the file's XMP, and neither does a messaging app or a document viewer. The embedded value is an input to publishing systems, not a substitute for markup.

"Stripping metadata broke the alt text on my site." Check the page source. If the attribute is still there, nothing broke. If it is empty, the strip did not reach into your database — something upstream was filling that field from the file, and the order needs changing.

"Keeping all the metadata is the accessible choice." Keeping all of it keeps the camera body, the lens, the serial number, the software history, the edit timestamps and, on a phone photo, usually the coordinates. There is no accessibility reason to publish any of that. If you want field-level control, you need a tool that edits metadata selectively; an all-or-nothing cleaner, including this one, cannot give you that.

"The description is the private part, so leaving it in is harmless." A good description can be the most revealing text attached to a picture. Two children in school uniform outside the house on Maple Street is a better alt text than most and a worse thing to ship to strangers than a GPS tag. Read your descriptions as a stranger would before deciding where they go.

"The caption in the file is the alt text." Captions and headlines answer what the picture is of and where it came from. Accessibility text answers what a reader needs to understand from it in this context. They overlap and are not the same, and IPTC defines them as separate properties for that reason.

"Alt text is for search engines." It has a side benefit there, which is why it is so often written as a keyword list. A keyword list read aloud is useless to the person it is supposed to serve. Write it for a human and let the side benefit happen on its own.

"My platform strips metadata on upload, so my published images are clean." Sometimes true, sometimes the opposite, occasionally both for different sizes of the same picture. As the WordPress example shows, a stripping routine can be written specifically to preserve IPTC, EXIF and XMP. Measure the file your site actually serves.

Where this site's tool fits, and where it does not

Stated plainly, because this page is about not misunderstanding a cleaning step. MetadataWipe accepts image/jpeg and image/png only, one file at a time. It decodes the image, draws it to a canvas at the original pixel dimensions, and exports a new file from that canvas — a fresh lossy encode for JPEG at a quality fixed in the code, a fresh PNG otherwise — appends -metadatawipe to the name and leaves your original file alone. All of it runs inside the page on your own device; nothing is sent anywhere.

Two consequences follow for this topic. First, the strip is complete and indiscriminate: the accessibility properties go the same way as the GPS tags, so copy any text you need before you clean. Second, the tool has nothing to do with your markup, so it cannot help you write an alt attribute and cannot damage one you have already written. Its built-in check is a heuristic rather than a parse — for a JPEG it looks near the start of the file for an APP1 segment carrying the Exif identifier and for the literal text GPS; for a PNG it walks chunk types looking for eXIf, tEXt, iTXt, zTXt, tIME and iCCP, which is where a PNG's XMP would normally sit. It reports presence, never values, so it will tell you that a text block exists and not what it says. For reading an accessibility description out of a file before you clear it, use a full metadata reader; for making sure the description reaches your readers, use the page.

Related guides

See also:

Frequently asked questions

Does removing EXIF data delete the alt text on my website?

No. The alt text your readers actually benefit from is an attribute in your HTML, and it lives in the page, not in the image file. You can strip every byte of metadata from a JPEG or PNG and the markup around it is untouched, so a screen reader announces exactly what it announced before. What a strip does remove is a separate, file-level field: the IPTC property Alt Text (Accessibility), which rides in the XMP block inside the picture. That value matters when some other system is going to read the file and copy the text out of it. If your alt attributes are already written in your CMS or your templates, stripping the file is a non-event for accessibility.

Where does an image's alt text actually live, in the file or in the page?

Both places can hold a description, and only one of them is read by a browser. The accessible name a screen reader announces for an image element is computed from the markup around it, which in practice means the alt attribute, or aria-label or aria-labelledby when those are present. There is no mechanism in a browser that opens the XMP packet inside a JPEG and hands that text to assistive technology, so an image with no alt attribute is not rescued by a description embedded in the file. The embedded field is best understood as a transport format: a way for the description to travel with the picture between a photographer, an agency, a picture desk and a publishing system, so that whoever builds the page does not have to invent the text from scratch.

My CMS filled in alt text automatically and now it is blank. What happened?

You very likely stripped the file before the upload rather than after, and the import had nothing left to read. WordPress core is a concrete example of this behaviour. Its helper wp_get_image_alttext, documented in wp-admin/includes/image.php as added in version 7.0.0 and present in the current development trunk, reads the file, locates the XMP packet, and runs an XPath query for Iptc4xmpCore:AltTextAccessibility, preferring a language entry that matches the site locale, then a partial match, then the x-default entry. The upload handlers in wp-admin/includes/media.php take that value and store it as the attachment meta key _wp_attachment_image_alt, which is the Alt Text box you see in the media library. The same reader also pulls the legacy IPTC caption into the attachment excerpt and the headline or title into the post title. Clear the metadata first and all of those arrive empty, which looks like a bug and is really just ordering. Copy the description out, or type it in afterwards.

Should I keep IPTC accessibility metadata in images I publish?

Keep the text, and decide separately whether the file is the right place to keep it. The text is valuable and someone should own it: write it once, store it where your site can reach it, and put it in the alt attribute of every page that uses the picture. Leaving it inside the published file as well is optional, and it has a cost, because the same XMP and EXIF blocks that carry a description also carry the camera body, the serial number, the editing history and often the coordinates. A description can leak on its own terms too, since a careful description of a photo may name a person, a street or a building. If the file is going to strangers, the safer default is to move the text into your publishing layer and ship a clean image. If the file is going to another professional system that expects to read it, keep the accessibility fields and clear the rest with a tool that lets you choose field by field.

Description saved and the alt attribute written? Clear the rest of the tag layer — scanned and rebuilt in your browser, nothing sent anywhere.

Try MetadataWipe free