PDF Tips

PDF Accessibility Explained: Tagged PDFs and Screen Readers

Understand what makes a PDF accessible to screen reader users, and why an untagged, scanned PDF can lock out readers with visual impairments.

By Maya OrtizSeptember 8, 20264 min read

Why some PDFs are invisible to screen readers

A sighted person looking at a PDF sees headings, paragraphs, tables, and images arranged on a page. A screen reader, by contrast, has no idea what any of that visual layout means unless the PDF explicitly tells it. Many PDFs — especially ones exported from scanners or older software — are essentially just a flat picture of text with no underlying structure at all. To a screen reader, a page like that is silent: there's nothing to read aloud.

This isn't a hypothetical edge case. Government forms, academic papers, product manuals, and business reports are all routinely shared as PDFs, and a meaningful share of readers rely on assistive technology to access them. If a PDF isn't built with accessibility in mind, it can quietly exclude a portion of its intended audience.

What makes a PDF "tagged"

A tagged PDF has an invisible structural layer sitting behind the visual layout. This layer tells assistive software:

  • Which text is a heading, and what level (H1, H2, H3)
  • Which blocks of text are paragraphs versus captions or footnotes
  • The correct reading order, which doesn't always match the visual left-to-right order (this matters a lot for multi-column layouts)
  • What an image represents, via alternative text
  • Which cells in a table belong to which row and column headers

Without these tags, a screen reader either reads content in the wrong order, skips it entirely, or announces "image" with no further context for a chart or diagram that might contain critical information.

The two most common accessibility failures

1. Scanned documents with no text layer. If a document was created by physically scanning a paper page, it's just a picture — even though it looks exactly like a normal document to a sighted reader. Running the file through OCR PDF adds an invisible, selectable text layer underneath the image, which is the essential first step toward accessibility. OCR alone doesn't create full tagging, but it's the difference between a completely silent document and one that can at least be read aloud, searched, and copied.

2. Reading order that doesn't match visual order. In multi-column layouts (like newsletters or two-column academic papers), a screen reader might jump from the top of column one straight across to the top of column two, reading gibberish, instead of finishing column one first. This has to be fixed in the source document structure rather than the PDF itself.

Practical steps to improve accessibility

  1. Start with a text-based source whenever possible. A PDF exported from Word or Google Docs carries much more structural information than one created by photographing a page.
  2. Run scanned documents through OCR using OCR PDF so there's at least a searchable, readable text layer.
  3. Use real headings in your source document, not just bold, larger text. Word and Google Docs both have heading styles (Heading 1, Heading 2) that carry through into the exported PDF's structure.
  4. Add alt text to images in your source document before exporting, describing what the image conveys.
  5. Keep tables simple, with a single header row, so the row/column relationships are unambiguous.

Common myths about PDF accessibility

"If it looks fine on screen, it's accessible." Visual appearance and underlying structure are completely separate. A document can look perfect and still be unreadable by a screen reader.

"OCR alone makes a document fully accessible." OCR adds a text layer, which is a major improvement over a blank image, but full accessibility also depends on proper heading structure, reading order, and alt text — things OCR doesn't add on its own.

"Accessibility only matters for government documents." Any PDF meant for a general audience — a course syllabus, a product spec sheet, a job application form — benefits from being usable by everyone who might need to read it.

Frequently asked questions

How can I check if a PDF has a real text layer? Try selecting text with your cursor. If you can highlight individual words, there's a text layer. If your selection just draws a box around the whole page, it's an image.

Does converting a scanned PDF with OCR change how it looks? No — OCR adds an invisible text layer behind the existing image; the visual appearance of the page stays exactly the same.

Is there a difference between "readable" and "accessible"? Yes. A document with OCR text is readable (searchable, selectable, and can be read aloud roughly in order). A fully accessible document also has correct structural tags, reading order, and alt text, which typically requires attention at the point of document creation.

Accessibility is rarely a single checkbox — it's a combination of good source-document habits and, where scanning is unavoidable, adding a proper text layer with OCR.

#accessibility#screen readers#tagged pdf#ocr

Related articles