How a scan becomes a readable PDF, why text recognition alone is not enough, where handwriting and poor originals set limits – and when creating the document anew is the better choice.
A scanned PDF contains only images of pages – for a screen reader it is empty. Text recognition (OCR, optical character recognition) places an invisible text layer behind the page images; the document then looks exactly the same but can be searched and read aloud (WCAG technique PDF7). OCR is therefore the prerequisite for accessibility, but not accessibility itself: afterwards the document still needs structure, language and a title – and recognised text that is actually correct.
How to recognise a scan
- Text cannot be selected or copied (selection test).
- Searching the document finds nothing, although the word is visible.
- The pages are slightly skewed, have punch holes, stamps or a grey border.
- Mixed forms: a digitally created document contains individual scanned pages, such as the signed last page or inserted attachments.
What OCR achieves
With cleanly printed originals in common typefaces, scanned straight, modern OCR generally recognises the text reliably. The right language is important: recognition uses dictionaries, and a German document with an English language setting produces more errors with umlauts and ß.
Text recognition is followed by structure: headings, paragraphs, lists and tables must be marked up, and running headers and footers marked as artefacts. With scans this is harder than with digitally created PDFs because there is no font information – headings can only be recognised by size and position. Automatically generated structures should therefore be spot-checked.
The limits
- Handwriting: handwritten notes, marginal comments and filled-in entries are recognised only unreliably. If they matter for the content, they have to be transcribed.
- Signatures: they are not ‘read’. If the signature is relevant, a short alt text is enough (‘Signature’, possibly with the printed name).
- Poor originals: faxes, carbon copies, low resolution, coloured paper, stamps over the text and strong skew significantly increase the error rate.
- Old typefaces: Fraktur (German blackletter) – for example in old by-laws or certificates – is recognised much less well by many OCR systems.
- Tables: in scans, lines and columns easily fall apart; header cells often have to be assigned by hand.
- Plans and maps: OCR picks out fragments of labels that only confuse when read aloud. Such pages are better treated as a figure with alt text.
- Recognition errors are invisible: a validator checks whether text is present – not whether it is correct. If a recognition error turns €1,800 into €7,800, the document passes every technical check and is still wrong. Amounts, dates and names therefore belong in the spot check.
Re-create, repair or exempt?
- Does the source file still exist? Then export again from it – that is almost always better than any OCR (Word guide).
- Does an exemption apply? Files published before 23 September 2018 are exempt under EU Directive 2016/2102, unless they are needed for active administrative processes. A by-law in force or a form in use is therefore generally not covered. Exempted documents belong in the accessibility statement.
- Is the document still needed? If not: take it off the website or archive it (see Keeping your website’s PDFs under control).
- Good original, typewritten text? OCR, structure, spot check – that is the normal case for archive documents.
- Poor original, a lot of handwriting, Fraktur, complex tables? Then re-keying is often quicker and more reliable than correcting – especially for short, frequently accessed documents.
- Reproductions from the archive? For items from heritage collections, the EU directive has its own exemption; how it has been transposed into the law of the German state concerned should be checked case by case.
Scanning better for the future
- Publish digital originals instead of ‘print, sign, scan’ – the signature is rarely needed online.
- If scanning is unavoidable: place the page straight, use sufficient resolution (at least 300 dpi is usual), greyscale rather than black and white for stamps and coloured paper.
- Switch on text recognition directly in the scanning workflow – with the right language.
- Remove blank pages and reverse sides, rotate pages correctly.
Common mistakes
- The scan is published without text recognition.
- ‘Searchable’ is confused with ‘accessible’ – the text layer is there, but tags are missing.
- The OCR text is never checked, and amounts and names are recognised incorrectly.
- Stamps, punch holes and page borders are read out as text instead of being artefacts.
- Rotated pages and double-page spreads stay as they came out of the scanner.
How DokAudit helps
DokAudit detects scanned pages and, in the fully automatic mode of the document hub, first runs text recognition, then structuring with headings, paragraphs, lists, tables and artefacts, plus title, language and bookmarks. The result is checked again with veraPDF and DokAudit’s own checks of the logical structure. No software can guarantee that the recognised text is correct in content – the test report points this out, and the spot check remains your task. If you only want to make a scan searchable, you will find a free tool under PDF OCR; an accessible result also requires the structure.
Sources:
As of: 09/2026. This article gives a general overview and is not legal advice. The legal and standards texts in force are authoritative; in individual cases, the law of the German states (Länder) may differ.