An embedded font without Unicode mapping: the printed appearance was correct, but copying and reading aloud produced gibberish. How DokAudit reconstructed the mapping – and why a simple copy test is worth doing for every document.
Anonymised case study from a real processing run; the name of the organisation and the document title have been changed.
A public notice from a local authority looked flawless – but anyone who copied text out of it got gibberish, and a screen reader read out nonsense. The cause: the embedded font did not reveal which letter was behind which glyph. DokAudit reconstructed this mapping automatically, step by step, from unambiguous methods through to AI as a last resort; afterwards, the text was readable and searchable, and the check was passed. The lesson applies to every PDF: ‘looks right’ does not mean ‘reads right’.
Starting point
The document was an official public notice (amtliche Bekanntmachung), exported from a layout program. As usual, the font was embedded as a subset – containing only the characters that actually occur in the document. On screen and in print everything was fine: the right font, the right letters, clean typesetting.
What the check found
A PDF does not store text as letters, but as references to the shapes of characters (glyphs) in a font. For a program to know that a particular shape is an ‘ä’, the font needs a mapping to Unicode – usually a ToUnicode table. In subset fonts, especially CID fonts without their own character map and without glyph names, this mapping is sometimes missing. The consequences:
- Copy and paste produces boxes, question marks or wrong letters.
- Search does not find words that are clearly visible on the page.
- Screen readers read out nonsense or skip the characters.
veraPDF reported the violation of rule 7.21.7-1 (‘The glyph can not be mapped to Unicode’); PAC also showed the error. In a few places, the glyph widths in the font dictionary also did not match those in the embedded font program (veraPDF rule 7.21.5-1).
Why this happens usually has to do with the export: when subset fonts are embedded, characters are renumbered, and not every program writes the mapping to Unicode in full. Particularly prone are special fonts for symbols, additional fonts that a program creates only for individual characters such as soft hyphens, and CID fonts whose glyphs have neither names nor a character map of their own. A checking tool detects the problem reliably – the naked eye never does.
What DokAudit did automatically
For every glyph used without a mapping, DokAudit determines the correct character – in a fixed order from inexpensive and unambiguous to elaborate:
- Outline matching: the shape of the glyph is compared with fonts whose mapping is known. If the outline matches, the character has been found.
- Learned mappings: shapes that have already been reliably identified in earlier documents are recognised again. Only a signature of the glyph shape and the corresponding character are stored for this, no document text.
- Shape analysis: simple shapes can be recognised geometrically – a horizontal bar is a hyphen, dash or underscore, plus full stop, bullet and space.
- AI as a last resort: only what is still open after that is identified – provided AI is permitted – collectively in one small image. The result is learned, so that the same shape is recognised without AI next time.
From the results, DokAudit generated the missing Unicode mapping, adjusted the glyph widths to the embedded font program where necessary and then checked the document again with veraPDF.
Result
Key figures of the font without Unicode mapping case study| Feature | Before | After |
|---|
| Copy and paste | gibberish | readable text |
|---|
| Search in the document | does not find visible words | works |
|---|
| Screen reader | reads nonsense or nothing | reads the text aloud |
|---|
| veraPDF 7.21.7-1 (Unicode mapping) | violation | passed |
|---|
| veraPDF 7.21.5-1 (glyph widths) | violation in a few places | passed |
|---|
What a person should still do
- Copy test at critical points: select paragraphs with umlauts, ß, € signs, section signs (§), dashes and ligatures such as ‘fi’ or ‘fl’, copy them and paste them into a text editor.
- Listening test: have a page read aloud by a screen reader (Test it yourself).
- Check inferred characters: mappings from shape analysis or AI are inferences. A spot check of amounts, dates and names makes sense.
- Fix the source: check the export settings and font choice in the layout program. Otherwise, if the next public notice is produced the same way, it will have the same problem.
Lessons learned
- ‘Looks right’ is not a test criterion. The printed appearance says nothing about whether a program can read the text.
- The copy test takes one minute. Select text, copy, paste – if nonsense comes out, the document is broken for screen readers and search, however good it looks.
- Unambiguous methods first. Outline comparison and learned mappings are traceable and solve most cases; AI is only there for the rest.
- Re-checking is part of it. Only the repeated check shows that the repair has worked – and that no new errors have been introduced.
Sources:
As of: 10/2026. This article gives a general overview and is not legal advice. The legal and standards texts in force are authoritative; in individual cases, the law of the German states (Länder) may differ.