Technology, openly explained
How DokAudit works.
Rules first, AI only where needed – and every change is re-measured. This page sets out which tools, algorithms and heuristics DokAudit uses for checking and repair, and what the AI does along the way.
- veraPDF: PDF/UA-1, PDF/UA-2, WCAG 2.2, PDF/A-2a
- Matterhorn checkpoints
- Repair loop until no findings remain
- 17 work steps: 12 deterministic, 3 heuristic, 2 with AI
Try it free · Integrations
Three kinds of work step
Every step on this page is labelled with its kind. Of 17 work steps, 12 are deterministic, 3 heuristic and 2 use AI.
- Deterministic: Fixed rules and standard checks – same file, same result.
- Heuristic: Rules of thumb based on layout, font and position – always re-measured.
- AI: Language and image model – only where rules are not enough; the result is checked again.
- Your choice: Per organisation: AI only with EU providers, or no AI at all. Documents are processed in Germany.
Five steps
- Check: Three checking tools, contrast, structure: veraPDF validates against PDF/UA-1, PDF/UA-2 and WCAG 2.2 in parallel; a second engine covers the Matterhorn checkpoints; contrast is measured on the rendered page.
- Repair: Fonts, structure, standard rules: fonts get Unicode mappings, the structure tree is built or repaired, tables get headers and scope, lists, links and form fields are tagged, artefacts are marked.
- Re-measure: After every repair round the file is validated again. The loop ends when no findings remain – or when only a human can decide.
- Mark: Only a file that passes gets the PDF/UA-1 identifier – and, on request, PDF/A-2a in the same file.
- Prove: A test report – itself accessible – documents the result, the repairs and anything left to review.
Checking
- veraPDF – PDF/UA-1, PDF/UA-2 and WCAG 2.2: Deterministic. The reference software of the PDF Association checks every file against three profiles in parallel; for the archive version also against PDF/A-2a. The rules are grouped by PAC's categories.
- Matterhorn checkpoints (pdfa11y): Deterministic. A second, independent checking tool covers the Matterhorn conditions that PAC also reports – such as table headers, lists, fonts and marked content.
- Logical structure: Deterministic. Our own check, like PAC's “Logical structure”: permitted nesting according to ISO 32000, headings without skipped levels, complete tables and lists, links with a Link element, meaningful alt text.
- Colour contrast on the rendered page: Deterministic. Text colour from the PDF, background from the rendered page (median of the pixels around each character). Contrast according to the WCAG formula, target 4.5:1, or 3:1 for large text. Logos and artefacts are excluded.
- Language and text layer: Heuristic. The document language is detected from function words (German, English, French, Italian, Spanish). DokAudit recognises scanned pages by their image share and missing text.
- Suitability for automatic tagging: Heuristic. Multi-column layout, rotated pages, forms, an unusually large number of font sizes – from these DokAudit decides whether rules are enough or AI structure recognition is needed.
- Score and traffic light: Deterministic. Starts at 100; each critical finding costs 25 points (missing tags 40), each notice 8. Without readable text, 20 at most. Green means: no finding open.
Repair – in this order
- Making fonts readable: Step 1 · deterministic. Missing Unicode mappings are added, non-embedded fonts are embedded as metrically identical substitutes with the original's character widths, and symbol fonts (Wingdings, Symbol) are mapped to real characters.
- Text recognition (OCR): Step 2 · deterministic. Only scanned pages get a text layer (Tesseract, German and English); existing text is left untouched.
- Generating structure (auto-tagging): Step 3 · heuristic. Text is wrapped in marked content, headings are determined by font size, tables are recognised from the column grid, and headers, footers and page numbers become artefacts. When in doubt, a paragraph rather than a wrong table.
- Structure for difficult layouts: Step 4 · AI. The AI receives the page image and numbered text blocks and answers only with roles and order – it does not invent any text. Its version is used only if veraPDF does not rate it worse than the rule-based version.
- Tidying up to the standard: Step 5 · deterministic. Table headers with scope, row and column spans, links and annotations in the structure tree, form fields with labels, bookmarks from the headings, title, language and tab order.
- Conformance loop: Step 6 · deterministic. Around 30 veraPDF rules each have a targeted repair assigned to them. DokAudit measures, repairs, measures again – and discards every round that is not better.
- Alt text: Step 7 · AI. Each image is cut out of the rendered page and described; purely decorative images become artefacts. Identical images are described only once. Note in the report: please proofread.
- Green loop up to 100: Step 8 · deterministic. Measured are veraPDF violations, PAC errors and PAC warnings. First our own rules apply, then what has been learned from earlier documents, and finally the AI with strictly defined operations. Every change is re-measured and kept only if it is better.
- Marking and archiving: Step 9 · deterministic. DokAudit sets the PDF/UA identifier only once the standard check has passed. Then PDF/A-2a: colour profile, metadata in standard form – adopted only if PDF/A-2a and PDF/UA-1 pass at the same time.
- Colour contrast (optional): Step 10 · deterministic. On request, text that is too light is darkened selectively in linear colour space; backgrounds stay the same. Optional, because the appearance changes.
What the AI does – and what it doesn't
AI helps with
- Alt text for images and graphics
- Structure recognition for multi-column or unusual layouts
- A cross-check against the page image: are the tables, headings and order right? Anything it finds appears in the report as “Please check”
- Headings, if the rules find none
- Labels for form fields without a tooltip
- Characters without a Unicode mapping – as the last step after outline matching and what has already been learned
Without AI, always by fixed rules
- Score and standard check
- Contrast measurement
- Text recognition (OCR)
- Font embedding
- PDF/A and PDF/UA-2
- Setting the PDF/UA identifier
Measured rather than trusted
Every AI result goes through the same checking tools. Anything that is not better is discarded.
Learned results before AI
Cases solved once (for example the characters of a font, or image motifs) are remembered by DokAudit, which then no longer needs AI for them.
Can be switched off, and limited
AI can be switched off per organisation; every document has a fixed AI budget. Without AI, everything runs on servers in Germany.
Data protection: for AI tasks DokAudit currently uses models from Anthropic; the provider is named as a sub-processor in the data processing agreement. Customer documents are not used for training. On request DokAudit works with an AI provider in the EU, with your own server or without any AI at all. View the data processing agreement (German)
How DokAudit differs
There are several good tools. The difference lies in how consistently results are re-measured, whether the outcome fits into an editorial team's day-to-day work – and what a page costs.
DokAudit, accessful and PAVE compared| Feature | DokAudit | accessful | PAVE 2.0 (ZHAW) |
|---|
| Approach | Check, repair and re-measure until no findings remain – fully automatic | Automatic correction with AI | Web tool: partly automatic, tags are reworked by hand in the browser |
|---|
| Checking within the process | veraPDF (PDF/UA-1, UA-2, WCAG 2.2) and Matterhorn checkpoints like PAC; every change is re-measured | Test log as evidence | Own check within the tool |
|---|
| Archive (PDF/A-2a) in the same file | yes, only if PDF/A-2a and PDF/UA-1 pass at the same time | not stated by the manufacturer | not stated by the manufacturer |
|---|
| Integration | REST API, webhooks and extensions for TYPO3 and Drupal – new PDFs are replaced automatically | REST API and webhooks | not stated by the manufacturer |
|---|
| Website scan and monitoring | yes, with pre-check portal and monthly monitoring | yes, Accessful Scan; check free of charge | not stated by the manufacturer |
|---|
| Price | from €0.50 per page (public sector plan), pilot free | prepaid packages from €1,950 for 780 pages, €2.00 to €2.50 per page | free for personal use |
|---|
| Hosting | Germany; AI can be switched off | Germany, also on-premises | ZHAW servers (Switzerland), documents stored for up to three weeks |
|---|
| Technology disclosed | yes – this page | main features | scientific publications |
|---|
Information on other providers according to their websites or publications, as of October 2026, without guarantee. More providers in the detailed comparison.
Limits – stated honestly
Machines check the technology of a document, not its meaning. Whether an alt text captures the essence of a chart, whether a heading hierarchy matches the meaning of the text or whether a link text is understandable is ultimately decided by a person. That is why every test report names the questions that are still open for exactly this document – no more and no less.
Tools and libraries
veraPDF · pdfa11y · pikepdf/qpdf · pdfminer.six · pdfplumber · PDFium · Poppler · OCRmyPDF/Tesseract · LibreOffice (Word → PDF) · fontTools · Docling · WeasyPrint
Standards: PDF/UA-1 (ISO 14289-1), PDF/UA-2 (ISO 14289-2), PDF/A-2a (ISO 19005-2), WCAG 2.2, EN 301 549, Matterhorn Protocol.