Beta · human review required
PDF remediation
Aelira scans PDF text and structure and may apply bounded fixes when the file exposes a safe target. Output is evidence for review, not a compliance certificate.
Release availability
Guide baseline: v0.9.11. The tagged PDF guide identifies v0.9.7 as the boundary for the immutable-source OCR, HTML safety and managed-publication controls described below. This Core release is available to self-host; this page does not claim the behavior is live in the hosted service.
Review the actual saved PDF
A completed original scan does not mean a remediated download exists. Supported verification compares original and saved bytes using the same deterministic settings. A lower measured score can be real; missing or unresolved evidence leaves the after-score unavailable. Partial output is withheld.
In Review, choose Show comparison to inspect actual original/saved page images and supported tagged text. MCID/ParentTree evidence is bounded and read-only; it is not a PDF order editor. Review mixed table/text placement separately from internal cell order and table semantics.
Scores, PDF comparison, approvals and missing outputBounded capabilities
| Area | What Core may do | Review boundary |
|---|---|---|
| Text and OCR | Profiles text and image presence per-page. Eligible zero-text image pages receive the current English-only OCR lane, and searchable text is preserved in the delivered PDF. | Proofread the recognized text and verify reading order and page meaning. |
| Structure | May update title and language metadata, bookmarks, and eligible heading, list, table, structure, and figure alt-text targets. | Complex relationships, tables, forms, reading order, and semantic intent can remain manual. |
| Validation | Re-scans output and can run implemented PDF checks before managed publication. | A passing check set does not certify PDF/UA or WCAG conformance; use the validators and assistive technology required by your organization. |
OCR, signatures, and the original
Remediation works from a private copy and refuses an output path that resolves to the input, so the original PDF remains immutable. Candidate generation and validation occur before pathname exposure; a failed attempt does not replace a prior valid output.
OCR suitability and output text are assessed per-page. Signed PDFs, XFA forms, indeterminate signature or language inspection, declared non-English documents that need OCR, image pages with partial direct text below the safe threshold, OCR engine refusal, and output pages without usable text fail closed for manual handling.
Accessible-HTML safety
PDF-derived titles and alt text are normalized and escaped for their HTML contexts. PyMuPDF page fragments are canonically rebuilt through a passive allowlist; active elements, event attributes, inline styles, unsafe URLs, comments, and malformed structures are removed or normalized.
Embedded images accept only validated PNG/JPEG data URLs. Byte, dimension, and pixel limits apply; Pillow performs structural verification, an independent reopen, and full pixel loading. PNG and JPEG terminal markers must occur at exact EOF, so trailing data, polyglots, spoofed or corrupt formats, decompression risks, and synthetic relative image requests are rejected.
Managed PDF publication
PDF and optional HTML candidates are created and checked through retained directory descriptors. The exact validated PDF is snapshotted into a private, unlinked output claim before the final pathname is exposed. The exact claimed stream is consumed by internal verification and by direct, queued, and Brightspace publication rather than reopening a mutable output path.
DB-first staging records the attempt before managed bytes become available. Size, SHA-256, MIME type, scan type, and filename are recomputed. Cancellation, ownership-fence loss, and commit failures abort only the staging attempt identified by its artifact ID and private publication token. Cleanup warnings, retained paths, and pending artifact recovery state stay explicit; no cleanup failure is reported as success.
Descriptor-bound managed publication depends on Unix/macOS/Linux filesystem APIs. Path-oriented compatibility remains for non-PDF formats and library consumers; this is not a Windows portability claim or a promise of one-and-only-one external effects.
Limits and evidence
- The PDF scan route accepts .pdf; MAX_FILE_SIZE_PDF defaults to 50 MiB. Content validation and the separate preview limits also apply.
- OCR can misread scans; generated alt text and inferred structure can be semantically wrong.
- Signatures and XFA are refused rather than silently invalidated.
- Format serialization is not lossless; compare visual, interactive, and structural behavior.
- PDF remediation is Beta and requires human review.