Documents
Document remediation covers PDF, Word, PowerPoint, Excel, and LaTeX source. It ships in github.com/Aelira-AI/aelira-core — nothing on this page is a hosted-only feature.
Where documents sit
Documents is one of four equal pillars in Aelira, alongside LMS integrations, web page scanning, and media. They cooperate through a shared scan model and findings format while each pillar can run independently.
Release availability
Guide baseline: v0.9.11. The tagged PDF guide identifies v0.9.7 as the boundary for the immutable-source OCR, HTML safety and managed-publication controls described here. This Core release is available to self-host; this page does not claim the behavior is live in Aelira's hosted service.
Coverage by format
This table is the boundary of what the document pipeline does today. Formats outside it are not handled, and remediation is in beta: output is a proposal for human review, never a signed-off document.
| Format | Checks | Can change | Human decision | Detail page |
|---|---|---|---|---|
| Tag structure, reading order, alt text presence, text layer on scanned pages, table structure. | May update metadata and bookmarks and apply eligible structure, heading, list, table, and figure alt-text changes when the file exposes a safe target. For eligible zero-text image pages, searchable text is preserved in the delivered PDF while the original PDF remains immutable. | OCR suitability and output text are assessed per-page. Signed or XFA files, partial direct text on image pages, non-English documents that need the current English-only OCR lane, and indeterminate inspection fail closed for manual handling. A person reviews OCR output, proposed alt text, and reading-order findings. | PDF remediation | |
| DOCX | Heading hierarchy, images and alt text, tables, fake lists, link text, language and title, SmartArt, and embedded objects. | May set eligible image alt text, apply heading styles, convert fake lists to list formatting, mark first-row headers, replace targeted link text, and set language and title. | Whether the resulting semantics express the author’s intent; SmartArt and embedded objects remain manual review and are not rewritten. | This page |
| PPTX | Image alt text, text/background contrast, slide titles, optional image-of-text OCR, and animation and media signals. | May set eligible image alt text, adjust text colour when foreground/background data is available, and add missing slide titles. | Reading-order findings produce guidance for manual review; Aelira does not structurally reorder shapes. Images of text and media captions or transcripts also remain review work. | PowerPoint scanner |
| XLSX | Sheet names, table and header structure, chart and image alternatives, merged cells, colour-only cues, frozen panes and navigation, pivot tables, and conditional formatting. | May rename a generic sheet, format or create a table with a first-row header, add a missing chart title, append text indicators for recognized colours, and freeze panes. | A person verifies inferred ranges and headers, chart-title meaning, colour semantics, formulas, pivots, and package fidelity. | This page |
| LaTeX | Document metadata, language declaration, figures, tables, equations, link text. | Preserves the .tex source and may improve metadata, language declaration, figure descriptions, table structure, equation output as MathML with ARIA labels, and link text. | Mathematical intent — a well-formed ARIA label can still describe the wrong thing. | LaTeX and MathML |
Word and Excel
DOCX, PPTX, and XLSX share the OOXML family but use separate processors, remediators, routes, and evidence. Word and Excel details stay here until dedicated pages exist.
Word (DOCX)
Checks heading hierarchy, images and alt text, tables, fake lists, link text, language and title, SmartArt, and embedded objects. Eligible changes include image alt text, heading styles, list formatting, first-row headers, targeted link text, language, and title. Semantics, SmartArt, and embedded objects remain human review work.
Excel (XLSX)
Checks sheet names, table and header structure, chart and image alternatives, merged cells, colour-only cues, frozen panes and navigation, pivot tables, and conditional formatting. May rename a generic sheet, format or create a table with a first-row header, add a missing chart title, append text indicators for recognized colours, and freeze panes. Recognized-colour text indicators can change cell text or value presentation. A person verifies inferred ranges and headers, chart-title meaning, colour semantics, formulas, pivots, and package fidelity.
Use .docx, .pptx and .xlsx. Legacy .doc, .ppt and .xls and macro-enabled .docm and .xlsm are not accepted by these scan routes. Renaming an extension does not convert a document. Serialization can normalize unsupported package parts, so compare the output in the intended Office application.
The release routes use a 50 MiB Word limit and a 100 MiB Excel limit when format-specific settings are absent. The PowerPoint setting defaults to 50 MiB. Content validation also applies. Read the format-specific upload contracts.
PDF output safety in this release
PDF-derived titles and alt text are normalized and escaped at their HTML contexts. PyMuPDF fragments are rebuilt through a passive allowlist that removes active elements, event attributes, inline styles, unsafe URLs, comments, and malformed structures. Embedded images accept only validated PNG/JPEG data URLs with bounded byte, dimension, and pixel checks.
Managed PDF output is prepared through descriptor-bound private candidates. The exact validated bytes are snapshotted into a private, unlinked output claim; internal verification and artifact publication consume the exact claimed stream. DB-first staging binds cancellation and recovery to the exact artifact attempt. Cleanup warnings and retained recovery state remain visible rather than being reported as success.
Human review is part of the pipeline
- A deterministic change does not automatically qualify for approval. Inspect the recorded outcome, verification evidence and current approval blockers. Review generated text and inferred semantics.
- If a source cannot be read, the issue stays open and is flagged for human review. Placeholder or invented text is never written in its place.
- Supported verification compares original and saved bytes using the same deterministic settings. A scan of the saved file alone does not establish the missing before-score. Partial output is withheld; manual, failed and unreported outcomes must remain explicit.
- Whether a document meets a standard is a determination the institution makes. The tooling produces findings, fixes, and re-scan results.
A managed Office artifact requires at least one fix, zero manual and failed issues, an output file and passed verification. When no artifact is published, do not assume a download exists. Follow the review workflow for scores, original findings, approvals and evidence exports.