PDF automation suite

Keep translations in place where possible, instead of reflowing the whole document

Built for PDF files

— Core Capabilities Available Now —

Layout-preserving PDF translation

By default, translations are written back to their original positions and the output remains a PDF. OCR can also process text in scanned files and images.

See how the layout is preserved

PDFDOCX

Convert PDF to editable Word

Reflow PDFs with a text layer into editable DOCX, ideal for heavy post-translation editing. Precise layout is not retained. For scanned-only files, use Keep Layout translation instead.

Extract readable Markdown

Extract headings, paragraphs, and tables from PDFs to create Markdown for reading or knowledge bases. Content is reflowed; original fonts and precise layout are not preserved.

PAC accessibility check and remediation

Address common PDF/UA-1 issues identified by PAC, including table headers and tags, for better screen reader access.

Free Tools

Local PDF utilities for merging, splitting, compressing, rotating, and more. Separate from the translation pipeline, files are not uploaded to the cloud.

FAQ

Will the layout change after translating a PDF?

The default strategy is “Keep Layout”: translations stay at the original coordinates where possible, and the output remains a PDF. Complex posters, comic panels, or outlined text are not guaranteed to match the original exactly in this version. We recommend comparing the downloaded file with the original.

Can scanned PDFs be translated?

Yes. Use Keep Layout translation. The system runs OCR on scanned pages and text in images, then writes the translation back in place. Reflow to Word does not currently support scanned-only files.

Can I translate a PDF into editable Word?

Yes. Choose the “Reflow” strategy when translating, and PDFs with a text layer are exported as editable DOCX. This is Reflow rather than Keep Layout, and suits cases where you need to keep editing.

Can I configure a professional industry glossary before translation?

Yes. The system includes built-in pre-translate analysis, which automatically extracts high-frequency terms and proper nouns from the document. Before translation, you can review them and configure the glossary to ensure consistency throughout the document.

Will layout shift when translating into right-to-left languages such as Arabic?

No. The system includes a built-in RTL auto-mirroring engine. When the target language is Arabic, Hebrew, Persian, or another RTL language, it automatically flips text frames and element alignment after translation, with no manual re-layout needed.

Will translating long documents over 100 pages cause a timeout error?

No. The backend uses a distributed asynchronous processing architecture that automatically splits large documents for parallel processing (fan-out), avoiding LLM token limits and network timeouts, and supporting seamless translation of hundred-megabyte long documents.