PDF tables, columns and Arabic: what Markdown conversion can preserve

Last updated:

Document to Markdown can infer headings, lists and simple aligned tables from digital PDF text, and retain Arabic Unicode text without reversing character strings. Complex columns, table associations and source glyph mappings remain best-effort, so compare the conversion with the original before using it for RAG or publishing it.

Verified on October 4, 2026 against the Actor's layout/conversion implementation, actual PDF and DOCX fixture tests, fixture provenance, live verification record, README and pricing configuration.

No OCR is included. Start with the text-layer or OCR decision guide if a PDF is scanned or has an unreliable hidden text layer.

What is established, and what is inferred

Layout featureCurrent behaviorWhat to verify
Headings and listsPDF font sizes and line markers drive Markdown heuristicsHeading hierarchy, missing bullets and false headings
Simple aligned PDF tablesConsecutive left-to-right rows with matching cell alignment become a Markdown table; first row becomes the headerHeader meaning, column association, units and totals
Arabic PDF tablesThe current PDF table detector excludes right-to-left rowsArabic text can remain without being a structured Markdown table
Multiple columnsLines/runs are ordered geometrically; there is no full semantic column reconstructionWhether adjacent columns interleave sentences or become a false table
Arabic textRuns use direction/coordinates; strings are not reversed, and presentation forms are normalizedSource word order, missing glyphs and mixed numbers/code
Rotated pagesExtracted with an additional rotation warningReading order and orientation
PDF imagesNo OCR or PDF image-alt extractionWhether essential table or figure content exists only as pixels
DOCX/HTML tablesExplicit document table structure is converted; a table without header cells uses its first row as header and emits a warningMerged cells, nested content and whether the first row is really a header

This distinction matters for other converters too. Maintainers have received reports of tables being treated as images, multi-column order problems, and Arabic logical-order and table-association failures. Those reports establish concrete failure modes, not comparative benchmarks or proof this Actor solves all of them.

Evidence from actual fixtures

The Actor's existing test suite uses real document parsers and genuine PDF/OOXML files. Synthetic fixtures test known content; they do not measure accuracy on arbitrary documents.

Tested caseObserved evidenceBoundary
Synthetic two-page English PDFHeading, bullet list, Item / Count table and second-page sentinel are recoveredOne simple digital layout, not a general complex-table benchmark
The same PDF with maxPages: 1Second-page text is excluded, a truncation warning appears and only one conversion unit is countedA bounded conversion intentionally omits later pages
Image-only/blank PDF fixtureEmpty Markdown, no-OCR warning, zero page unitsChecks absence of text; does not test OCR accuracy
Mozilla Arabic CID-font PDFArabic words survive without string reversal or replacement glyphs; forms are normalizedThe known source/font mapping imperfection remains
Genuine Arabic DOCXمرحباً بالعالم, الإصدار API 2026, heading/list structure and an Arabic table remain in logical orderThis stronger structured-source evidence is for DOCX, not arbitrary Arabic PDFs
Arabic mixed numeric/currency runsA coordinate-level regression keeps the Arabic phrase with its numeric content in the expected orderA focused layout control, not a full-document benchmark

The live Arabic PDF conversion recorded on October 2, 2026 returned four repeated lines including انواع اخلطوط العربية from Mozilla's ArabicCIDTrueType.pdf fixture. The imperfect word اخلطوط is not silently corrected to الخطوط. This is a known extraction/source-mapping limitation, and the layout warning remains.

Arabic DOCX and HTML use logical document text, whereas PDF text relies on glyph maps and positioned runs. A correct RTL preview cannot repair wrong characters or lost table relationships in the underlying text.

Sample the difficult pages

Run the source PDF with a page cap that includes the pages you want to inspect. maxPages selects the first pages; it does not offer arbitrary page ranges.

json
{
  "urls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"],
  "maxPages": 20,
  "includeImagesAlt": false,
  "chunkSize": 0
}

This W3C PDF is a connectivity/text-layer smoke test, not a table or Arabic benchmark. Replace it with a trusted public document representative of your own layout, and retain the full Markdown while reviewing it.

  1. For tables, compare each header, row label, value and unit with the source. A misplaced column can make a plausible sentence factually wrong.
  2. For columns, read the final line of one paragraph and the first line of the next. Check sidebars, footnotes and page transitions for interleaving.
  3. For Arabic, compare known phrases and proper names, then mixed Latin identifiers, numbers and punctuation. Check actual characters as well as their visual direction.
  4. Inspect every warning and the PDF pages count. Look for scanned/empty pages, rotation and maxPages truncation.
  5. Enable chunks only after conversion passes review. Verify that retrieval keeps labels with values; fixed-size chunks can split tables even when conversion was correct.

Use Waraq's Arabic Markdown editor to read the output locally alongside the source. Its RTL display helps review; it does not certify PDF reading order or fix glyph mapping.

For agent-driven ingestion, DoneLatch's evidence workflow can require a recorded check after the final change. Useful checks assert a known header/value relationship or phrase order and include a negative example that would fail if columns were swapped; a simple “Markdown is nonempty” check misses these failures.

Limits and cost remain the same

The Actor accepts 1–20 public document URLs; PDF extraction defaults to the first 20 pages and allows maxPages from 1–200. Each document transfer/decompressed content is capped at 50 MiB. There is no special table-repair or Arabic-OCR option.

PDF pricing is $0.0045 per inspected page containing extracted text, plus $0.0005 per run at the default 512 MB, with the start event scaling by memory. Blank/scanned pages and failed conversions have no page event. DOCX/HTML/TXT/Markdown use one event per started 2,000-word block per document. A readable but imperfect page can still incur a conversion unit; the fee does not certify semantic fidelity.

Try Document to Markdown on Apify and review a representative table, column boundary and Arabic phrase before ingesting the full collection.

FAQ

Does Arabic PDF support mean perfect Arabic reading order?

No. The Actor preserves Unicode text without reversing strings, but PDF glyph maps and positioned runs can still produce incorrect words or order. Review a known phrase and the mixed numbers/code in your own document.

Are Arabic PDF tables converted into Markdown tables?

The current PDF table detector excludes right-to-left rows. Arabic text can be extracted, but structured table recovery is not promised. Arabic DOCX tables have separate positive fixture evidence.

Can a two-column page be mistaken for a table?

Yes. Aligned text runs drive the PDF heuristic, so columns can interleave or resemble table cells. Compare paragraph transitions with the original before indexing.

Does chunking fix a broken table or OCR layer?

No. Chunks split the Markdown already extracted. Correct source text and layout first, then verify that chunk boundaries preserve the context needed for retrieval.