PDF to Markdown for RAG: check the text layer before choosing OCR

Last updated:

Document to Markdown extracts an existing PDF text layer and converts public digital documents into Markdown with optional overlapping RAG chunks. It has no OCR: route image-only scans to a separate OCR workflow, then check the recognized text before indexing it.

Verified on October 4, 2026 against the Actor's converter, layout logic, options, billing tests, real PDF/DOCX fixtures, README and pricing configuration.

No OCR is performed on scanned pages or embedded images. A returned row or successful run does not establish that all pages contain usable text.

Decide before ingestion

Selecting and copying a few lines in a PDF viewer is a useful first check, not a guarantee: a scanned PDF can have a poor hidden OCR layer, and selectable text can still be in the wrong order. Compare extracted words and page order with the original.

Source conditionRecommended next step
Digital PDF with readable selectable textTry text-layer extraction; review headings, order and tables
Image-only scan with no selectable textOCR first using a separate product; this Actor cannot recognize the image
Mixed PDF with text pages and scansInspect every page warning; text pages do not make the scanned pages recoverable
Scan with a hidden OCR layerCheck recognition errors and logical order; this Actor reads that layer without re-OCR
DOCX or structured HTML is availablePrefer the structured source when layout fidelity matters
Tables, multiple columns or Arabic with suspect glyphsUse the layout and Arabic verification guide before embedding

PDF.js exposes document/page text for extraction; recovering an image's characters requires another recognition step. The Actor builds Markdown from PDF.js text items and coordinates, as reflected in PDF.js documentation. Reported multi-column extraction problems also show why readable characters and correct reading order are separate checks.

Convert a small sample first

bash
curl --fail-with-body -X POST "https://api.apify.com/v2/acts/dropin-apis~document-to-markdown/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"],"maxPages":2,"includeImagesAlt":false,"chunkSize":1200,"chunkOverlap":100}'

This is a batch Actor using Apify's synchronous run-and-dataset endpoint, not a persistent conversion service. For slow jobs, start a run asynchronously and retrieve its dataset after completion. The PDF API overview covers the full workflow.

FormatCurrent conversion
PDFPDF.js text layer; inferred headings, lists and simple aligned tables
DOCXMammoth document structure, including headings, lists and tables
HTMLStatic HTML article/main extraction with Readability and Turndown; no browser rendering
TXTUTF-8 text, with HTML characters escaped
MarkdownText retained with Unicode normalization

HTML/DOCX image alt text can be retained through includeImagesAlt; PDF image alt text is not extracted. The Actor does not load page scripts or external document resources.

What the output proves

Each row contains url, contentType, title, markdown, pages, wordCount, chunks and warnings. pages counts inspected PDF pages, including blank/scanned ones; it is null for non-PDF formats.

A textless PDF page generates a warning that the page has no extractable text and OCR is unavailable. An entirely image-only fixture returns empty Markdown, zero words, zero custom page units and an explicit warning. It is a valid empty extraction, not an OCR success.

Every PDF carries a warning that headings, reading order and tables are inferred. If maxPages cuts the source short, a separate warning gives the source page count and selected cap. Do not discard warnings merely because some Markdown was returned.

Chunks do not repair extraction

chunkSize: 0 disables splitting. Otherwise, choose 200–20,000 Unicode code points, with chunkOverlap at most half the chunk size. The default overlap is the smaller of 100 or 20% of the chosen size.

Each chunk has index, start, end and text. Offsets refer to Unicode code points in the returned Markdown, with end exclusive; they are not byte offsets, PDF coordinates or token counts. Fixed-size splitting can cut across a table or heading, so inspect chunk boundaries for documents where those structures matter.

Preserve source URL, document title, warnings and extraction date alongside embeddings. Keep the full Markdown as the offset reference. Neither a word count nor a chunk array proves that retrieval will return a correct answer.

Bounds and price

LimitCurrent value
URLs1–20 public HTTP(S) URLs per run, sequential; duplicate URLs/fragments removed
PDF pagesFirst 20 by default; maxPages from 1–200 per PDF
Document bytes50 MiB transferred/decompressed, including bounded DOCX ZIP expansion
Network/conversion deadlines60 seconds for document fetch/DNS/retries; 45 seconds for conversion in an isolated worker
OutputTwo million Markdown characters and 8 MiB per result item
Default whole-run timeout240 seconds; slow batches should be split

Fetching obeys robots rules, validates redirect destinations and pins public DNS addresses. It can reject a public URL when robots rules are denied or unavailable. It does not bypass logins or CAPTCHAs. Extracted text remains untrusted data for any downstream agent.

Pricing is $0.0045 per page-converted unit, plus $0.0005 per Actor start at the default 512 MB; the start charge scales with selected memory. PDF units count inspected pages containing extracted text. Blank/scanned pages and failed conversions have no conversion event, but the run-start charge still applies.

DOCX, HTML, TXT and Markdown use one unit per started 2,000-word block per document. Chunk overlap adds no event. pages and chunk count are therefore not billing-unit counts. At default memory, two text PDF pages cost $0.0095 including the start; an image-only scan has no page-conversion charge and still gets no recognized text.

Run Document to Markdown on Apify for a digital document sample; choose a separate OCR service for scans.

FAQ

Can this Actor turn a scanned PDF into Markdown?

It cannot recognize image-only scans. Pages with no extractable text return warnings and no page-conversion units. Use OCR separately before indexing recognized text.

Does a successful run mean every PDF page was extracted?

No. maxPages can truncate a PDF, scanned pages can be empty, failed URLs can be skipped, and a spending cap can stop the batch. Compare requested URLs, inspected pages and warnings before ingesting results.

Are RAG chunk offsets tokens or PDF page coordinates?

Neither. start and end count Unicode code points in the returned Markdown, with an exclusive end. Preserve that Markdown so you can reconstruct the chunk text.

Do overlapping chunks increase the bill?

No. PDF billing counts pages with extracted text; other documents count started 2,000-word blocks. Chunk overlap adds no conversion event.