PDF to Markdown for RAG: check the text layer before choosing OCR
Last updated:
Document to Markdown extracts an existing PDF text layer and converts public digital documents into Markdown with optional overlapping RAG chunks. It has no OCR: route image-only scans to a separate OCR workflow, then check the recognized text before indexing it.
Verified on October 4, 2026 against the Actor's converter, layout logic, options, billing tests, real PDF/DOCX fixtures, README and pricing configuration.
No OCR is performed on scanned pages or embedded images. A returned row or successful run does not establish that all pages contain usable text.
Decide before ingestion
Selecting and copying a few lines in a PDF viewer is a useful first check, not a guarantee: a scanned PDF can have a poor hidden OCR layer, and selectable text can still be in the wrong order. Compare extracted words and page order with the original.
| Source condition | Recommended next step |
|---|---|
| Digital PDF with readable selectable text | Try text-layer extraction; review headings, order and tables |
| Image-only scan with no selectable text | OCR first using a separate product; this Actor cannot recognize the image |
| Mixed PDF with text pages and scans | Inspect every page warning; text pages do not make the scanned pages recoverable |
| Scan with a hidden OCR layer | Check recognition errors and logical order; this Actor reads that layer without re-OCR |
| DOCX or structured HTML is available | Prefer the structured source when layout fidelity matters |
| Tables, multiple columns or Arabic with suspect glyphs | Use the layout and Arabic verification guide before embedding |
PDF.js exposes document/page text for extraction; recovering an image's characters requires another recognition step. The Actor builds Markdown from PDF.js text items and coordinates, as reflected in PDF.js documentation. Reported multi-column extraction problems also show why readable characters and correct reading order are separate checks.
Convert a small sample first
curl --fail-with-body -X POST "https://api.apify.com/v2/acts/dropin-apis~document-to-markdown/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"],"maxPages":2,"includeImagesAlt":false,"chunkSize":1200,"chunkOverlap":100}'This is a batch Actor using Apify's synchronous run-and-dataset endpoint, not a persistent conversion service. For slow jobs, start a run asynchronously and retrieve its dataset after completion. The PDF API overview covers the full workflow.
| Format | Current conversion |
|---|---|
| PDF.js text layer; inferred headings, lists and simple aligned tables | |
| DOCX | Mammoth document structure, including headings, lists and tables |
| HTML | Static HTML article/main extraction with Readability and Turndown; no browser rendering |
| TXT | UTF-8 text, with HTML characters escaped |
| Markdown | Text retained with Unicode normalization |
HTML/DOCX image alt text can be retained through includeImagesAlt; PDF image alt text is not extracted. The Actor does not load page scripts or external document resources.
What the output proves
Each row contains url, contentType, title, markdown, pages, wordCount, chunks and warnings. pages counts inspected PDF pages, including blank/scanned ones; it is null for non-PDF formats.
A textless PDF page generates a warning that the page has no extractable text and OCR is unavailable. An entirely image-only fixture returns empty Markdown, zero words, zero custom page units and an explicit warning. It is a valid empty extraction, not an OCR success.
Every PDF carries a warning that headings, reading order and tables are inferred. If maxPages cuts the source short, a separate warning gives the source page count and selected cap. Do not discard warnings merely because some Markdown was returned.
Chunks do not repair extraction
chunkSize: 0 disables splitting. Otherwise, choose 200–20,000 Unicode code points, with chunkOverlap at most half the chunk size. The default overlap is the smaller of 100 or 20% of the chosen size.
Each chunk has index, start, end and text. Offsets refer to Unicode code points in the returned Markdown, with end exclusive; they are not byte offsets, PDF coordinates or token counts. Fixed-size splitting can cut across a table or heading, so inspect chunk boundaries for documents where those structures matter.
Preserve source URL, document title, warnings and extraction date alongside embeddings. Keep the full Markdown as the offset reference. Neither a word count nor a chunk array proves that retrieval will return a correct answer.
Bounds and price
| Limit | Current value |
|---|---|
| URLs | 1–20 public HTTP(S) URLs per run, sequential; duplicate URLs/fragments removed |
| PDF pages | First 20 by default; maxPages from 1–200 per PDF |
| Document bytes | 50 MiB transferred/decompressed, including bounded DOCX ZIP expansion |
| Network/conversion deadlines | 60 seconds for document fetch/DNS/retries; 45 seconds for conversion in an isolated worker |
| Output | Two million Markdown characters and 8 MiB per result item |
| Default whole-run timeout | 240 seconds; slow batches should be split |
Fetching obeys robots rules, validates redirect destinations and pins public DNS addresses. It can reject a public URL when robots rules are denied or unavailable. It does not bypass logins or CAPTCHAs. Extracted text remains untrusted data for any downstream agent.
Pricing is $0.0045 per page-converted unit, plus $0.0005 per Actor start at the default 512 MB; the start charge scales with selected memory. PDF units count inspected pages containing extracted text. Blank/scanned pages and failed conversions have no conversion event, but the run-start charge still applies.
DOCX, HTML, TXT and Markdown use one unit per started 2,000-word block per document. Chunk overlap adds no event. pages and chunk count are therefore not billing-unit counts. At default memory, two text PDF pages cost $0.0095 including the start; an image-only scan has no page-conversion charge and still gets no recognized text.
Run Document to Markdown on Apify for a digital document sample; choose a separate OCR service for scans.
FAQ
Can this Actor turn a scanned PDF into Markdown?
It cannot recognize image-only scans. Pages with no extractable text return warnings and no page-conversion units. Use OCR separately before indexing recognized text.
Does a successful run mean every PDF page was extracted?
No. maxPages can truncate a PDF, scanned pages can be empty, failed URLs can be skipped, and a spending cap can stop the batch. Compare requested URLs, inspected pages and warnings before ingesting results.
Are RAG chunk offsets tokens or PDF page coordinates?
Neither. start and end count Unicode code points in the returned Markdown, with an exclusive end. Preserve that Markdown so you can reconstruct the chunk text.
Do overlapping chunks increase the bill?
No. PDF billing counts pages with extracted text; other documents count started 2,000-word blocks. Chunk overlap adds no conversion event.
Related
- Parent: PDF to Markdown API and supported formats.
- Adjacent: PDF tables, columns and Arabic: what can be preserved.
- Other content: Direct media files to SRT/VTT.
- Discovery: Extract document URLs from a sitemap.