PDF to Markdown API, with ready-to-embed RAG chunks
Last updated:
This Actor converts public PDF, DOCX, HTML, TXT and Markdown URLs into Markdown, with word count, source title and optional overlapping chunks for retrieval-augmented generation. It reads each document's text layer; it does not perform OCR on scanned pages.
What it converts
| Format | How | Limits |
|---|---|---|
| PDF.js text layer; inferred headings, lists and simple aligned tables | First 20 pages by default, maximum 200. Complex multi-column layouts are best effort. | |
| DOCX | Mammoth: headings, paragraphs, lists and tables | Public Word files only, no external files or image downloads. |
| HTML | Mozilla Readability article extraction, then Turndown Markdown | Static HTML only — no browser rendering, login or CAPTCHA bypass. |
| TXT / MD | UTF-8 text | TXT is escaped as plain text; Markdown is kept as-is. |
Example input
{
"urls": [
"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
"https://example.com/"
],
"maxPages": 2,
"includeImagesAlt": true,
"chunkSize": 1200,
"chunkOverlap": 100
}urls accepts 1–20 public HTTP(S) URLs, processed sequentially. chunkSize: 0 disables chunking; otherwise 200–20,000 Unicode code points, with overlap at most half the chunk size.
Example output
{
"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
"contentType": "application/pdf",
"title": "dummy.pdf",
"markdown": "<!-- Page 1 -->\n\nDummy PDF file",
"pages": 1,
"wordCount": 3,
"chunks": [],
"warnings": ["PDF headings, reading order and tables are inferred from the text layer; complex layouts may need review."]
}Each chunk, when chunkSize is set, carries index, start, end, text, with offsets counted in Unicode code points of the returned Markdown.
How it compares
| Option | OCR | Arabic text | RAG chunking | Price |
|---|---|---|---|---|
| This Actor | No (text layer only) | Yes — logical Unicode order, not reversed | Built in, overlap configurable | $0.0045 per converted page (PDF), or per 2,000-word block (HTML/DOCX/TXT/MD) |
| A full OCR pipeline | Yes | Varies | Usually a separate step | Typically higher per-page cost |
| Copy-pasting manually | N/A | Manual | Manual | Your time |
This Actor does not do OCR — a scanned or image-only PDF page returns an explicit warning and no text, with no hidden OCR surcharge. Use a dedicated OCR product for scans.
Run it from code
curl -X POST "https://api.apify.com/v2/acts/dropin-apis~document-to-markdown/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://example.com/"],"chunkSize":1200}'In Apify's MCP server, search for document-to-markdown and call it with call-actor using the input above.
Pricing
$0.0045 per converted page-equivalent, plus $0.0005 per Actor start at the default 512 MB. A one-page PDF costs $0.005; a two-page text PDF costs $0.0095; ten short HTML documents in one run cost $0.0455. Failed conversions have no page charge.
Open Document to Markdown on Apify
FAQ
Can it read scanned PDFs or images?
No. A scanned or blank PDF page returns an explicit warning and no text — there is no hidden OCR surcharge. Use a dedicated OCR product for scanned documents.
Does it preserve every PDF table and column?
No. PDF text coordinates infer heading sizes, lists and simple aligned tables; complex multi-column pages, rotated content and inaccurate font maps can need review. DOCX and HTML use their own document structure directly.
Does it handle Arabic text correctly?
Yes for text-layer content. Arabic DOCX and HTML retain logical Unicode text, including mixed Latin words and numbers. PDF.js extracts Arabic from real text-layer PDFs without reversing characters, though arbitrary PDF glyph-mapping errors in a source file cannot be universally repaired.
What does one page cost?
$0.0045 per converted page for PDF (only pages with extracted text are charged), or one event per started 2,000-word block for HTML/DOCX/TXT/MD, plus a one-time $0.0005 Actor-start fee per run at default memory.
Is this a persistent API I can call anytime?
It's a batch Actor: start a run with your URLs, then read the dataset. It needs no subscription, external OCR key, or separately hosted server, and works from the Apify Console, the run API above, or an MCP-connected agent.