Parsing a nested, compressed sitemap index completely — not just the first file

Last updated:

*A single sitemap.xml fetch is often not the whole site map: the Sitemaps protocol itself caps one sitemap file at 50,000 URLs and 50MB uncompressed, so any larger site publishes a sitemap index — a file that lists other sitemap files instead of URLs — and that index can itself be large, nested, and gzip-compressed.* Getting every URL means recursing into every child file the index points to, unzipping the compressed ones, and stopping at sane limits — not just reading the first file and calling it done.

Why a naive sitemap fetch misses URLs

Two real, separate failure modes show up here. First, large sites' sitemap indexes can list hundreds of child sitemap files — one developer's reported case needed to fetch over 400 individual child sitemaps from a single index with curl and xargs by hand (Stack Overflow), which doesn't scale as a manual or ad hoc script. Second, sitemap files are commonly served gzip-compressed (.xml.gz) to save bandwidth, and parsing that without writing a temp file to disk first is its own real implementation question developers have asked about directly (Stack Overflow). A script that only handles a flat, uncompressed sitemap.xml silently stops short on both counts.

How this Actor actually does it

Verified against this Actor's own source, not just described from the README: the crawl keeps two separate Sets — one of sitemap-file URLs it has already queued or fetched, one of page URLs it has already emitted — so a child file referenced twice by a sloppy index (or, in a pathological index, a cycle) is only ever fetched once, and a page URL that appears in two different sitemap files is only emitted as one output row, not two. Gzip files are detected by their magic bytes and decompressed with Node's zlib.gunzipSync directly in memory — no temp file ever touches disk, answering the "without downloading it to disk" question directly. Files are fetched one at a time into the active batch specifically so maxUrls can stop the run precisely, without needlessly downloading one extra child file past the point you actually needed.

json
{
  "startUrls": [{ "url": "https://www.gnu.org" }],
  "maxUrls": 500
}
json
{
  "url": "https://www.gnu.org/software/libc/index.html",
  "lastmod": "2026-08-14",
  "sourceSitemap": "https://www.gnu.org/sitemap0.xml.gz",
  "domain": "www.gnu.org"
}

Both halves of this were checked against real, live sites during development, not only synthetic fixtures: a real sitemap index (apify.com) for recursion, and a real gzip-compressed child sitemap (gnu.org's sitemap0.xml.gz) for the decompression path — sourceSitemap above shows that exact .gz child file as the row's origin.

The caps, and why they exist

At most 50 sitemap files per run, and 20MB per file after unpacking — generous relative to the protocol's own 50,000-URL/50MB-per-file ceiling, but a real, enforced stop so a misconfigured or hostile index can't turn one run into an unbounded crawl. maxUrls itself is clamped to a top of 50,000, matching the protocol's own per-file limit rather than an arbitrary round number. Fetches run at 5 concurrent requests, with up to 4 attempts and exponential backoff on HTTP 429 or any 5xx, and a per-host Crawl-delay from robots.txt is honored (up to 5 seconds) rather than ignored.

Limits

Public http/https sitemaps only — private hosts (localhost, 10.x, 192.168.x) are refused outright. A path disallowed for User-agent: * in robots.txt is never fetched, index or not. The run fails (not a silent empty success) if every input produces zero URLs across every candidate sitemap tried.

Pricing

Pay per event — the primary event is $0.0004 per URL ($0.40 per 1,000), charged when a row is saved. No monthly fee, and no separate charge for the sitemap-file fetches themselves, only for the page-URL rows that result.

Open Sitemap URL Extractor & Scraper on Apify

FAQ

Will this follow a sitemap index recursively, or only read the top-level file?

Recursively. An index's child files are queued and fetched one at a time (up to 50 files per run) until maxUrls is reached or the queue empties, not stopped after the first file.

Does it download a gzip sitemap to disk to unzip it?

No. A response starting with the gzip magic bytes is decompressed directly in memory with Node's built-in zlib, never written to a temporary file.

What happens if the same child sitemap is listed twice in an index?

It's fetched once. An internal set of already-seen sitemap-file URLs skips a repeat, so a redundant or even cyclical index reference can't double work or loop forever.

What if the same page URL appears in two different sitemap files?

It's emitted as one output row, not two — a separate internal set of already-emitted page URLs deduplicates across every sitemap file in the run, not just within one file.

Why does the sitemaps.org spec matter here?

It's the reason sitemap indexes exist at all: the protocol caps a single sitemap file at 50,000 URLs and 50MB uncompressed, so any site with more pages than that must split across multiple files behind an index — which is exactly the structure this Actor is built to recurse into.

How is this different from the main Sitemap URL Extractor page?

That page covers the full Actor end to end — every input field, the complete output schema, and code samples. This page goes deeper on the specific mechanics of nested indexes and gzip decompression for anyone whose sitemap isn't a single flat file. See also Extract Structured Data from a URL for turning the pages a sitemap lists into validated data once you have the URL list.