Extract structured data from a URL — six fixed schemas, every value backed by evidence
Last updated:
This API turns a public page into validated JSON with proof for every field — not a raw JSON-LD dump, and not a prompt-based AI scraper that can quietly invent a value. It supports six fixed schemas (product, article, event, organization, local business, documentation), checks JSON-LD first, then microdata, OpenGraph and restrained page-text rules, and never calls an LLM.
Example input
{
"urls": [
"https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
],
"schemaType": "product"
}Send 1–20 public HTML URLs of the same schema type per run — mixed page types need separate runs.
Example output (real result)
{
"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
"schemaType": "product",
"data": {
"name": "A Light in the Attic",
"price": "£51.77",
"currency": "GBP",
"availability": "In stock"
},
"evidence": {
"name": { "method": "html", "pointer": "article h1, main h1, h1", "snippet": "A Light in the Attic" }
},
"absentFields": [],
"source": { "contentType": "text/html; charset=utf-8", "bytes": 6812, "sha256": "..." }
}Every field — found or not — gets an evidence record. A field the page genuinely doesn't have comes back null with {"method":"absent","pointer":null,"snippet":null}, listed in absentFields, rather than a guessed value. Evidence shows which publisher text or metadata produced a value; it does not independently verify the publisher's claim is true.
What it extracts
| Schema | Fixed fields |
|---|---|
product | name, price, currency, availability |
article | headline, datePublished, description, section |
event | name, startDate, endDate, location |
organization | name, description, organizationType, location |
localBusiness | name, category, city, openingHours |
documentation | title, description, product, language |
How it compares
| Option | Price | What you get |
|---|---|---|
| This Actor | $0.009 per validated page | Six fixed, schema-validated contracts with per-field evidence; deterministic, no LLM; charged only on a successful, grounded result |
apify/ai-web-scraper | $0.03/page on its FREE tier ($0.02–0.025 on higher tiers) | Arbitrary prompts and arbitrary schemas — more flexible, no evidence contract, no fixed validation |
ninhothedev/json-ld-extractor | $0.0005/result + $0.00005 start | Raw JSON-LD copy — cheaper, but no normalization, no non-JSON-LD fallback (microdata/OpenGraph), no evidence, no null-on-missing discipline |
logiover/json-ld-schema-meta-tag-extractor | $0.005/result FREE ($0.0035–0.0045 higher tiers) | Similar raw-extraction tier to the above |
| Writing your own parser | Your time | Full control, but you maintain JSON-LD/microdata/OpenGraph fallback logic and schema validation yourself |
This Actor sits deliberately between the two tiers: 70% below the general AI Web Scraper's FREE-tier page price, because it answers a narrower question (one of six documented schemas, not any prompt) with a guarantee raw JSON-LD extraction doesn't make (checked, validated output with evidence, not just "whatever JSON-LD happened to be on the page"). Full competitor pricing sourced from the public Apify Store search API — see actors/structured-data-to-json/VERIFY.md for the complete table and methodology.
Why not a universal AI scraper?
A 20-page measurement gate (6 schema types, 80 fields, hand-labeled ground truth) found 97.5% field accuracy from deterministic extraction alone — JSON-LD, microdata, OpenGraph and restrained text rules, no model. A local quantized model was tested on the same fixtures and added zero correct fields while only increasing latency, so it's not included in this version. This isn't "AI-powered" marketing; it's the opposite finding, reported honestly.
Limits
- Six supported schema types only — no arbitrary prompts, no custom schema. Use a general AI scraper if you need either.
- No JavaScript rendering. A fact that only exists after client-side rendering stays
null, or the page can fail its identity check. - Public pages only. No cookies, login, form submission, CAPTCHA solving, Cloudflare bypass or browser fingerprinting.
robots.txtis respected, with policy-fetch errors failing closed (no access on an unreadable robots file, not open access). - No personal data. No person email, phone, profile or similar personal-data fields are extracted by this Actor.
- Up to 2 MiB per page after decompression; DNS resolved and private/reserved addresses refused before connecting, rechecked on every redirect.
Pricing
$0.009 per successfully validated page, pay per event. A page is charged only when it produces a dataset item with a grounded identity field and schema-valid output — invalid URLs, private destinations, robots denials, login/challenge pages, HTTP failures and empty results are never charged. Apify's standard small Actor-start event applies on top.
Open Structured Data to JSON on Apify
FAQ
Does it extract any field I describe in a prompt?
No. It supports the six documented schemas only (product, article, event, organization, local business, documentation). Use a general AI scraper for arbitrary prompts or custom fields.
Does it just copy all the JSON-LD from a page?
No. It selects and normalizes only the documented fields for the chosen schema, supplements them from grounded page metadata (microdata, OpenGraph, semantic HTML) when JSON-LD is incomplete, validates the final object against a fixed schema, and attaches an evidence record to every field.
What happens when a value is missing from the page?
The field comes back null, its evidence method is "absent", and it's listed in absentFields. The Actor does not guess or fill in a plausible-looking value.
Can it scrape private or logged-in pages?
No. Public pages only — no credentials, cookies, or login flows are accepted or attempted.
Is this affiliated with Schema.org or the sites it reads?
No. Schema.org names describe a public vocabulary of types; this is an independent Apify Actor with no affiliation to Schema.org or any site it extracts from.
Can an AI agent use this through Apify MCP?
Yes. An agent can find and call this Actor through Apify's MCP server, then inspect data, evidence and absentFields before trusting a value — the fixed, validated contract is deliberately easier for an agent to check than free-form prompt-generated JSON.