Extract structured data from a URL — six fixed schemas, every value backed by evidence

Last updated:

This API turns a public page into validated JSON with proof for every field — not a raw JSON-LD dump, and not a prompt-based AI scraper that can quietly invent a value. It supports six fixed schemas (product, article, event, organization, local business, documentation), checks JSON-LD first, then microdata, OpenGraph and restrained page-text rules, and never calls an LLM.

Example input

json
{
  "urls": [
    "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
  ],
  "schemaType": "product"
}

Send 1–20 public HTML URLs of the same schema type per run — mixed page types need separate runs.

Example output (real result)

json
{
  "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
  "schemaType": "product",
  "data": {
    "name": "A Light in the Attic",
    "price": "£51.77",
    "currency": "GBP",
    "availability": "In stock"
  },
  "evidence": {
    "name": { "method": "html", "pointer": "article h1, main h1, h1", "snippet": "A Light in the Attic" }
  },
  "absentFields": [],
  "source": { "contentType": "text/html; charset=utf-8", "bytes": 6812, "sha256": "..." }
}

Every field — found or not — gets an evidence record. A field the page genuinely doesn't have comes back null with {"method":"absent","pointer":null,"snippet":null}, listed in absentFields, rather than a guessed value. Evidence shows which publisher text or metadata produced a value; it does not independently verify the publisher's claim is true.

What it extracts

SchemaFixed fields
productname, price, currency, availability
articleheadline, datePublished, description, section
eventname, startDate, endDate, location
organizationname, description, organizationType, location
localBusinessname, category, city, openingHours
documentationtitle, description, product, language

How it compares

OptionPriceWhat you get
This Actor$0.009 per validated pageSix fixed, schema-validated contracts with per-field evidence; deterministic, no LLM; charged only on a successful, grounded result
apify/ai-web-scraper$0.03/page on its FREE tier ($0.02–0.025 on higher tiers)Arbitrary prompts and arbitrary schemas — more flexible, no evidence contract, no fixed validation
ninhothedev/json-ld-extractor$0.0005/result + $0.00005 startRaw JSON-LD copy — cheaper, but no normalization, no non-JSON-LD fallback (microdata/OpenGraph), no evidence, no null-on-missing discipline
logiover/json-ld-schema-meta-tag-extractor$0.005/result FREE ($0.0035–0.0045 higher tiers)Similar raw-extraction tier to the above
Writing your own parserYour timeFull control, but you maintain JSON-LD/microdata/OpenGraph fallback logic and schema validation yourself

This Actor sits deliberately between the two tiers: 70% below the general AI Web Scraper's FREE-tier page price, because it answers a narrower question (one of six documented schemas, not any prompt) with a guarantee raw JSON-LD extraction doesn't make (checked, validated output with evidence, not just "whatever JSON-LD happened to be on the page"). Full competitor pricing sourced from the public Apify Store search API — see actors/structured-data-to-json/VERIFY.md for the complete table and methodology.

Why not a universal AI scraper?

A 20-page measurement gate (6 schema types, 80 fields, hand-labeled ground truth) found 97.5% field accuracy from deterministic extraction alone — JSON-LD, microdata, OpenGraph and restrained text rules, no model. A local quantized model was tested on the same fixtures and added zero correct fields while only increasing latency, so it's not included in this version. This isn't "AI-powered" marketing; it's the opposite finding, reported honestly.

Limits

Pricing

$0.009 per successfully validated page, pay per event. A page is charged only when it produces a dataset item with a grounded identity field and schema-valid output — invalid URLs, private destinations, robots denials, login/challenge pages, HTTP failures and empty results are never charged. Apify's standard small Actor-start event applies on top.

Open Structured Data to JSON on Apify

FAQ

Does it extract any field I describe in a prompt?

No. It supports the six documented schemas only (product, article, event, organization, local business, documentation). Use a general AI scraper for arbitrary prompts or custom fields.

Does it just copy all the JSON-LD from a page?

No. It selects and normalizes only the documented fields for the chosen schema, supplements them from grounded page metadata (microdata, OpenGraph, semantic HTML) when JSON-LD is incomplete, validates the final object against a fixed schema, and attaches an evidence record to every field.

What happens when a value is missing from the page?

The field comes back null, its evidence method is "absent", and it's listed in absentFields. The Actor does not guess or fill in a plausible-looking value.

Can it scrape private or logged-in pages?

No. Public pages only — no credentials, cookies, or login flows are accepted or attempted.

Is this affiliated with Schema.org or the sites it reads?

No. Schema.org names describe a public vocabulary of types; this is an independent Apify Actor with no affiliation to Schema.org or any site it extracts from.

Can an AI agent use this through Apify MCP?

Yes. An agent can find and call this Actor through Apify's MCP server, then inspect data, evidence and absentFields before trusting a value — the fixed, validated contract is deliberately easier for an agent to check than free-form prompt-generated JSON.