Why a Woolworths product page comes back empty with requests + BeautifulSoup

Last updated:

If requests.get() plus BeautifulSoup on a Woolworths product or search page returns an empty product list, that's not a bug in your selectors — the page's HTML genuinely doesn't contain the product data. Woolworths loads products through separate data requests after the initial page loads, so a plain HTML fetch only ever sees the empty shell the page starts with, no matter how the CSS selectors are written.

Why your scraper sees nothing

This is a real, reported problem, not an edge case: developers hit the same wall scraping other dynamically-loaded retail sites, where a direct GET returns a near-empty document and the actual listings arrive afterward via requests a plain HTTP client never makes (Stack Overflow, a second, independent report of the same class of problem). requests and BeautifulSoup only ever see whatever HTML the server returns for that one request — if the real product data arrives through a second, separate call the browser makes afterward, a plain HTTP client never makes that second call and never sees that data.

There are two common fixes: drive a real browser (Selenium/Playwright) so those follow-up requests actually fire, or call the same data endpoint the page itself calls, directly. This Actor takes the second path — it talks to Woolworths' public product-search data endpoint directly, the same one the website's own frontend calls, and returns clean, structured rows without needing a browser at all.

What you get instead

One normalized row per product — price, was-price, special flag, unit price, stock status — with no browser automation needed:

json
{
  "stockcode": "888140",
  "name": "Woolworths Full Cream Milk",
  "brand": "Woolworths",
  "packageSize": "3L",
  "price": 4.95,
  "wasPrice": 4.95,
  "isOnSpecial": false,
  "cupPrice": 1.65,
  "cupMeasure": "1L",
  "cupString": "$1.65 / 1L",
  "isAvailable": true,
  "isInStock": true,
  "searchTerm": "milk",
  "scrapedAt": "2026-10-02T00:14:46.726Z"
}

The second obstacle: datacenter IPs

Getting past the dynamic-loading problem isn't the whole story. On Apify's own infrastructure, Woolworths returns HTTP 403 to requests from Apify's datacenter IP ranges — confirmed with a real live run, not assumed — even when the data-endpoint call itself is otherwise correct. This Actor defaults to Apify's Residential proxy in Australia, the one option confirmed to actually work on the platform it runs on, and automatically switches to it the first time it sees an HTTP 403 with no proxy configured. You don't pay extra for this: proxy bandwidth is covered by the Actor's own pay-per-event price, not billed separately.

Limits

Category browsing isn't implemented in this version — categoryUrls is accepted but has no effect; use searchTerms. Up to 36 items per page, Woolworths' own hard limit (confirmed live: requesting a larger page size returns an explicit "Page size should not be greater than the limit: 36" error from Woolworths itself). Only public catalog/pricing data — no customer accounts, no checkout, no personal data of any kind.

Pricing

Pay per event — $0.001 per product ($1 per 1,000), no subscription, proxy cost included in that price.

Open Woolworths AU Products on Apify

FAQ

Is there an official Woolworths API?

No public product/pricing API exists for this. This Actor calls Woolworths' own public product-search data endpoint — the same one the website's frontend uses — and returns the response as clean, structured rows.

Why does my own requests + BeautifulSoup script get an empty product list?

Because the page's initial HTML doesn't contain the product data — it loads products through a separate request after the page itself loads. A plain HTTP client that only fetches the page's HTML never makes that second request, so it never sees the products, regardless of how correct your parsing code is.

Would Selenium or Playwright fix it?

Yes, driving a real browser would let those follow-up requests fire and the data would appear in the rendered DOM — but that's slower and heavier than calling the same data endpoint directly, which is what this Actor does instead.

Do I need my own proxy?

No. Apify Residential AU proxy is the default and is already included in the Actor's price. If you turn it off, the Actor detects the resulting HTTP 403 and switches back to it automatically, rather than returning an empty result.

Can I scrape a specific Woolworths category page instead of a search term?

Not in this version. Woolworths' category-browsing endpoint requires a request shape (CategoryId/Url/FormatObject) that this Actor doesn't yet construct — use searchTerms instead, which covers the same products from the search side.

Where do I find the full input/output reference?

See the main Woolworths AU Products page for every input field, the complete output schema, and code samples.