POST
/scrapeScrape a page without JavaScript rendering
Scrapes a URL using Chrome browser TLS fingerprinting without executing JavaScript. Use this endpoint for pages that do not require client-side rendering, and set headers, geo, or proxy to customize the request. The response includes the HTML body and request metadata.
- IdempotentThe SDK sends
Idempotency-Key, so a retried request is only applied once.
JSON request containing the URL to scrape and optional request, proxy, retry, and extraction settings.
urlstringrequired
URL to scrape
headersarray<string>optional
Custom headers to send with the request. By default, regular Chrome browser headers are sent to the target URL.
retryNumintegeroptional
Amount of attempts.
geostringoptional
Geo location for basic proxy pools (you can purchase premium ScrapeNinja proxies for wider country selection and higher proxy quality). [Read more about ScrapeNinja proxy setup](https://scrapeninja.net/docs/proxy-setup/)
proxystringoptional
Premium or your own proxy URL (overrides `geo` field). [Read more about ScrapeNinja proxy setup](https://scrapeninja.net/docs/proxy-setup/)
followRedirectsintegeroptional
Follow redirects.
timeoutintegeroptional
Timeout per attempt, in seconds. Each retry will take [timeout] number of seconds.
textNotExpectedarray<string>optional
Text which will trigger a retry from another proxy address.
statusNotExpectedarray<integer>optional
HTTP response statuses which will trigger a retry from another proxy address.
extractorstringoptional
Custom JS function to extract JSON values from scraped HTML. Write&test your own extractor on https://scrapeninja.net/cheerio-sandbox/
200Returns the HTML body and an `info` object containing the response status code, final URL, and headers.
infoobjectoptional
bodystringoptional
HTML body of the rendered page.
Error handling
url is required and must identify the page to scrape. When supplied, proxy overrides geo; timeout defaults to 10 seconds per attempt and statusNotExpected defaults to 403 and 502.