Automatically extract clean article text and other data from news articles, blog posts and other text-heavy pages.
- Provider
- Diffbot Extract API (Web Extraction)
- Input summary
- {"name":"url","type":"string","required":true,"description":"Target URL to extract"}, {"name":"fields","type":"string","required":false,"description":"Specify optional fields to be returned from any fully-extracted pages (e.g. `fields=querystring,links`)","enum":["links","extlinks","meta","querystring","breadcrumb","quote..., {"name":"timeout","type":"integer","required":false,"description":"Sets a value in milliseconds to wait for the retrieval/fetch of content from the requested URL. The default timeout for the third-party response is 30 seconds (30000)."}, {"name":"callback","type":"string","required":false,"description":"Use for jsonp requests. Needed for cross-domain ajax."}, {"name":"proxy","type":"string","required":false,"description":"Specify an IP address of a [custom proxy](https://docs.diffbot.com/reference/using-proxies#how-to-use-proxies) that will be used to fetch the target page. (Ex: `&proxy` or `..., {"name":"proxyAuth","type":"string","required":false,"description":"Used to specify the authentication parameters that will be used with a custom proxy specified in the &proxy parameter. (Ex: `proxyAuth=username:password`)"}, {"name":"useProxy","type":"string","required":false,"description":"Set to `default` to use [Diffbot's datacenter proxy](https://docs.diffbot.com/reference/using-proxies#how-to-use-proxies) for this request. `none` will instruct Extract t..., {"name":"paging","type":"boolean","required":false,"description":"Pass `paging=false` to disable automatic concatenation multiple-page articles."}
- Output summary
- Not available
- Billing
- 1
- Freshness
- Not available
- Verification
- Not available