OnPage API Resources

‌‌
This endpoint will provide you with a list of resources, including images, scripts, stylesheets, and broken elements.
You will get a detailed overview of every resource found on the crawled pages.

If you would like to receive a list of pages that contain a specific resource, please refer to the Pages By Resource endpoint.

checked POST
Pricing

Your account will not be charged for using this function. You can get the results of the task within the next 30 days for free.
The cost can be calculated on the Pricing page.

All POST data should be sent in the JSON format (UTF-8 encoding). The task setting is done using the POST method. When setting a task, you should send all task parameters in the task array of the generic POST array.

Description of the fields for setting a task:

Field nameTypeDescription
idstring

ID of the task
required field
you can get this ID in the response of the Task POST endpoint
example:
"07131248-1535-0216-1000-17384017ad04"

urlstring

page URL
optional field
specify this field if you want to get the resources for a specific page
note that to obtain resource's meta from a particular URL, you should specify the URL in this field;
if you do not indicate a url when setting a task, resource's meta in the results will be returned based on the data from the page where our crawler first saw the resource

limitinteger

the maximum number of returned resources
optional field
default value: 100
maximum value: 1000

offsetinteger

offset in the results array of returned resources
optional field
default value: 0
maximum value: 2000000
if you specify the 10 value, the first ten resources in the results array will be omitted and the data will be provided for the successive resources

filtersarray

array of results filtering parameters
optional field
you can add several filters at once (8 filters maximum)
you should set a logical operator and, or between the conditions
the following operators are supported:
regex, not_regex, <, <=, >, >=, =, <>, in, not_in, like, not_like
you can use the % operator with like and not_like to match any string of zero or more characters
example:
["resource_type","=","stylesheet"]

[["resource_type","=","image"],
"and",["checks.is_https","=",false]]

[["fetch_timing.duration_time",">",1],"and",[["total_transfer_size",">",100],"or",["checks.high_loading_time","=",true]]]

The full list of possible filters is available by this link.

relevant_pages_filtersarray

filter the resources by relevant pages
optional field
you can use this field to obtain resources from pages matching to the defined parameters
you can apply the same filters here as available for the pages endpoint
you can add several filters at once (8 filters maximum)
you should set a logical operator and, or between the conditions
the following operators are supported:
regex, not_regex, <, <=, >, >=, =, <>, in, not_in, like, not_like
you can use the % operator with like and not_like to match any string of zero or more characters
example:
["checks.no_image_title","=",true]

order_byarray

results sorting rules
optional field
you can use the same values as in the filters array to sort the results
possible sorting types:
asc - results will be sorted in the ascending order
desc - results will be sorted in the descending order
you should use a comma to set up a sorting type
example:
["size,desc"]
note that you can set no more than three sorting rules in a single request
you should use a comma to separate several sorting rules
example:
["size,desc","fetch_timing.fetch_end,desc"]

search_after_tokenstring

token for subsequent requests
optional field
provided in the identical filed of the response to each request;
use this parameter to avoid timeouts while trying to obtain over 20,000 results in a single request;
by specifying the unique search_after_token value from the response array, you will get the subsequent results of the initial task;
search_after_token values are unique for each subsequent task ;
Note: if the search_after_token is specified in the request, all other parameters should be identical to the previous request

tagstring

user-defined task identifier
optional field
the character limit is 255
you can use this parameter to identify the task and match it with the result
you will find the specified tag value in the data object of the response


‌‌‌‌‌‌
As a response of the API server, you will receive JSON-encoded data containing a tasks array with the information specific to the set tasks.

Description of the fields in the results array:

Field nameTypeDescription
versionstring

the current version of the API

status_codeinteger

general status code
you can find the full list of the response codes here
Note: we strongly recommend designing a necessary system for handling related exceptional or error conditions

status_messagestring

general informational message
you can find the full list of general informational messages here

timestring

execution time, seconds

costfloat

total tasks cost, USD

tasks_countinteger

the number of tasks in the tasks array

tasks_errorinteger

the number of tasks in the tasks array returned with an error

tasksarray

array of tasks

    idstring

task identifier
unique task identifier in our system in the UUID format

    status_codeinteger

status code of the task
generated by DataForSEO; can be within the following range: 10000-60000
you can find the full list of the response codes here

    status_messagestring

informational message of the task
you can find the full list of general informational messages here

    timestring

execution time, seconds

    costfloat

cost of the task, USD

    result_countinteger

number of elements in the result array

    patharray

URL path

    dataobject

contains the same parameters that you specified in the POST request

    resultarray

array of results

        crawl_progressstring

status of the crawling session
possible values: in_progress, finished

        crawl_statusobject

details of the crawling session

            max_crawl_pagesinteger

maximum number of pages to crawl
indicates the max_crawl_pages limit you specified when setting a task

            pages_in_queueinteger

number of pages that are currently in the crawling queue

            pages_crawledinteger

number of crawled pages

        total_items_countinteger

total number of relevant items crawled

        items_countinteger

number of items in the results array

        itemsarray

items array

            resource_typestring

type of the returned resource
possible types: script, image, stylesheet, broken

            metaobject

resource properties
the value depends on the resource_type
note that if you do not indicate a url when setting a task, resource's meta is returned based on the data from the page where our crawler first saw the resource;
to obtain resource's meta from a particular url, specify that URL when setting a task

                alternative_textstring

content of the image alt attribute
the value depends on the resource_type

                titlestring

title

                original_widthinteger

original image width in px

                original_heightinteger

original image height in px

                widthinteger

image width in px

                heightinteger

image height in px

            status_codeinteger

status code of the page where a given resource is located

            locationstring

location header
indicates the URL to redirect a page to

            urlstring

resource URL

            sizeinteger

resource size
indicates the size of a given resource measured in bytes

            encoded_sizeinteger

resource size after encoding
indicates the size of the encoded resource measured in bytes

            total_transfer_sizeinteger

compressed resource size
indicates the compressed size of a given resource in bytes

            fetch_timestring

date and time when a resource was fetched
in the UTC format: “yyyy-mm-dd hh-mm-ss +00:00”
example:
2021-02-17 13:54:15 +00:00

            fetch_timingobject

resource fething time range

                duration_timeinteger

indicates how many milliseconds it took to fetch a resource

                fetch_startinteger

time to start downloading the resource
the amount of time a browser needs to start downloading a resource

                fetch_endinteger

time to complete downloading the resource
the amount of time a browser needs to complete downloading a resource

            cache_controlobject

instructions for caching

                cachableboolean

indicates whether the resource is cacheable

                ttlinteger

time to live
the amount of time it takes for the browser to cache a resource; measured in milliseconds

            checksobject

resource check-ups
contents of the array depend on the resource_type

                no_content_encodingboolean

resource with no content encoding
indicates whether a page has no compression algorithm of the content;
available for items with the following resource_type: script, image, stylesheet, broken

                high_loading_timeboolean

resource with high loading time
indicates whether a resource loading time exceeds 3 seconds;
available for items with the following resource_type: script, image, stylesheet, broken

                is_redirectboolean

resource with redirects
indicates whether a page with this resource has 3XX redirects to other pages;
available for items with the following resource_type: script, image, stylesheet, broken

                is_4xx_codeboolean

resource with with 4xx status code
indicates whether a page with this resource has 4XX response code

                is_5xx_codeboolean

resource with 5xx status code
indicates whethera page with this resource has 5XX response code

                is_brokenboolean

broken resource
indicates whether a page with this resource returns 4xx, 5xx response codes or has broken elements inside the resource;
available for items with the following resource_type: script, image, stylesheet, broken

                is_wwwboolean

page with www
indicates whether a page with this resource is on a www subdomain;
available for items with the following resource_type: script, image, stylesheet, broken

                is_httpsboolean

page with the https protocol
available for items with the following resource_type: script, image, stylesheet, broken

                is_httpboolean

page with the http protocol
available for items with the following resource_type: script, image, stylesheet, broken

                original_size_displayedboolean

image desplayes in its original size
indicates whether the image is displayed in its original size;
available for items with the following resource_type: image

                is_minifiedboolean

resource is minified
indicates whether the content of a stylesheet or script is minified;
available for items with the following resource_type: stylesheet, script

                has_redirectboolean

resource has a redirect
available for items with the following resource_type: script, image;
if the resource_type is image, this field will indicate whether other pages and/or resources have redirects pointing at the image;
if the resource_type is script, this field will indicate whether the script contains a redirect

                has_subrequestsboolean

resource contains subrequests
indicates whether the content of a stylesheet or script contain additional requests;
available for items with the following resource_type: stylesheet, script

                from_sitemapboolean

resource was found on website's sitemap
if true, the resource was found on the sitemap of the website

            resource_errorsobject

resource errors and warnings

                errorsarray

resource errors

                    lineinteger

line where the error was found

                    columninteger

column where the error was found

                    messagestring

text message of the error
the full list of possible HTML errors can be found here

                    status_codeinteger

status code of the error
possible values:
0 — Unidentified Error;
501 — Html Parse Error;
1501 — JS Parse Error;
2501 — CSS Parse Error;
3501 — Image Parse Error;
3502 — Image Scale Is Zero;
3503 — Image Size Is Zero;
3504 — Image Format Invalid

                warningsarray

resource warnings

                    lineinteger

line the warning relates to
note that if "line": 0, the warning relates to the whole page

                    columninteger

column the warning relates to
note that if "column": 0, the warning relates to the whole page

                    messagestring

text message of the warning
possible messages:
"Has node with more than 60 childs." - HTML page has at least 1 tag nesting over 60 tags of the same level
"Has more that 1500 nodes." - DOM tree contains over 1,500 elements
"HTML depth more than 32 tags." - DOM depth exceeds 32 nodes

                    status_codeinteger

status code of the warning
possible values:
0 — Unidentified Warning;
1 — Has node with more than 60 childs;
2 — Has more that 1500 nodes;
3 — HTML depth more than 32 tags

                content_encodingstring

type of encoding

                media_typestring

types of media used to display a resource

                accept_typestring

indicates the expected type of resource
for example, if "resource_type": "broken", accept_type will indicate the type of the broken resource
possible values:
any, none, image, sitemap, robots, script, stylesheet, redirect, html, text, other, font

                serverstring

server version

                last_modifiedobject

contains data on changes related to the resource
if there is no data, the value will be null

                    headerstring

date and time when the header was last modified
in the UTC format: "yyyy-mm-dd hh-mm-ss +00:00"
example:
2019-11-15 12:57:46 +00:00
if there is no data, the value will be null

                    sitemapstring

date and time when the sitemap was last modified
in the UTC format: "yyyy-mm-dd hh-mm-ss +00:00"
example:
2019-11-15 12:57:46 +00:00
if there is no data, the value will be null

                    meta_tagstring

date and time when the meta tag was last modified
in the UTC format: "yyyy-mm-dd hh-mm-ss +00:00"
example:
2019-11-15 12:57:46 +00:00
if there is no data, the value will be null


‌‌

Instead of ‘login’ and ‘password’ use your credentials from https://app.dataforseo.com/api-access

# Instead of 'login' and 'password' use your credentials from https://app.dataforseo.com/api-access 
login="login" 
password="password" 
cred="$(printf ${login}:${password} | base64)" 
curl --location --request POST "https://api.dataforseo.com/v3/on_page/resources" 
--header "Authorization: Basic ${cred}"  
--header "Content-Type: application/json" 
--data-raw '[
  {
    "id": "07281559-0695-0216-0000-c269be8b7592",
    "filters": [
      ["resource_type", "=", "image"],
      "and",
      ["size", ">", 100000]
    ],
    "order_by": ["size,desc"],
    "limit": 10
  }
]'
<?php
// You can download this file from here https://cdn.dataforseo.com/v3/examples/php/php_RestClient.zip
require('RestClient.php');
$api_url = 'https://api.dataforseo.com/';
// Instead of 'login' and 'password' use your credentials from https://app.dataforseo.com/api-access
$client = new RestClient($api_url, null, 'login', 'password');

$post_array = array();
// simple way to get a result
$post_array[] = array(
   "id" => "07281559-0695-0216-0000-c269be8b7592",
   "filters" => [
      ["resource_type", "=", "image"],
      "and",
      ["size", ">", 100000]
   ],
   "order_by" => ["size,desc"],
   "limit" => 10
);
try {
   // POST /v3/on_page/resources
   // the full list of possible parameters is available in documentation
   $result = $client->post('/v3/on_page/resources', $post_array);
   print_r($result);
   // do something with post result
} catch (RestClientException $e) {
   echo "n";
   print "HTTP code: {$e->getHttpCode()}n";
   print "Error code: {$e->getCode()}n";
   print "Message: {$e->getMessage()}n";
   print  $e->getTraceAsString();
   echo "n";
}
$client = null;
?>
const post_array = [];

post_array.push({
  "id": "07281559-0695-0216-0000-c269be8b7592",
  "filters": [
    ["resource_type", "=", "image"],
    "and",
    ["size", ">", 100000]
  ],
  "order_by": ["size,desc"],
  "limit": 10
});

const axios = require('axios');

axios({
  method: 'post',
  url: 'https://api.dataforseo.com/v3/on_page/resources',
  auth: {
    username: 'login',
    password: 'password'
  },
  data: post_array,
  headers: {
    'content-type': 'application/json'
  }
}).then(function (response) {
  var result = response['data']['tasks'];
  // Result data
  console.log(result);
}).catch(function (error) {
  console.log(error);
});
from random import Random
from client import RestClient
# You can download this file from here https://api.dataforseo.com/v3/_examples/python/_python_Client.zip
client = RestClient("login", "password")

post_data = dict()
# simple way to get a result
post_data[len(post_data)] = dict(
    id="07281559-0695-0216-0000-c269be8b7592",
    filters=[
        ["resource_type", "=", "image"],
        "and", 
        ["size", ">", 100000]
    ],
    order_by=["size,desc"],
    limit=10
)
# POST /v3/on_page/resources
# the full list of possible parameters is available in documentation
response = client.post("/v3/on_page/resources", post_data)
# you can find the full list of the response codes here https://docs.dataforseo.com/v3/appendix/errors
if response["status_code"] == 20000:
    print(response)
    # do something with result
else:
    print("error. Code: %d Message: %s" % (response["status_code"], response["status_message"]))
using Newtonsoft.Json;
using System;
using System.Collections.Generic;
using System.Net.Http;
using System.Net.Http.Headers;
using System.Text;
using System.Threading.Tasks;

namespace DataForSeoDemos
{
    public static partial class Demos
    {
        public static async Task on_page_resources()
        {
            var httpClient = new HttpClient
            {
                BaseAddress = new Uri("https://api.dataforseo.com/"),
                // Instead of 'login' and 'password' use your credentials from https://app.dataforseo.com/api-access
                DefaultRequestHeaders = { Authorization = new AuthenticationHeaderValue("Basic", Convert.ToBase64String(Encoding.ASCII.GetBytes("login:password"))) }
            };
            var postData = new List<object>();
            // simple way to get a result
            postData.Add(new
            {
                id = "07281559-0695-0216-0000-c269be8b7592",
                filters = new object[]
                {
                    new object[] { "resource_type", "=", "image" },
                    "and",
                    new object[] { "size", ">", 100000 }
                },
                order_by = new object[] { "size,desc" },
                limit = 10
            });
            // POST /v3/on_page/resources
            // the full list of possible parameters is available in documentation
            var taskPostResponse = await httpClient.PostAsync("/v3/on_page/resources", new StringContent(JsonConvert.SerializeObject(postData)));
            var result = JsonConvert.DeserializeObject<dynamic>(await taskPostResponse.Content.ReadAsStringAsync());
            // you can find the full list of the response codes here https://docs.dataforseo.com/v3/appendix/errors
            if (result.status_code == 20000)
            {
                // do something with result
                Console.WriteLine(result);
            }
            else
                Console.WriteLine($"error. Code: {result.status_code} Message: {result.status_message}");
        }
    }
}

The above command returns JSON structured like this:

{
  "version": "0.1.20200805",
  "status_code": 20000,
  "status_message": "Ok.",
  "time": "4.8323 sec.",
  "cost": 0,
  "tasks_count": 1,
  "tasks_error": 0,
  "tasks": [
    {
      "id": "08091838-1535-0216-0000-cee52596d188",
      "status_code": 20000,
      "status_message": "Ok.",
      "time": "4.7640 sec.",
      "cost": 0,
      "result_count": 1,
      "path": [
        "v3",
        "on_page",
        "resources"
      ],
      "data": {
        "api": "on_page",
        "function": "resources",
        "limit": 100
      },
      "result": [
        {
          "crawl_progress": "finished",
          "crawl_status": {
            "max_crawl_pages": 10,
            "pages_in_queue": 0,
            "pages_crawled": 10
          },
          "total_items_count": 15,
          "items_count": 100,
          "items": [
            {
              "resource_type": "stylesheet",
              "meta": null,
              "status_code": 200,
              "location": null,
              "url": "https://dataforseo.com/wp-content/themes/startit/assets/css/style_dynamic.css?ver=1566565352",
              "size": 1216,
              "encoded_size": 0,
              "total_transfer_size": 275,
              "fetch_time": "2021-02-17 13:54:15 +00:00",
              "fetch_timing": {
                "duration_time": 0,
                "fetch_start": 0,
                "fetch_end": 0
              },
              "cache_control": {
                "cachable": false,
                "ttl": 0
              },
              "checks": {
                "no_content_encoding": false,
                "high_loading_time": false,
                "is_redirect": false,
                "is_4xx_code": false,
                "is_5xx_code": false,
                "is_broken": false,
                "is_www": false,
                "is_https": true,
                "is_http": false,
                "is_minified": false,
                "has_subrequests": false
              },
              "content_encoding": "gzip",
              "media_type": "text/css",
              "accept_type": "stylesheet",
              "server": "nginx/1.10.1 (Ubuntu)",
              "last_modified": {
                "header": "2021-10-21 14:11:10 +00:00",
                "sitemap": null,
                "meta_tag": "2021-03-15 00:00:00 +00:00"
              }
            },
            {
              "resource_type": "script",
              "meta": null,
              "status_code": 200,
              "location": null,
              "url": "https://dataforseo.com/wp-includes/js/jquery/jquery-migrate.min.js?ver=1.4.1",
              "size": 10056,
              "encoded_size": 0,
              "total_transfer_size": 385,
              "fetch_time": "2021-02-17 13:54:15 +00:00",
              "fetch_timing": {
                "duration_time": 0,
                "fetch_start": 0,
                "fetch_end": 0
              },
              "cache_control": {
                "cachable": true,
                "ttl": 2592000
              },
              "checks": {
                "no_content_encoding": false,
                "high_loading_time": false,
                "is_redirect": false,
                "is_4xx_code": false,
                "is_5xx_code": false,
                "is_broken": false,
                "is_www": false,
                "is_https": true,
                "is_http": false,
                "is_minified": false,
                "has_redirect": false,
                "has_subrequests": false,
                "from_sitemap": false
              },
              "content_encoding": "gzip",
              "media_type": "application/javascript",
              "accept_type": "script",
              "server": "nginx/1.10.1 (Ubuntu)",
              "last_modified": {
                "header": "2021-10-21 14:11:10 +00:00",
                "sitemap": null,
                "meta_tag": "2021-03-15 00:00:00 +00:00"
              }
            },
            {
              "resource_type": "image",
              "meta": {
                "alternative_text": "sean-cooney-review",
                "title": null,
                "original_width": 250,
                "original_height": 250,
                "width": 250,
                "height": 250
              },
              "status_code": 200,
              "location": null,
              "url": "https://dataforseo.com/wp-content/uploads/2020/06/sean-cooney-review-1.png",
              "size": 94695,
              "encoded_size": 94695,
              "total_transfer_size": 95036,
              "fetch_time": "2021-02-17 13:54:15 +00:00",
              "fetch_timing": {
                "duration_time": 0,
                "fetch_start": 0,
                "fetch_end": 0
              },
              "cache_control": {
                "cachable": true,
                "ttl": 2592000
              },
              "checks": {
                "no_content_encoding": true,
                "high_loading_time": false,
                "is_redirect": false,
                "is_4xx_code": false,
                "is_5xx_code": false,
                "is_broken": false,
                "is_www": false,
                "is_https": true,
                "is_http": false,
                "has_redirect": false,
                "original_size_displayed": true
              },
              "content_encoding": null,
              "media_type": "image/png",
              "accept_type": "image",
              "server": "nginx/1.10.1 (Ubuntu)",
              "last_modified": {
                "header": "2021-10-21 14:11:10 +00:00",
                "sitemap": null,
                "meta_tag": "2021-03-15 00:00:00 +00:00"
              }
            },
            {
              "resource_type": "broken",
              "meta": null,
              "status_code": 404,
              "location": null,
              "url": "https://dataforseo.com/css/bootstrap.css",
              "size": 75924,
              "encoded_size": 0,
              "total_transfer_size": 464,
              "fetch_time": "2021-02-17 13:54:15 +00:00",
              "fetch_timing": {
                "duration_time": 0,
                "fetch_start": 0,
                "fetch_end": 0
              },
              "cache_control": {
                "cachable": true,
                "ttl": 0
              },
              "checks": {
                "no_content_encoding": false,
                "high_loading_time": false,
                "is_redirect": false,
                "is_4xx_code": true,
                "is_5xx_code": false,
                "is_broken": true,
                "is_www": false,
                "is_https": true,
                "is_http": false
              },
              "content_encoding": "gzip",
              "media_type": "text/html",
              "accept_type": "stylesheet",
              "server": "nginx/1.10.1 (Ubuntu)",
              "last_modified": {
                "header": "2021-10-21 14:11:10 +00:00",
                "sitemap": null,
                "meta_tag": "2021-03-15 00:00:00 +00:00"
              }
            }
          ]
        }
      ]
    }
  ]
}