Diffbot Web Robot

A summary of the Diffbot Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the Diffbot robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the Diffbot web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

www.vntweb.co.uk 35.226.243.86 - - [03/Sep/2019:20:08:54 +0100] "GET /wp-content/themes/vntweb-17/images/vntweblogo_wht2.png HTTP/1.1" 404 2076 "-" "Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.1.2) Gecko/20090729 Firefox/3.5.2 (.NET CLR 3.5.30729; Diffbot/0.1; +http://www.diffbot.com)"

HTTP User Agent

Diffbot/0.1

IP Addresses

The observed IP address was 35.226.243.86.

WHOIS DNS command gives the following information about the IP address:

NetRange:35.208.0.0 – 35.247.255.255
NetName:GOOGLE-CLOUD
OrgName:Google LLC
Address:1600 Amphitheatre Parkway
City:Mountain View
StateProv:CA
PostalCode:94043
Country:US
Updated:2017-12-21

As can be seen from the above the observed IP address is a part of a block assigned to Google Cloud.

Owner

Diffbot Technologies Corp.

Country

USA

Exclusion

The user-agent string includes a reference to the website http://www.diffbot.com.

The referenced website doesn’t provide information about the robot and whether it supports the robots exclusion text and also obeys the crawl delay.

The Diffbot website provides no details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1

Details are given about both preventing the robot from indexing the website and how to adjust its crawl rate.

Their advice is to include the following entry in the robots.txt file to prevent DotBot from visiting your site

User-agent: Diffbot
Disallow: / 

Also to control the frequency of Diffbot visiting your site, setting a minimum acceptable delay between consecutive requests can be set with the following added to the robots.txt file:

User-agent: Diffbot
Crawl-Delay: 10

In this example, taken from their website the delay has been set to 10 seconds.

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

Diffbot allows web pages to be extracted as structured data.

Products:

  • Extraction APIs
  • Crawlbot
  • knowledge Graph

Diffbot provides the following examples of uses

  • Sales and Marketing
    Assisting the closure of more deals and level up up marketing with better data.
  • Business intelligence
    helping to produce meaningful insights with the Knowledge Graph derived accurate, complete and deep data.
  • Recruiting
    aids the sourcing of ideal candidates