ia_archiver Web Robot

A summary of the ia_archiver Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the ia_archiver robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the ia_archiver web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

vntweb.co.uk 54.165.90.203 - - [31/Mar/2019:13:28:38 +0100] "GET /robots.txt HTTP/1.0" 301 323 "-" "ia_archiver"

HTTP User Agent

ia_archiver

IP Addresses

The observed IP address was 54.165.90.203.

WHOIS DNS command gives the following information about the IP address:

NetRange:54.160.0.0 – 54.175.255.255
NetName:AMAZON-2011L
Organization:Amazon Technologies Inc.
Address:410 Terry Ave N.
City:Seattle
StateProv::WA
PostalCode:98109
Country:US
Updated:2017-01-28

As can be seen from the above the observed IP address is a part of a block assigned to Amazon Technologies Inc.

Owner

Amazon

Country

US

Exclusion

The user-agent string doesn’t include a reference to a website. However, there is a support posting which provides information regards the crawling of the robot and its standards support.

https://support.alexa.com/hc/en-us/articles/200450194-Alexa-s-Web-and-Site-Audit-Crawlers

The website linked above provides details about how the ia_archiver robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1

Details are given about both preventing the robot from indexing the website but not how to adjust its crawl rate.

Their advice is to include the following entry in the robots.txt file to prevent ia_archiver from visiting your site

User-agent: ia_archiver
Disallow: / 

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

The Alexa robot is used to gather information for the Alexa website SEO analysis service.

There are further links provided about the Alexa robot and supporting references for the identification of the robot and which IP addresses are used.