W3C_Validator Web Robot

A summary of the W3C_Validator Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the W3C_Validator robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the W3C_Validator web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

vntweb.co.uk 128.30.52.72 - - [08/Apr/2019:07:56:10 +0100] "GET / HTTP/1.1" 301 313 "-" "W3C_Validator/1.3 http://validator.w3.org/services"

HTTP User Agent

W3C_Validator/1.3

IP Addresses

The observed IP address was 128.30.52.72.

WHOIS DNS command gives the following information about the IP address:

NetRange:128.30.0.0 – 128.30.255.255
NetName:MIT-NET
Organization:Massachusetts Institute of Technology (MIT-2)
Address:Room W92-167
Address:77 Massachusetts Avenue
City:Cambridge
StateProv::MA
PostalCode:02139-4307
Country:US
Updated:2017-01-28

As can be seen from the above the observed IP address is a part of a block assigned to

Owner

W3C

Country

USA

Exclusion

The user-agent string includes a reference to the website http://validator.w3.org/services.

The referenced website doesn’t confirm that the bot supports the robots exclusion text and also obeys the crawl delay.

The suggestion is to implement exclusion via firewall configuration based upon eithe the IP address or the user agent.

Further Info

The W3C Validator Internet robot is triggered to visit a website by visiting the validation service run by W3C at http://validator.w3.org/

The validator can be used to minimise the errors within a website. For a website comprising a number theird party components it may not be possible to acheive an error free analysis.

More information is available about the Nu HTML Checker on a dedicated web page http://validator.w3.org/about.