A summary of the Gluten Free Crawler Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.
Who owns the Gluten Free Crawler robot? Is it a good or a bad robot? And why is it visiting your website?
Shown below is a sample log file entry for the Gluten Free Crawler web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.
Server Log File
www.vntweb.co.uk 104.131.147.112 - - [24/Jun/2019:15:16:17 +0100] "GET / HTTP/1.1" 200 9510 "-" "Mozilla/5.0 (compatible; Gluten Free Crawler/1.0; +http://glutenfreepleasure.com/)"
HTTP User Agent
Gluten Free Crawler/1.0
IP Addresses
The observed IP address was 104.131.147.112.
WHOIS DNS command gives the following information about the IP address:
| NetRange: | 104.131.0.0 – 104.131.255.255 |
| NetName: | DIGITALOCEAN-9 |
| OrgName: | DigitalOcean, LLC |
| Address: | 101 Ave of the Americas |
| City: | New York |
| StateProv: | NY |
| PostalCode: | 10013 |
| Country: | US |
| Updated: | 2019-02-04 |
As can be seen from the above the observed IP address is a part of a block assigned to Digital Ocean.
Owner
Josh
Country
USA
Exclusion
The user-agent string includes a reference to the website http://glutenfreepleasure.com/.
The referenced website doesn’t confirm whether the bot supports the robots exclusion text and also obeys the crawl delay.
The standard is described at http://www.robotstxt.org/wc/exclusion.html#robotstxt.
An email address is provided to contact the owner of the robot to either block it from visiting your website or to provide feedback about its operation.
Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.
Further Info
According to the website its a private crawling project using a handy domain name.


