tracemyfile Web Robot

A summary of the tracemyfile Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the tracemyfile robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the tracemyfile web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

www.vntweb.co.uk 91.106.199.104 - - [31/Aug/2019:16:50:34 +0100] "GET /robots.txt HTTP/1.1" 200 114 "-" "Mozilla/5.0 (compatible; tracemyfile/1.0; +bot@tracemyfile.com)"

HTTP User Agent

tracemyfile/1.0

IP Addresses

The observed IP address was 91.106.199.104.

WHOIS DNS command gives the following information about the IP address:

inetnum:91.106.199.0 – 91.106.199.255
netname:CNH-CC-KNA1-05
address:City Network Hosting AB
address:Borgmastaregatan 18
address:SE-371 34 Karlskrona
Country:SE
Updated:2013-12-09T14:43:43Z

As can be seen from the above the observed IP address is a part of a block assigned to Amazon Technologies Inc.

Owner

Tracemyfile Ltd

Country

UK

Exclusion

The user-agent string includes a reference to the website https://tracemyfile.com.

The referenced website doesn’t provide information about their web robot. About whether the bot supports the robots exclusion text and obeys the crawl delay.

The Tracemyfile website provides NO details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1

Details are given about preventing the robot from indexing the website but no mention about how to adjust its crawl rate.

with no details given, you may wish to try the following entry in the robots.txt file to prevent tracemyfile from visiting your site

User-agent: tracemyfile
Disallow: / 

You may wish to trial controlling the frequency of tracemyfile visiting your site, setting a minimum acceptable delay between consecutive requests can be set with the following added to the robots.txt file:

User-agent: tracemyfile
Crawl-Delay: 10

In this example the delay has been set to 10 seconds.

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

With Tracemyfile its possible to find who is making use of your files.

On the home page there is the option to upload an image file. This is then checked against their database of images to provide a report on its usage.

For a matching image the list of instances is shown. details include the website; country; last observed; and whether the image has been modified.