MJ12bot Web Robot

A summary of the MJ12bot Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the MJ12bot robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the MJ12bot web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

vntweb.co.uk 162.210.196.98 - - [19/Mar/2019:03:14:34 +0000] "GET /robots.txt HTTP/1.1" 301 323 "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; http://mj12bot.com/)"

HTTP User Agent

MJ12bot/v1.4.8

IP Addresses

The observed IP address was 162.210.196.98.

WHOIS DNS command gives the following information about the IP address:

NetRange:162.210.192.0 – 162.210.199.255
OrgName:Leaseweb USA, Inc.
Address:9480 Innovation Dr
City:Manassas
StateProv:VA
PostalCode:20109
Country:US
Updated:2017-01-28

As can be seen from the above the observed IP address is a part of a block assigned to Lease Web.

Owner

Majestic

Country

UK

Exclusion

The user-agent string includes a reference to the website http://mj12bot.com/. A whole domain for details about the robot.

The referenced website confirms that the bot supports the robots exclusion text and also obeys the crawl delay.

The Majestic website provides details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1.

Details are given about both preventing the robot from indexing the website and how to adjust its crawl rate.

Their advice is to include the following entry in the robots.txt file to prevent MJ12bot from visiting your site

User-agent: MJ12bot
Disallow: / 

Also to control the frequency of MJ12bot visiting your site, setting a minimum acceptable delay between consecutive requests can be set with the following added to the robots.txt file:

User-agent: MJ12bot
Crawl-Delay: 5

In this example, taken from their website the delay has been set to 5 seconds.

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

The website confirms that the bot belongs to Majestic, a UK based specialist search engine.

The robot is used to map the relationships between websites in support of the Majestic website’s link database.

The Majestic website provides a search engine for: SEO professionals; media analysts; entrepreneurs and developers.

A number of tools are provided including:

  • Site explorer – explore a domain or url in detail with this tool
  • Backlink history checker – determine the number of backlinks determined for given domains and subdomains.
  • Search explorer – search their index for keywords to see the title and url where it appears.
  • Link intelligence API – the API allows raw data to be retrieved

More information is available on the help centre page of the Majestic website.