MojeekBot Web Robot

A summary of the MojeekBot Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the MojeekBot robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the MojeekBot web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

vntweb.co.uk 5.102.173.71 - - [19/Mar/2019:12:33:46 +0000] "GET /how-to-add-target-blank-to-wordpress-menu-custom-link/ HTTP/1.1" 200 8164 "-" "Mozilla/5.0 (compatible; MojeekBot/0.6; +https://www.mojeek.com/bot.html)"

HTTP User Agent

MojeekBot/0.6

IP Addresses

The observed IP address was 5.102.173.71.

WHOIS DNS command gives the following information about the IP address:

inetnum:5.102.173.64 – 5.102.173.79
org-name:CUSTDC_CLIENT_MOJEEK
descr:Mojeek
address:Custodian DataCentre Ltd.
address:Maidstone TV Studios
address:Maidstone TV Studios
address:UNITED KINGDOM
last-modified:2013-07-18T14:02:26Z

As can be seen from the above the observed IP address is a part of a block assigned to CustodianDC Ltd.

nslookup DNS command gives

71.173.102.5.in-addr.arpa name = crawl-5-102-173-71.mojeek.com.

Owner

Mojeek Ltd.

Country

UK

Verification

The Mojeek website confirms the IP address is confirmed as 5.102.173.71.

Exclusion

The user-agent string includes a reference to the websitehttps://www.mojeek.com/bot.html.

The referenced website confirms that the bot supports the robots exclusion text and also obeys the crawl delay.

The Mojeek website provides details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1.

Details are given about both preventing robot from indexing the website and how to verify that it is the MojeekBot visiting your site. A maximum crawl rate is defined as not more than 1 page per 4 seconds.

Whilst their website doesn’t provide information about specifically blocking their robot, it does give generic configuration details. You may wish to try including the following entry in the robots.txt file to prevent MojeekBot from visiting your site

User-agent: MojeekBot
Disallow: / 

Also to control the frequency of MojeekBot visiting your site, setting a minimum acceptable delay between consecutive requests can be set. Specific details are not provide, you may wish to try the following added to the robots.txt file:

User-agent: MojeekBot
Crawl-Delay: 10

In this example, the delay has been set to 10 seconds. Some robots have a maximum delay value. This information isn’t provided on the robot’s web page.

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

Mojeek was created in 2004 as a global alternative search engine based in the UK.

The software which with its own unique code, not derived from pre-existing software.

Mojeek doesn’t track its users and the results retrieved are its own not retrieved from another search engine.