Qwantify/Bleriot Web Robot

A summary of the Qwantify/Bleriot Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the Qwantify/Bleriot robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the Qwantify/Bleriot web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

vntweb.co.uk 91.242.162.85 - - [08/Apr/2019:14:17:55 +0100] "GET /robots.txt HTTP/1.1" 200 119 "-" "Mozilla/5.0 (compatible; Qwantify/Bleriot/1.1; +https://help.qwant.com/bot)"

HTTP User Agent

Qwantify/Bleriot/1.1

IP Addresses

The observed IP address was 91.242.162.85.

WHOIS DNS command gives the following information about the IP address:

inetnum:91.242.162.0 – 91.242.162.255
netnum:QWANT-NET
org-name:QWANT SAS
address:Qwant SAS, 28 Rue de l’Universite, 75007 Paris, FR
country:FR
last-modified:2015-08-20T09:59:58Z

As can be seen from the above the observed IP address is a part of a block assigned to Qwant SAS.

nslookup DNS command gives

** server can’t find 85.162.242.91.in-addr.arpa: NXDOMAIN

Owner

Qwant SAS

Country

France

Verification

The Slack website confirms example uses of their robot user-agent string:

Mozilla/5.0 (compatible; Qwantify/Bleriot/1.1; +https://help.qwant.com/bot)

Exclusion

The user-agent string includes a reference to the website https://help.qwant.com/bot.

The referenced website confirms that the bot supports the robots exclusion text and also obeys the crawl delay.

The Qwant website provides details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1.

Details are given about both preventing the robot from indexing the website and how to adjust its crawl rate.

Their advice is to include the following entry in the robots.txt file to prevent Qwantify/Bleriot from visiting your site.

User-agent: Bleriot
Disallow: / 

Also to control the frequency of Qwantify/Bleriot visiting your site, setting a minimum acceptable delay between consecutive requests can be set with the following added to the robots.txt file:

User-agent: Bleriot
Crawl-Delay: 10

In this example the delay has been set to 10 seconds. If not set the crawl delay has a default of 2 seconds.

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

On this page there are details about the origin of the web crawler which Qwant use. It has been developed from the opensource project by Qwant Research named Mermoz. The code for this can be downloaded from GitHub.

The site also confirms that the name of the bot will be changed to Bleriot.

A confirmation of the user agent to be used in the robots.txt file is not given on the referenced page. I have used a reference to their revised naming. You may need to experiment with alternatives or contact them (details on the page).

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

Qwant is a search engine alternative to Google, Bing, DuckDuckGo. At the time of writing it was the default search engine for the Brave browser.

The Qwant website has an overview of the history, ethos and operation.

More References

Bloomberg’s view of Qwant SAS https://www.bloomberg.com/research/stocks/private/snapshot.asp?privcapId=217713343

A wikipedia summary page https://en.wikipedia.org/wiki/Qwant