VelenPublicWebCrawler Web Robot

A summary of the VelenPublicWebCrawler Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the VelenPublicWebCrawler robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the VelenPublicWebCrawler web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

vntweb.co.uk 51.75.5.39 - - [09/Apr/2019:16:19:23 +0100] "GET /slow-the-googlebot-crawl-rate HTTP/1.1" 301 334 "-" "Mozilla/5.0 (compatible; VelenPublicWebCrawler/1.0; +https://velen.io)"

HTTP User Agent

VelenPublicWebCrawler/1.0

IP Addresses

The observed IP address was 51.75.5.39.

WHOIS DNS command gives the following information about the IP address:

inetnum:51.75.4.0 – 51.75.7.255
netname:PCI-GRA5
org-name:OVH SAS
address:2 rue Kellermann
address:59100
address:Roubaix
address:FRANCE
last-modified:2017-10-30T14:40:06Z

As can be seen from the above the observed IP address is a part of a block assigned to OVH.

nslookup DNS command gives

39.5.75.51.in-addr.arpa name = ip-51-75-5.eu.

Owner

Unknown

Country

Unknown – maybe Germany

Exclusion

The user-agent string includes a reference to the website https://velen.io/ (no longer available in 2021).

The referenced website confirms that the bot supports the robots exclusion text and also obeys the crawl delay.

The Velen Crawler website provides details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1.

Whilst not providing examples of how to prevent the bot from crawling a website the user agent is confirmed as VelenPublicWebCrawler. Also the crawl rate for the bot is given as not more than one page at a time and 2 seconds apart.

The example below is to be included in the robots.txt file to prevent VelenPublicWebCrawler from visiting your site

User-agent: VelenPublicWebCrawler
Disallow: / 

The website confirms acceptance of the standards therefore to control the frequency of VelenPublicWebCrawler visiting your site, setting a minimum acceptable delay between consecutive requests can be set with the following added to the robots.txt file:

User-agent: VelenPublicWebCrawler
Crawl-Delay: 10

In this example the delay has been set to 10 seconds. The website doesn’t confirm the delay which may be configured. Some bots and crawlers have a maximum value.

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

The website referenced in the log file entry is a simple single page with details about the crawler and a contact email address.

The page says that the crawler is written in Go for Hunter.

The text presents this as a new project with the aim of analysing public internet pages to build business datasets and machine learning models with which to better understand the web.

Looking at the certificate for the website. It is shown as having a common name of ssl390662.cloudflaressl.com. The general title is Velen — Crawler last modified on 14th February 2019.