Sogou web spider Web Robot

A summary of the Sogou web spider Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.

Who owns the Sogou web spider robot? Is it a good or a bad robot? And why is it visiting your website?

Shown below is a sample log file entry for the Sogou web spider web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.

Server Log File

vntweb.co.uk 123.126.113.174 - - [23/Mar/2019:13:23:24 +0000] "GET /robots.txt HTTP/1.1" 301 323 "-" "Sogou web spider/4.0(+http://www.sogou.com/docs/help/webmasters.htm#07)"

HTTP User Agent

Sogou web spider/4.0

IP Addresses

The observed IP address was 123.126.113.174.

WHOIS DNS command gives the following information about the IP address:

inetnum:123.112.0.0 – 123.127.255.255
netname:UNICOM-BJ
descr:China Unicom Beijing province network
address:No.21,Financial Street
address:Beijing,100033
address:P.R.China
last-modified:2017-10-23T05:59:13Z

As can be seen from the above the observed IP address is a part of a block assigned to China Unicom.

nslookup DNS command gives

174.113.126.123.in-addr.arpa name = sogouspider-123-126-113-174.crawl.sogou.com.

Owner

Sogou

Country

P.R.China

Exclusion

The user-agent string includes a reference to the website http://www.sogou.com/docs/help/webmasters.htm#07.

The referenced website confirms that the bot supports the robots exclusion text and also obeys the crawl delay.

The Sogou website provides details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1.

However, details are not given about both configuring the robots.txt file for their robot to prevent it visiting all or a part of your website and to limit the crawl rate.

You may wish to try adding the following in the robots.txt file to block the Sogou web spider:

User-agent: Sogou web spider
Disallow: / 

Also to control the frequency of Sogou web spider visiting your site try adding the following to seta minimum acceptable delay between consecutive requests

User-agent: Sogou web spider
Crawl-Delay: 10

In this example the delay has been set to 10 seconds.

As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.

Take care making changes to the robots.txt file. A misunderstanding in configuration or an error in configuration can lead to important search engines excluding your website.

Further Info

Sogou is a Chinese search engine, with the website claiming that it has become the world’s first search engine with a Chinese webpage volume of 10 billion.

Sogou has applications aimed at both the web and the desktop.

The web application offers web search including specialisation of music, pictures, videos, news, and maps.

The desktop application provides Pinyin input and has dual-core browser.