A summary of the bl.uk_ldfc_bot Internet robot. Including details for the owner, description, HTTP user agent and whether this robot adheres to the robot exclusion standard.
Who owns the bl.uk_ldfc_bot robot? Is it a good or a bad robot? And why is it visiting your website?
Shown below is a sample log file entry for the bl.uk_ldfc_bot web robot. It’s derived from an Apache web server log file. From the log entry information about how the robot identifies itself, HTTP User Agent, and where it is hosted are given.
Server Log File
www.vntweb.co.uk 194.66.232.89 - - [11/Aug/2019:16:30:53 +0100] "GET /robots.txt HTTP/1.1" 200 116 "-" "bl.uk_ldfc_bot/3.4.0-20190418 (+http://www.bl.uk/aboutus/legaldeposit/websites/websites/faqswebmaster/index.html)"
HTTP User Agent
bl.uk_ldfc_bot/3.4.0-20190418
IP Addresses
The observed IP address was 194.66.232.89.
WHOIS DNS command gives the following information about the IP address:
| inetnum: | 194.66.224.0 – 194.66.239.255 |
| netname: | BRIT-LIBRARY |
| descr: | British Library |
| address: | The British Library |
| address: | 96 Euston Road |
| address: | London NW1 2DB |
| address: | United Kingdom |
| country: | GB |
| last-modified: | 2016-06-01T14:22:49Z |
As can be seen from the above the observed IP address is a part of a block assigned to British Library via JANET (UK).
Owner
British Library
Country
UK
Exclusion
The user-agent string includes a reference to the website www.bl.uk/aboutus/legaldeposit/websites/websites/faqswebmaster/index.html. This link redirects to the page www.bl.uk/legal-deposit/web-archiving# which provides information for webmasters about the use of the robot.
The referenced website confirms that the bot supports the robots exclusion text but makes no mention of the crawl delay, but does make mention of following politeness rules to avoid a harmful impact upon the performance of the crawled website.
The website provides details about how their robot complies with the robots.txt exclusion standard, which was described at http://www.robotstxt.org/wc/exclusion.html#robotstxt, but is not currently available. Information is available on the same website https://www.robotstxt.org/robotstxt.html and also on the w3c website at https://www.w3.org/TR/html4/appendix/notes.html#h-B.4.1.1.
Details are given about preventing the robot from indexing the website but no mention about how to adjust its crawl rate.
No specific details are given about adding a reference to their robot in the robots.txt file to prevent bl.uk_ldfc_bot from visiting your site
User-agent: bl.uk_ldfc_bot Disallow: /
You may wish to trial controlling the frequency of bl.uk_ldfc_bot visiting your site, setting a minimum acceptable delay between consecutive requests can be set with the following added to the robots.txt file:
User-agent: bl.uk_ldfc_bot Crawl-Delay: 10
In this example the delay has been set to 10 seconds.
As is common with website crawlers there is a delay between changes made to the robots.txt file and the change being implemented.
Further Info
The British Library collects UK published material from the Internet under legal deposit.
UK information is defined by the use of a UK based server, domain or contact information.
More information about the collection is to be found on their UK Web Archive page.


