June 2, 201016 yr Hi I have a visitor analytics software for my website that tracks crawlers, and it seems that most crawlers visit two pages of my site most extensively: the index one and /robots.txt The thing is that I haven't created such a file for my website (so I guess that the crawlers just look for such a file and drop on a 404 error, right?). Another thing is that on the website of my hosters they say that it is recommended that I consider creating a robots.txt file. Now from what I know, the robots.txt file is for excluding all or certain crawlers to crawl specific pages of your site. But I don't want to block any crawler, so why would I make this file? So my questions are: what exactly is the robots.txt file? Does it have another purpose than excluding crawlers? Why do all crawlers look for this file on my website?
June 2, 201016 yr Author I guess I should post this question in the "SEO (Search Engine Optimisation) & Search Engines" section, sorry.
June 2, 201016 yr Most web spiders/bots will check for a robots.txt file which tells them what they shouldn't index. e.g User-agent: * Disallow: /cgi-bin/ Disallow: /private/ The robotstxt.org website will tell you everything you need to know.
June 2, 201016 yr Author I see, so should I make one then to include my exclusions like the example you gave? I took it that only the "Public" directory could be accessed by visitors/crawlers? Do I have to exclude all other directories manually by making a robotst.txt file? Thanks it would be good if you could give me quick answers for this question, but I'll check the link you gave me anyway.
June 3, 201016 yr I see, so should I make one then to include my exclusions like the example you gave? I took it that only the "Public" directory could be accessed by visitors/crawlers? Do I have to exclude all other directories manually by making a robotst.txt file? Thanks it would be good if you could give me quick answers for this question, but I'll check the link you gave me anyway. Robots will grab anything they can get their hands on within your http root folder (public_html on apache, wwwroot on iis) if you don't specifically exclude them. If you don't want to exclude anything within those folders then do the following code to allow all to access everything. User-agent: * Disallow: Even if you're not excluding anything its still worth having a robots.txt because the behaviour of web robots is undefined when they receive a 404, some carry on regardless, but some will blacklist your entire site.
June 9, 201016 yr By all means create a robots.txt file, but be careful which directories you put in the robots.txt. Anyone can see it and in some cases it can just point someone directly to the folder you don't want them to see. Make sure the folder is not browseable -- a quick fix (while you work on a proper secure solution) is to put a blank index.html file in any directory you don't want browsed (images, stats, etc). ---- St Albans Web Design Web Design, Development and Email Marketing
Create an account or sign in to comment