Forum Moderators: open
looksmart-sv-fw.looksmart.com - - [08/Jan/2003:15:12:35 -0800] "GET /1Reference.html HTTP/1.1" 403 225 "-" "grub-client"
looksmart-sv-fw.looksmart.com - - [08/Jan/2003:15:12:35 -0800] "GET /1Reference.html HTTP/1.1" 403 225 "-" "grub-client"
So, if one bans all Grubs/grub-clients then how would one allow Looksmart?
My robots.txt allows all.
Pendanticist.
what do you mean?
As I've been led to believe, Grubs are inherantly bad bots that download/scrape or otherwise 'take' your material.
Do a site search for that nearly perfect .htaccess ban list and you'll find Grub(s) right in there with the rest of them.
Naturally, I was surprised to see LS using such a commonly known-to-be bad bot.
Pendanticist.
But there are too many different, outdated, and broken versions of their client software in circulation, and they don't seem to have built any way into the system to disable the old ones from their central server, and to force their users to upgrade.
I have also come to suspect that some unrelated bad bots have started to piggy back on their basically intact reputation (despite what some people say above), faking their UA. In the end, even though I have no problems with what they're trying to do, I still had to ban them.
Regarding grub-client, I'm not sure how honorable grub.org actually is. I've been in contact with Kord Campbell of Grub who assured me (after much pulling of teeth) that they would take my domain off of their search lists as of 6 Jan. Lo and behold, the grubs are coming again as of the 18th.
You need to read their policies real careful. Regardless of their intent to promote a well-searched web, they are also *selling* the content. As well, there's no assurance that a grub-client is actually their's.
Since it's distributed, blocking IPs is fruitless. Using an .htaccess method or integrated php method to block the agent itself is probably the best (if not only) way.
FWIW, I've emailed pendanticist, bird and fineware in an attempt to track down any issues that may be caused by faults in the grubclient crawler and the system that we run in general.
Pendanticist's Issue
--------------------
At this time, it appears that pendanticist may have our client, user-agent name grub-client, confused with the nicknames of other automated agents out on the Internet. To my knowledge none of the people that run our client are actually pulling data from our system, so I'm not so sure how our agent could be at fault for someone else stealing his intellectual property and putting it on their website. Keep in mind that each client does not have easy local access to the data that is pulled back by the software and that in order to access the data, one would have to have access to our database.
Pendanticist uses the terms "downloader, scraper, ripper, and grub" to describe hostile bots that roam the Internet presumably trolling for information that they can exploit. Nobody likes these vermin for obvious reasons, and I can understand the need to block these bots on an agent by agent basis. My only regret here is that our name has somehow been incorporated into describing those other not-so-nice agents.
Grub-client and the Grub Project do not utilize the data that is gathered by its system for malicious or dubious purposes. On occasion we have other search engine projects pull data from us, but we always check out what they are doing with the data before we authorize them to use it. Grub will NOT sell data to people who utilize it for IP theft, spamming purposes, or any other activity that would adversely affect those that run the sites.
Crawling Behavior
-----------------
Our software was designed to be run by website maintainers - people like you - to enable you to crawl your own content, with your own internal bandwidth. The very reason that I started this project was because I was sick of the multitude of crawlers that plagued my sites and burned my bandwidth.
That said, Grub has on occasion been known to do some things that are less than friendly to unsuspecting websites.
Among the infractions we are guilty of are mangling the user-agent names (corrupted stack bug), requesting more than one URL at a time (sequential database entries), and blindly ignoring valid robots.txt files (our stupid, crash prone robot code). Over the past two years the three of us have strived to correct these behaviors, but the fact is that we still need *some* work on certain aspects of the code to have a polished product.
Fineware brings up the recent issue where he had requested to have his URLs removed, and then found us once again crawling him, which was due to us having to rollback the database prior to when I removed his URLs from the database. I have again removed his URLs from the system, and have put plans in place to ensure we keep track of requests to remove URLs from the system. This seems only fair as we already track requests to be *put* into the system! ;)
Robots.txt Respect
-------------------
As for the robots.txt tracker, it works for the most part, but it is very load intensive on the system and causes the occasional crash which prevents it from operating on a continuous basis, or at 100% accuracy.
We have made this a top priority to fix, and we expect to see a performance and stability boost out of the grubbot process in the next few weeks. We realize that, for now, performance of our current robot crawler is less than stellar, so we would be more than happy to manually deal with any issues a site may have concerning the crawlers activities.
Reverse Name Lookup Confusion
-------------------------------
Addressing the issues raised concerning Looksmart's client, people need to be aware that just because a site, or network is running our client, does not indicate that they are actually pulling or utilizing data from the system, or the sites they crawl.
The person at Looksmart is simply running the client, helping out like everyone else involved with the project, to crawl the URLs that are in our database. Furthermore, Looksmart's search engine, Wisenut, does not currently utilize the data that is returned by our distributed network, nor do we supply them with any other type of meta data concerning the URLs that we track. In other words, blocking grub-client doesn't equate to blocking Looksmart.
I hope this long winded explanation sheds some light on what we are trying to do. Please feel free to contact me if you have any other questions!
Kord Campbell
kord@grub.org
President
Grub, Inc.
Clearly, what you've posted was not taken from here, but rather from a stickmail communication we had.
If that isn't a blatant infraction of the TOS in WebmasterWorld...it sure as hell outta be!
However, just to keep the record straight, I've decided to post my entire message to you as verification.
Greetings Kord Campbell,I saw your post and was wondering if you could provide me some data that illustrates some of this "bad" behavior that you mention concerning the Grub client, and/or other bots that might be posing as us.I'm not sure how someone could "piggy-back" on our system, but people do strange things all the time. There may, or may not be something that we can do to prevent this.
We would like to take whatever actions necessary with the system to avoid being labled as a "bad bot" out in the real world. We want webmasters utilizing and running the bot, not being afraid of it.
[webmasterworld.com...]
This thread has 199 messages and spans 14 pages, so it make take a bit of reading.
As for me (a relative newbie to the forums since Apr 27, 2002), I started banning Grubs sometime in the middle of that thread and did so based on all the pages I kept seeing being crawled by grubs.
My domain is an academically focused .com that I've worked very hard since '97 to get where it is today. Additionally, I've traced back many of those who've run 'grubs' and found their new additions to have come from my work.
I don't trust any grub client and I ban any downloader, scraper, ripper, grub I ever find. As it goes, there have been times I missed banning a grub only to find the next day that the grub scarfed every file I had and in very short order.
While I respect your right to develope applications, in turn, you should respect my right to deny that applications use on my domain, whatever it's original intentions were.
I sure don't know what else to say Kord, other than urge you to post to the forums (I see you just joined today) and see what other, more knowledgable webmasters have to say.
Pendanticist.
I stand by everything I've said and this Sir, is the last I intend to say on the matter...to you, or anyone else.
It's my domain and I'll do with it as I see fit!
Pendanticist.
[webmasterworld.com...]
As I've been led to believe, Grubs are inherently bad bots that download/scrape or otherwise 'take' your material.Do a site search for that nearly perfect .htaccess ban list and you'll find Grub(s) right in there with the rest of them.
Naturally, I was surprised to see LS using such a commonly known-to-be bad bot.
Pendanticist.
1. Looksmart isn't currently "using" our crawling system - they are simply contributing to the project by having one of their workstations run as a crawling node. They don't the data we collect either, so you could block grub-client without blocking Looksmart's crawlers.
2. grub-client!= grubs. I think that using the terms "scrape" and "take" in this context is implying that Grub has less than honorable intentions with the data. The fact is that we don't use the data any differently than any other normal bot, and just because we are in someone's recommended .htaccess file doesn't mean that blocking us is still justified, or that our system is malicious in nature.
Granted our robots.txt functions may be a bit flaky, but we are trying fix that - and address it here. Just for the record, we have ~35M URLs that we track, ~2.1M hosts that we know about, and ~21,000 lines of robots text from those hosts that have a robot file and that are in our database.
Pendanticist,
We have never indicated to you, or others that we have the right to crawl your site and you are absolutely correct when you say it is your site, and your system. In short, it is a privilege for us to crawl other's sites - not a right.
However, I don't think that blaming us for things we aren't doing is your right as well. If we are doing something wrong, then we'll own it and fix it - if we aren't, then you can expect us to continue to talk about it here.
Kord
[grub.org...]
We've tested the script today, and it seems to be working ok. If anyone has any problems with it, please let us know by emailing support@grub.org.
We should have another page up in a few days that will allow you to request complete removal of your URLs from our system. This form will keep sites that don't want to be added to the system in a separate table than the URLs that we track, and should the database ever be rebuilt, or damaged, we can rebuild it with your pages excluded.
Kord Campbell