Forum Moderators: Robert Charlton & goodroi
[news.com.com...]
A Nevada federal court has ruled that the cached versions of Web pages that Google stores and offers as a part of many search results are not copyright infringement.
Clearly, the court did not understand what real caching is and what Google calls caching. I do not thing Googles meets the crieteria for caching:
The material described in paragraph (1) is transmitted to the subsequent users described in paragraph (1)(C) without modification to its content from the manner in which the material was transmitted from the person described in paragraph (1)(A) {FN104: 17 U.S.C. §512(b)(2)(A)}
What this does I think, is effectively neuters all copyright laws on the internet today. It is the wild-wild west again.
With all that Google has done that is good - I don't know how we could be so far apart on this one issue.
Blake Field: (who brought the suit):
[blakeswritings.com...]
Unlike Google's caching, which may benefit Google indirectly, framing of third-party pages with ads offers direct revenues to the framer.
One could argue that framebreaking code is a quick-and-easy defense, but it isn't that simple, because framebreaking code can cause other problems and needs to be implemented carefully (unlike Google's benign "nocache" tag).
As far as I know, the only major case involving framing was the TotalNews lawsuit of 1997, which was brought by the NY Times, Washington Post, CNN, Reuters, and a few other companies after TotalNews.com framed their stories with its own ads and navigation scheme. Unfortunately, that case was settled out of court, so no precedent was set--and today, at least one major Web property--About.com, which ironically is now owned by the New York Times Company--is profiting from other the content of other Web sites by running ads above their content.
He didn't argue it at all because the judgement was not based on that point.
> could have used no cache.
As a lawyer, he knew he could not put a proprietary tag on his site.
<meta name="googlebot" content="noarchive">
is implicity copyright Google and may be covered by G's patents.
From the legal precedings:
With this knowledge, Field set out to get his copyrighted works included in Google’s index, and to have Google provide “Cached” links to Web pages containing those works.
Field created a robots.txt file for his site and set the permissions within this file to allow all robots to visit and index all of the pages on the site.
When Google learned that Field had filed (but not served) his complaint, Google promptly removed the “Cached” links to all of the pages of his site. See MacGillivray Decl. ¶2; see also Countercls. ¶22; Ans. to Countercls. ¶22. Google also wrote to Field explaining that Google had no desire to provide “Cached” links to Field’s pages if Field did not want them to appear.
It is now clear. If you do not know of a standard, such as "NOFOLLOW, NOINDEX", that is too bad. The Web has just moved from teenager to adult. No longer the uninitiated may just throw up some material and presume everything is fine. Look for the next de facto standard to protect/erode your business.
If you're really, really worried about this, why not simply put a "nocache" tag on your pages?
... because that is simply saying that what they are doing is okay, and that I will comply with "The Law of Google".
Google caches everything it comes in contact with without regards to the knowledge level of the webmaster (and it is not their requirement to know that Google will copy their work if they don't "opt out")
This includes artistic works, web based games, text based content... everything. I am not comfortable with this, never have been, and never will be. - and I shouldn't have to risk economical internet suicide to prevent it.
It is now clear. If you do not know of a standard, such as "NOFOLLOW, NOINDEX", that is too bad.
Actually, that is the one thing that I think is very unclear. Field knew of the meta tags and robots.txt. If someone that does not have a knowledge of such things were to file the lawsuit I would be willing to bet that it would not have been decided at the summary judgement stage.
I really think that it was the estoppel and the implied license from the robots.txt that killed the case and will keep it dead.
If the ruling stands it will almost certainly hurt future cases by those that do not know about robots.txt and the meta tags, but with good lawyers they should at least be able to argue about the differences between this case and their own.
Like I said, this is Google's dream case on this sort of issue. The plantiff behaves like the bad guy, so google becomes the good guy by default.
It is sad, that this has happened this way. There is something to be said about Google profiting on non-tagged material.
I would have preferred a grandma with no technical knowledge discovering this on her poems or love letters. Then it might have ended differently. I believe Mr. Field, by admitting that he was fully aware of the tags, and his steps of filling copyrights, revealed his intentions.
I stand corrected BigDave. You are right. That is unclear.
I'd be more comfortable if users could only access the cached version if the original site is down. That would be a service to both users and webmasters.
4
Then it might have ended differently. I believe Mr. Field, by admitting that he was fully aware of the tags, and his steps of filling copyrights, revealed his intentions.
It does not seem to matter much what his 'intentions' were. In my mind, this is very cut and dried:
Google has a copy of someone else's work without their expressed permission.
That is copyright violation. Period.
Opt out is not expressed permission. Opt-in, however, is, and I believe Google should make inclusion 'opt-in'.
I'd be more comfortable if users could only access the cached version if the original site is down. That would be a service to both users and webmasters.
That's not a bad idea. One difficulty would be determining whether the original site is down. (Sometimes a site is invisible to some users but not to others.)
According to the undisputed testimony of Google’s Internet expert, Dr. John Levine, Web site publishers typically communicate their permissions to Internet search engines (such as Google) using "meta-tags." A Web site publisher can instruct a search engine not to cache the publisher’s Web site by using a "no-archive" meta-tag. According to Dr. Levine, the "noarchive" meta-tag is a highly publicized and well-known industry standard.
Since when did "noarchive" become an industry standard? Sounds to me like Google's expert misrepresented the facts.
Field even uses one of those defacto standards (robots.txt) to specifically allow google to crawl his site with the knowledge that it will cache those pages.
All the legitimate search engines follow those "standards" and millions of sites use them. They qualify as a standard.
<added>The decision mentions "undisputed". That does not mean that no one can dispute it, it means that Field did not dispute it in his filings or at the hearing. Again, that is where you want to make sure you have a good lawyer and your own exper witnesses.
This legal ruling is absolutely correct. Remember, when you make your content public on the Web, it is called "to publish your site"! When you make your content publicly available and you don't define any restrictions on your robots.txt file, that means your content can be cached.
If he wants Google not to cache his pages, there's a far simpler remedy than launching a federal lawsuit. Thus the focus of the suit is less on copyright infringement and more on the plaintiff's feeling that he shouldn't have to take any action to prevent cacheing.
Since when did "noarchive" become an industry standard? Sounds to me like Google's expert misrepresented the facts.
Google caches everything it comes in contact with without regards to the knowledge level of the webmaster (and it is not their requirement to know that Google will copy their work if they don't "opt out")
[edited by: mcavic at 9:50 pm (utc) on Jan. 26, 2006]
it is incumbent upon webmaster's to understand and know them all.
That is not what the ruling said. The ruling was largely based on the fact that Field KNEW how to keep it from being archived, yet did not do it.
By his actions and admissions he wanted google to cache it so he could sue them.
If he does a robots noarchive expecting it to work on MSN and it does not, even if it works on all others, then he very well might have a case against MSN using the argument of Google's expert in this case.
The camp which holds the belief G's cache violates copyright, period, over and out, done may desire to evaluate anew this belief rather than dig in heels and stomp and snort.
I think G gets a well played and played well nod here.
I also think the disposition of an appeal, assuming an appeal is taken by the non-prevailing party, will prove an interesting read.
Re: volatilegx's post containing precedents:
Agreeing with the analysis in Netcom, we hold that the automatic copying, storage, and transmission of copyrighted materials, when instigated by others, does not render an ISP strictly liable for copyright infringement
As each reference noted. This group of precedents refers to whether the ISP can be held liable when one of their customers commits copyright violations.
Re: Brett_Tabke's comment:
Search Engines are not caching your page - they are republishing it with their banner advertisement on the top
Well ... they are, in fact, caching it. They are also, in fact, adding an additional brand (theirs) to the version of the page when it is viewed from within their cache. They didn't change the page, and hence the framing arguments. Fortunately, all of the promise held by the original page is still intact, including links to the current version and any other links to your other 'real' (not the cached) pages, and your 'buy stuff from us now' messages. If you're a plumber, the visitor viewing the cached page still has your phone number ... right?
Rather than Google selling/re-publishing/profiting from your complete previously published work (which is the whole site, not just one page, right?), they are basically sticking a copy of one of your pages (at a time) up on the front of their newsstand. You may choose to shop from their offerings while looking at the poster, but if the poster has any interest to you, you would be inclined to follow its promise independent of where you first found it. Google doesn't sell plumbing services, yet. Are they in competition with you? I don't see any ads except for those already embedded in the pages, so in reality they are only benefiting from their branding, not from anything that would cost you a sale.
Re: HeatherR's comment:
Google has a copy of someone else's work without their expressed permission. That is copyright violation. Period.
Nope. There's wiggle room, as many have stated.
This is a little 'cache 22', if I may. We like the search engines, they bring us customers, we just don't want to have to do anything to directly address the technology we all know they use ... and goodness knows we don't want them to make any kind of money!
Seriously, folks, the big picture is that this is a very narrow ruling that does not address scrapers or router caching or even browser caching. This only addresses the situation brought up by Mr. Blake with regard to his relationship with Google. The question is nowhere near settled, but for Mr. Blake, it is. (Are we certain he's not a Google contractor, getting some jurisprudence onto the books for their benefit?)
I can't just go around the internet copying images and expect a link to the original authors works to suffice(the bold is mine)
What concern's me is that this seems to have become accepted. Look at all the scraper sites. They put in a link to me when they copy and publish a snippet from my site. Google links to me when they publish the cached version of my page. Neither one resembles what I thought 'fair use' meant.
I don't have a problem with my site being cached but publishing it must be a copyright violation. I could live with search engines doing it but it seems to be opening it up to everyone else.
First search engines used a snippet to describe my site. Now thousands of scrapers do as well. They will just claim they are providing search information. So now hundreds if not thousands of pages on the internet are using info from my sites.
So maybe the whole thing needs to be revisited. I don't have any answers but I'd sure like to see a solution.
The question is nowhere near settled, but for Mr. Blake, it is.
Exactly why the claim of "the most important decision" yada yada yada is overstating it a bit.
open season
(Are we certain he's not a Google contractor, getting some jurisprudence onto the books for their benefit?)
Maybe google will show their appreciation by bestowing higher PR on Blake's site.