Forum Moderators: not2easy

Message Too Old, No Replies

Copyright Infringement

It is not just small scale and it certainly isn't new

         

iamlost

10:47 pm on Apr 8, 2005 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



It is always interesting to watch the masses wake up to some "new" problem. The current webmaster bug-a-boo is copyright infringement.

Read the thread titles in this forum over the last several months and see the same topics addressed and re-addressed ad nauseum.

Imagine being in the content generation/copywriting business and not understanding copyright. How the heck can you play the game if you haven't the foggiest what the rules are? If it weren't so sad I'd ROFL.

Site scraping is not new! I saw my first scrape-script in 1996. I know a couple of very prosperous types who saw the "future" of Adsence and similar programs and went content stealing in a big way beginning over two years ago. And have never been caught and likely never will be:

All domains owned through shell companies registered in tax havens with contact person being each firms lawyer. Each site targets a very particular niche. They continually run SE queries for each niche/keyword/keyphrase. Their own bots happily go out and act just like any any other search bot returning pages of other peoples sweat and content, parse the html, index the results and add to the database. New index words are "logged" to be used to generate DB queries. The queries generate all the sentences from all the scraped sites with that particular index word. All automated.

Initially, at this point, they had to write their own "original" content and dump it into a page template. Now they use a program that aggregates the queries into auto-generated sentences and they just edit and paste into the template.

From scratch, one to two pages of brand new content uploaded per person per day is doing well. These folk manage five to ten pages per hour each. And it is "their" words not "copied" words so how will you ever know? Many of their sites are considered authoritative in that niche. And they laugh all the way to the bank.

I expect many others are doing something similar.

And now some SEs are returning "answers" to "queries" rather than just links to sites (New Google Answering Facts [webmasterworld.com]). Remember that you gave (at least implicit and perhaps explicit) permission for SEs to scrape your pages and list them. You might want to be real clear on what copyright you are granting the SEs along with the scraping and cacheing and listing.

And if a couple of ordinary folk can morph others content into their own just think what the SEs could do. I have several clients whose Home/Major Category index pages are the only pages shown to bots. They take their content copyright very very seriously. When a SE does list an interior page an immediate removal request is issued by their lawyers.

Web search is mutating and you better know the rules and what some are doing to bend/break them or the future will be an unpleasant surprise. But not new.

BigDave

11:19 pm on Apr 8, 2005 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member



Wow! Where to start.

Imagine being in the content generation/copywriting business and not understanding copyright.

I totally agree.

The only problem is that what you describe is in no way copyright infringement unless you REALLY stretch the definition of derivative work even farther than Darl McBride.

All domains owned through shell companies registered in tax havens with contact person being each firms lawyer.

Tax haven doesn't matter. I suspect that they are countries with high privacy standards that are not members of the Berne Convention.

If they are not part of the Berne Convention, then your work gets no protection under that national laws of that country. You have no copyright in that country, so any copying done in that country would not be copyright infringement.

And it is "their" words not "copied" words so how will you ever know?

So if it is not "copied", and it is just some sort of automated gibberish, then how is it copyright infringement?

And now some SEs are returning "answers" to "queries" rather than just links to sites (New Google Answering Facts)

You cannot copyright facts.

Remember that you gave (at least implicit and perhaps explicit) permission for SEs to scrape your pages and list them.

Actually, much of their permission comes from what copyright does not grant to the rights holder.

Reading a published webpage is not copying.

The results that they serve up are probably covered under fair use.

The cached copy is possibly covered under the provisions of the DMCA.

They mostly follow your restrictions on crawling because they are nice ;-)

You might want to be real clear on what copyright you are granting the SEs along with the scraping and cacheing and listing.

Are you suggesting trying to put restriction on rights that were not granted you by copyright law?

Basically the scraping that you are talking about (other than the SE part) is the generation of crap pages. And if they make it to the point that actual people consider their site authoritative, without committing copyright infringement, then I would chalk up their actions to research rather than call it scraping.

iamlost

3:57 am on Apr 9, 2005 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



The only problem is that what you describe is in no way copyright infringement unless you REALLY stretch the definition of derivative work even farther than Darl McBride.

I should have used a different thread title. Perhaps "Automated Content Morphing". The title does skew how you read the message. But it did catch your attention :-)

Countries with strict privacy laws tend to also be tax havens and privacy and tax shelters are likely equally why these countries were/are selected and certainly why I used the term. Perhaps I should have used Bruce Sterling's term "Data Haven" but while more accurate I thought it likely too esoteric without an accompanying explanation. Perhaps I was wrong.

If they are not part of the Berne Convention, then your work gets no protection under that national laws of that country.

Where the server is is more to the point. Server charges in many/most/all Data Havens are rather steep. Just as most of the world's spam originates in the USA so there also reside most of the scrape generated sites I have tracked.

The results that they serve up are probably covered under fair use.

Cached pages aside, as results are currently presented I agree.

The cached copy is possibly covered under the provisions of the DMCA.

The DMCA is a national law. Would depend on where a case was brought as to its relevance.
Presenting cached complete copies in place of links coupled with a short description certainly exceeds fair use.

They mostly follow your restrictions on crawling because they are nice

They are not "nice" and haven't been for years. SE bots often ignore robots.txt. Cloaking is alive and well :-)

Are you suggesting trying to put restriction on rights that were not granted you by copyright law?

No. I am suggesting that one should be clear what ones rights are and, equally important, what they are not. My detailed example was to show how people already are bypassing copyright and that there is not much anyone can do about it.
Page scraping is unlikely to be eliminated but industrial strength automated scraping can be mitigated if at the potential displeasure of the SEs.

Basically the scraping that you are talking about (other than the SE part) is the generation of crap pages. And if they make it to the point that actual people consider their site authoritative, without committing copyright infringement, then I would chalk up their actions to research rather than call it scraping.

The pages I mention are not crap, unfortunately.
I agree that it is actually automated research but it is accomplished by an initial scrape. That it is massaged/morphed prior to presentation simply illustrates the difference between a script-kiddy and a professional.

I am not being a Luddite nor crying wolf, just pointing out a present reality and a potential future.