Forum Moderators: open

Message Too Old, No Replies

auto categorizer

categorize any website on the internet

         

monsterisp

5:46 am on Feb 18, 2004 (gmt 0)

10+ Year Member



We built a script that will categorize any website (by spidering and analyzing the page) into a dmoz like directory with hundreds of categories and sub categories. There is ALOT of SMART AI technology built into how it determines where to place the content. Do you guys see any applicable use for this technology? Has someone else developed it already? It could take the entire internet and categorize every website into 1 large directory. The code is already in place but we are in a beta phase right now...

IITian

7:35 pm on Mar 2, 2004 (gmt 0)

10+ Year Member



Results look promising. I found that when I was check out small websites of a few pages, it returned the results almost instantly. However, when I gave it websites containing a few hundred pages, it got stuck. Maybe, it has something to do with the server.

I wondered, if it has the technology to sample just a few high-level pages (or just one page) to deternime a close enough category and stop.

giga

4:27 am on Mar 4, 2004 (gmt 0)

10+ Year Member



I have some code that I may be able to hack up to do this
any luck?

econman

5:26 pm on Mar 5, 2004 (gmt 0)

10+ Year Member



As to your orginal question, yes, I think this could be a valuable technology that might have multiple uses, including ones that don't require you to immediately reach the goal of categorizing milllions of websites.

I've sent a sticky note indicating an interest in using/licensing the technology as part of a project we are thinking about pursuing, but haven't received a reply. Let me know if you need me to send another sticky note.

giga

9:13 pm on Mar 5, 2004 (gmt 0)

10+ Year Member



Hi econman, we can utilize this formula to categorize any amount of data, in any format (even non html files). It was origionally created to categorize millions of text strings into 1 massive organized directory.
Please sticky me with your intentions, and perhaps we can create a demo for you, tailored for your specific application (to prove its potential).

Thank you,
Giga

[edited by: skibum at 7:52 pm (utc) on Mar. 13, 2004]

MarshallClark

11:05 am on Mar 6, 2004 (gmt 0)



Hi Giga,

I'm talking to a potential client that may be able to benefit from your technology. They have a very large database of articles that I'd like to see converted into a directory. Could your software be populated with data other than a pre-existing directory? Could I instead provide my own hierarchial category list and have it work from that? If a category list was used (instead of a directory of sites) would I need to associate keywords with each category first? Thanks for any info.

giga

11:38 am on Mar 6, 2004 (gmt 0)

10+ Year Member



Could your software be populated with data other than a pre-existing directory?

The purpose of our categorize is to populate any directory structure with any data in any format. The mathmatical logic formula is universal in nature, and thus we have a script that will categorize any amount of data into structured categories.

Could I instead provide my own hierarchial category list and have it work from that?

Definatly, wit ill work on any hierarchial directory structure provided by you, we have also developed a version that will automatically create its own sub categories.

If a category list was used (instead of a directory of sites) would I need to associate keywords with each category first?

We should talk offline regarding this step.

Thanks
Giga

econman

2:07 pm on Mar 6, 2004 (gmt 0)

10+ Year Member



I tried a URL for a relatively small site with a well-defined scope, and within less than a minute it gave me back a suggested directory location.

The suggested directory structure and location seemed reasonable and appropriate. However, it also offered numerous other suggestions, most of which were not as appropriate (because they were too specialized, and/or only reflected a small part of the overall content on the test URL I was trying).

I'm a bit unsure about what is going on, and thus how to interpret the results. Were all of the suggested locations taken from DMOZ or some other directory structure that existed prior to my submitting my sample request? Or were the suggestions invented on the fly, in response to my request?

Stated another way, are the suggestions provided to the user choices for sublocations within a pre-existing directory structure that has been predefined (whether by you or by DMOZ)?

If I were judging how well the script placed the URL within the offered set of options, the top choice was certainly the best.

If I am judging how many of the suggested options were appropriate for the URL, I would say relatively few.

If I am judging whether the suggested directory structure was appropriate/good (e.g. this is a computer generated structure and the question is whether it does as good a job as the humans at Yahoo or DMOZ) I would need to test it some more to pass judgment.

2. Can this technology be applied to geogrpahy rather than, or in addition to, subject matters? For instance, could it be adapted to the problem of mass-evaluating whether URLs contain content that is primarily related to a specific geographic area (e.g. finding Dallas Texas portals, travel guides devoted to Texas) or whether it is operated by an entity that is located in or focusing on a specific geographic area (e.g. plumbers and chess clubs that are located in Dallas)?

3. Is this software capable of drawing a distinction between a website and a web page?

As an aside, this is one of the biggest weaknesses with Google and other search engines. They don't draw any such distinction. Most of the results shown to users are based upon the overall site content, but other results are for individual pages. There are perhaps 6 million active domains. When Google talks about indexing 4 billion webpages, they are referring to the pages within those 6 million sites, but their results don't adequately distinguish between the two concepts.

Sometimes I want to find an article or other information about a specific topic (I'm looking for the best webPAGE about discussing widgets) and other times I'm looking for the best webSITES for doing research, or making a purchase, related to a topic (I'm looking for the best place to learn about or purchase widgets.)

There is a big difference between trying to identify the best page within a book, or the best book. Google, et al, fail to recognize this fundamental distinction. If your tool could draw such a distinction efficiently, it might be helpful to someone trying to build a specialized "niche" search engine.

4. I didn't see, or at least didn't recognize, the keyword suggestion tool. Can you clarify where this is, and how it is supposed to work?

giga

3:18 am on Mar 7, 2004 (gmt 0)

10+ Year Member



ANSWERS TO YOUR QUESTIONS

I tried a URL for a relatively small site with a well-defined scope, and within less than a minute it gave me back a suggested directory location.
The suggested directory structure and location seemed reasonable and appropriate. However, it also offered numerous other suggestions, most of which were not as appropriate (because they were too specialized, and/or only reflected a small part of the overall content on the test URL I was trying).

**The suggested matches are SIGNIFICALLY weaker potential matches than the primary match. Perhaps we should remove the suggestions (as they were more of a diagnostic tool for us to gauage relevancy and to tweak the algo)

I'm a bit unsure about what is going on, and thus how to interpret the results. Were all of the suggested locations taken from DMOZ or some other directory structure that existed prior to my submitting my sample request?

**Yes from DMOZ in this particular demo.

Or were the suggestions invented on the fly, in response to my request?

**No, but an interesting concept :)

Stated another way, are the suggestions provided to the user choices for sublocations within a pre-existing directory structure that has been predefined (whether by you or by DMOZ)?

**Yes all 450,000+ categories have been pre-defined by DMOZ, however we are developing a version that will create its own sub categories automatically based on the end users suggested top level topics.

If I were judging how well the script placed the URL within the offered set of options, the top choice was certainly the best.

**As stated above it should have a significally higher score. The others are just placed in, to avoid a complete failure to categorize.

If I am judging how many of the suggested options were appropriate for the URL, I would say relatively few.

**We plan to slim down if not remove the suggetions all together.

If I am judging whether the suggested directory structure was appropriate/good (e.g. this is a computer generated structure and the question is whether it does as good a job as the humans at Yahoo or DMOZ) I would need to test it some more to pass judgment.

**In some ways it does make a better match for some sites than a dmoz editor might suggest, in other instances it completely misses the mark, however what makes this unique is that with a high degree of accuracy it can categorize millions of pages in a timely fashion without human intervention.

2. Can this technology be applied to geogrpahy rather than, or in addition to, subject matters? For instance, could it be adapted to the problem of mass-evaluating whether URLs contain content that is primarily related to a specific geographic area (e.g. finding Dallas Texas portals, travel guides devoted to Texas) or whether it is operated by an entity that is located in or focusing on a specific geographic area (e.g. plumbers and chess clubs that are located in Dallas)?

**It already does. Try putting in a popular local college. The results are very much determined by the sturcutre of the dmoz subcategories. In other words if there is a website in Dallas TX and no subcategory provided by DMOZ for a Dallas TX listing to match against in its particular field, well.. the categoriezer cannot find an appropriate match. In some ways it is only as good as the directory it is provided with. In this unique case it is DMOZ

3. Is this software capable of drawing a distinction between a website and a web page?

**That would be a very easy addition to the code.

As an aside, this is one of the biggest weaknesses with Google and other search engines. They don't draw any such distinction. Most of the results shown to users are based upon the overall site content, but other results are for individual pages. There are perhaps 6 million active domains. When Google talks about indexing 4 billion webpages, they are referring to the pages within those 6 million sites, but their results don't adequately distinguish between the two concepts.

**Traditionally directories typically only index the homepage url, and not individual pages.

Sometimes I want to find an article or other information about a specific topic (I'm looking for the best webPAGE about discussing widgets) and other times I'm looking for the best webSITES for doing research, or making a purchase, related to a topic (I'm looking for the best place to learn about or purchase widgets.)

There is a big difference between trying to identify the best page within a book, or the best book. Google, et al, fail to recognize this fundamental distinction. If your tool could draw such a distinction efficiently, it might be helpful to someone trying to build a specialized "niche" search engine.

*That is one of our goals.

4. I didn't see, or at least didn't recognize, the keyword suggestion tool. Can you clarify where this is, and how it is supposed to work?

*We will put this feature back online momentarily.

Thank you for all your interest!

giga

8:28 pm on Mar 10, 2004 (gmt 0)

10+ Year Member



Any other questions we can answer?

boselecta

8:42 pm on Mar 10, 2004 (gmt 0)

10+ Year Member



I recommend you take a look into the TREC reference collections.

They are used as standard benchmarks for the effectiveness of auto-categorisation algorithms and consist of a training set of pre-classified docs and a test set of docs that your software should then attempt to classify correctly.
The Reuters 21578 collection is one common example.

Using these sets you can begin to quantify just how good your software is (or not) and compare it with the performance of others.

giga

4:40 am on Mar 11, 2004 (gmt 0)

10+ Year Member



That is exactly the direction we needed to hear. Thank you very much for your post. Since we are unfamiliar with anyone doing anything similar it was hard for us to compare to similar sets, however this will make a great test.

Thank you,
Giga

ps. sticky us if you would like to chat!

giga

12:33 am on Mar 12, 2004 (gmt 0)

10+ Year Member



We have removed the adult category from the visible listings on the demo as per WebmasterWorlds TOS.

a1call

8:37 pm on Mar 12, 2004 (gmt 0)

10+ Year Member



Hi,
>Any other questions we can answer?

I have a web directory. Many links would be suitable for more than one category. Can your software Check all links and put them in all other categories that they relate to or only one category will be suggested.
How can I give it a try?
Thanks

Marcia

9:50 am on Mar 13, 2004 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member



I'd love to go beyond the mystery here and see what this is all about. Can you please stickymail me the URL, so I can have a look also? I kind of feel seriously left out on this deal. ;)

Thanks, much appreciated!

giga

10:31 am on Mar 13, 2004 (gmt 0)

10+ Year Member



[65.87.196.76...]

provide any feedback! btw this works on any set of data, not just websites. We are applying it to a seperate project of ours and its doing incredible results. Thank you for your interest everyone. (helps motivate us) :)

giga

10:33 am on Mar 13, 2004 (gmt 0)

10+ Year Member



I have a web directory. Many links would be suitable for more than one category. Can your software Check all links and put them in all other categories that they relate to or only one category will be suggested.

Well it works as if a human reviewer were attempting to place each link (semi-artifical logic at play). If a link had a apparent and common sense placement into one of your categories or sub-categories it would very likely be placed. However if it had no relation to any of your topics it would fail to place (as there would be no relevant category for it).

econman

11:26 am on Mar 13, 2004 (gmt 0)

10+ Year Member



Please send me a note via sticky mail with an email address or other way I can reach you. I tried sending you a note via sticky mail a few days ago, but didn't receive a reply.

giga

12:46 pm on Mar 13, 2004 (gmt 0)

10+ Year Member



Sent the sticky. :)

giga

6:28 am on Mar 24, 2004 (gmt 0)

10+ Year Member



Anyone have any suggestions or possible uses for this technology? It is flexable beyond dmoz's structure, and works on any data (not just websites).

giga

5:25 pm on Mar 24, 2004 (gmt 0)

10+ Year Member



I am recieving sticky requests everyday for the url,

here is it

**** [65.87.196.76...] *****

giga

3:19 am on Apr 2, 2004 (gmt 0)

10+ Year Member



Anyone think perhaps this technology could help with google's adsense relevancy, or similar relevancy determining pay per click systems? In theory the script could spider a page, learn what its about, match that against a database of advertisers, and display only relevant ads based on the "theme" of the website.

Giga

econman

10:51 am on Apr 9, 2004 (gmt 0)

10+ Year Member



I've been mulling over your concept, and it's potential usefulness for a commercial application. From this perspective, the key questions are:

1. Have you done any analysis to estimate the accuracy of your automated sorting process? It doesn't need to be 99% accurate, but the more accurate the better. Have you done any methodical testing to compare how its' sorting compares to what humans would do? (note: you could set up a simple test by asking your software to sort 30 web pages or websites into a classification schema; separately, ask at least 3 different humans to sort the same 30 web pages or websites into the same classification schema. Then compare results.

2. How much computing power does it take to accomplish the sorting? Stated differently, if you devoted a decent quality PC with an adequate amount of RAM to the task of running your software full blast around the clock, how many individual pages within a website could it classify in a 24 hour period, or per hour? How many entire websites (of average size and complexity) could it classify per 24 hours or per hour?

Depending on the answers to these questions, it might or might not be useful (both for my project, and for other applications).

giga

6:04 pm on Apr 10, 2004 (gmt 0)

10+ Year Member




1. Have you done any analysis to estimate the accuracy of your automated sorting process? It doesn't need to be 99% accurate, but the more accurate the better. Have you done any methodical testing to compare how its' sorting compares to what humans would do? (note: you could set up a simple test by asking your software to sort 30 web pages or websites into a classification schema; separately, ask at least 3 different humans to sort the same 30 web pages or websites into the same classification schema. Then compare results.

2. How much computing power does it take to accomplish the sorting? Stated differently, if you devoted a decent quality PC with an adequate amount of RAM to the task of running your software full blast around the clock, how many individual pages within a website could it classify in a 24 hour period, or per hour? How many entire websites (of average size and complexity) could it classify per 24 hours or per hour?

1. We have never attempted to go head on head with a human as to accuracy. This would be an EXCELLENT methodology from which to compare the two. In regards to accuracy, we believe we are generally satisfied with the results it is producing. I think the next step may be to follow your advice and attempt a test. I would suggest someone, anyone either sticky me (or post, not sure about the tos though..) 30 random websites. We then take an idependant possibly even DMOZ editor and he attempts to "quickly" categorize all 30 sites. We record the time it takes, and the final results of his attempts for each site. Then run our categorizer on the same pool of data. IF you would like to assist us (or anyone) in this academic endeavor to ensure it is 100% unbiased and objective in nature, we gladly welcome the challenge (i'm curious how it compares as well!).

2. The computing power is an issue. We have found a way to build a clustering of computers to help assist in the task of categorizing websites. IN other words several servers on our network can each work on multiple tasks at 1 time, and all contribute their final results to the total tally. As we have stated from the beginning, many many many probability scenarios, and calculations are being made for each and every site that we attempt to categorize as the computer tries not only to categorize the website, but also to understand the general theme of what the site is about. I am not the programmer, so I do not know specific details regarding computing power, however i do know that there is alot of applications going on in the background as it attempts to find relevancy, also a clustering of servers greatly increases computation time.

-Giga

giga

7:22 am on Apr 11, 2004 (gmt 0)

10+ Year Member



If anyone has any suggestions or joint partnership ideas on how to either A. Capitolize on this endeavor, or B. Simply get it out to the public as a valuable resource please post. We have the server space, and a piece of technology that can do what a dmoz editor can do but faster, and "possible" more efficiently... not only that but this technology can also mimic adsense's own relevancy ad server. Finally we were thinking about making it to auto develop directories for end users (for free) so that it could be placed on their niche websites and would only spider and place websites related to whatever directory structure the end users decide. However as of late, interest has wanned from this project :(

tombola

8:47 am on Apr 11, 2004 (gmt 0)

10+ Year Member



giga, though you're bumping up this thread regularly to get fresh attention for your auto categorizer, you say that "interest has wanned from this project".

If you offer a good product/service, people here will show their interest, but it seems they're not interested in pure hype.

giga

10:53 pm on Apr 11, 2004 (gmt 0)

10+ Year Member



LOL, "pure hype"? Please explain, from where I sit I'm looking at a functioning prototype. I bump the tread because the categorizer is doing somthing alot more powerful (what we believe) than even google or any of these other engines are doing right now, i see no one offering such relevancy. Screw it if you guys want to see a simple search engine, thats easy at this point. I just dont care to spend years growing old building somthing everyone else already has. I thought this tech could be used by someone on top to grow even taller.

Honestly your right, i'm talking to myself and hear only my own echo reply. We'll move on from this forum since it feels like no one here really "gets it".

[65.87.196.76...]

This 86 message thread spans 3 pages: 86