Forum Moderators: open
I wondered, if it has the technology to sample just a few high-level pages (or just one page) to deternime a close enough category and stop.
I've sent a sticky note indicating an interest in using/licensing the technology as part of a project we are thinking about pursuing, but haven't received a reply. Let me know if you need me to send another sticky note.
Thank you,
Giga
[edited by: skibum at 7:52 pm (utc) on Mar. 13, 2004]
I'm talking to a potential client that may be able to benefit from your technology. They have a very large database of articles that I'd like to see converted into a directory. Could your software be populated with data other than a pre-existing directory? Could I instead provide my own hierarchial category list and have it work from that? If a category list was used (instead of a directory of sites) would I need to associate keywords with each category first? Thanks for any info.
Could your software be populated with data other than a pre-existing directory?
The purpose of our categorize is to populate any directory structure with any data in any format. The mathmatical logic formula is universal in nature, and thus we have a script that will categorize any amount of data into structured categories.
Could I instead provide my own hierarchial category list and have it work from that?
Definatly, wit ill work on any hierarchial directory structure provided by you, we have also developed a version that will automatically create its own sub categories.
If a category list was used (instead of a directory of sites) would I need to associate keywords with each category first?
We should talk offline regarding this step.
Thanks
Giga
The suggested directory structure and location seemed reasonable and appropriate. However, it also offered numerous other suggestions, most of which were not as appropriate (because they were too specialized, and/or only reflected a small part of the overall content on the test URL I was trying).
I'm a bit unsure about what is going on, and thus how to interpret the results. Were all of the suggested locations taken from DMOZ or some other directory structure that existed prior to my submitting my sample request? Or were the suggestions invented on the fly, in response to my request?
Stated another way, are the suggestions provided to the user choices for sublocations within a pre-existing directory structure that has been predefined (whether by you or by DMOZ)?
If I were judging how well the script placed the URL within the offered set of options, the top choice was certainly the best.
If I am judging how many of the suggested options were appropriate for the URL, I would say relatively few.
If I am judging whether the suggested directory structure was appropriate/good (e.g. this is a computer generated structure and the question is whether it does as good a job as the humans at Yahoo or DMOZ) I would need to test it some more to pass judgment.
2. Can this technology be applied to geogrpahy rather than, or in addition to, subject matters? For instance, could it be adapted to the problem of mass-evaluating whether URLs contain content that is primarily related to a specific geographic area (e.g. finding Dallas Texas portals, travel guides devoted to Texas) or whether it is operated by an entity that is located in or focusing on a specific geographic area (e.g. plumbers and chess clubs that are located in Dallas)?
3. Is this software capable of drawing a distinction between a website and a web page?
As an aside, this is one of the biggest weaknesses with Google and other search engines. They don't draw any such distinction. Most of the results shown to users are based upon the overall site content, but other results are for individual pages. There are perhaps 6 million active domains. When Google talks about indexing 4 billion webpages, they are referring to the pages within those 6 million sites, but their results don't adequately distinguish between the two concepts.
Sometimes I want to find an article or other information about a specific topic (I'm looking for the best webPAGE about discussing widgets) and other times I'm looking for the best webSITES for doing research, or making a purchase, related to a topic (I'm looking for the best place to learn about or purchase widgets.)
There is a big difference between trying to identify the best page within a book, or the best book. Google, et al, fail to recognize this fundamental distinction. If your tool could draw such a distinction efficiently, it might be helpful to someone trying to build a specialized "niche" search engine.
4. I didn't see, or at least didn't recognize, the keyword suggestion tool. Can you clarify where this is, and how it is supposed to work?
I tried a URL for a relatively small site with a well-defined scope, and within less than a minute it gave me back a suggested directory location.
The suggested directory structure and location seemed reasonable and appropriate. However, it also offered numerous other suggestions, most of which were not as appropriate (because they were too specialized, and/or only reflected a small part of the overall content on the test URL I was trying).
**The suggested matches are SIGNIFICALLY weaker potential matches than the primary match. Perhaps we should remove the suggestions (as they were more of a diagnostic tool for us to gauage relevancy and to tweak the algo)
I'm a bit unsure about what is going on, and thus how to interpret the results. Were all of the suggested locations taken from DMOZ or some other directory structure that existed prior to my submitting my sample request?
**Yes from DMOZ in this particular demo.
Or were the suggestions invented on the fly, in response to my request?
**No, but an interesting concept :)
Stated another way, are the suggestions provided to the user choices for sublocations within a pre-existing directory structure that has been predefined (whether by you or by DMOZ)?
**Yes all 450,000+ categories have been pre-defined by DMOZ, however we are developing a version that will create its own sub categories automatically based on the end users suggested top level topics.
If I were judging how well the script placed the URL within the offered set of options, the top choice was certainly the best.
**As stated above it should have a significally higher score. The others are just placed in, to avoid a complete failure to categorize.
If I am judging how many of the suggested options were appropriate for the URL, I would say relatively few.
**We plan to slim down if not remove the suggetions all together.
If I am judging whether the suggested directory structure was appropriate/good (e.g. this is a computer generated structure and the question is whether it does as good a job as the humans at Yahoo or DMOZ) I would need to test it some more to pass judgment.
**In some ways it does make a better match for some sites than a dmoz editor might suggest, in other instances it completely misses the mark, however what makes this unique is that with a high degree of accuracy it can categorize millions of pages in a timely fashion without human intervention.
2. Can this technology be applied to geogrpahy rather than, or in addition to, subject matters? For instance, could it be adapted to the problem of mass-evaluating whether URLs contain content that is primarily related to a specific geographic area (e.g. finding Dallas Texas portals, travel guides devoted to Texas) or whether it is operated by an entity that is located in or focusing on a specific geographic area (e.g. plumbers and chess clubs that are located in Dallas)?
**It already does. Try putting in a popular local college. The results are very much determined by the sturcutre of the dmoz subcategories. In other words if there is a website in Dallas TX and no subcategory provided by DMOZ for a Dallas TX listing to match against in its particular field, well.. the categoriezer cannot find an appropriate match. In some ways it is only as good as the directory it is provided with. In this unique case it is DMOZ
3. Is this software capable of drawing a distinction between a website and a web page?
**That would be a very easy addition to the code.
As an aside, this is one of the biggest weaknesses with Google and other search engines. They don't draw any such distinction. Most of the results shown to users are based upon the overall site content, but other results are for individual pages. There are perhaps 6 million active domains. When Google talks about indexing 4 billion webpages, they are referring to the pages within those 6 million sites, but their results don't adequately distinguish between the two concepts.
**Traditionally directories typically only index the homepage url, and not individual pages.
Sometimes I want to find an article or other information about a specific topic (I'm looking for the best webPAGE about discussing widgets) and other times I'm looking for the best webSITES for doing research, or making a purchase, related to a topic (I'm looking for the best place to learn about or purchase widgets.)
There is a big difference between trying to identify the best page within a book, or the best book. Google, et al, fail to recognize this fundamental distinction. If your tool could draw such a distinction efficiently, it might be helpful to someone trying to build a specialized "niche" search engine.
*That is one of our goals.
4. I didn't see, or at least didn't recognize, the keyword suggestion tool. Can you clarify where this is, and how it is supposed to work?
*We will put this feature back online momentarily.
Thank you for all your interest!
They are used as standard benchmarks for the effectiveness of auto-categorisation algorithms and consist of a training set of pre-classified docs and a test set of docs that your software should then attempt to classify correctly.
The Reuters 21578 collection is one common example.
Using these sets you can begin to quantify just how good your software is (or not) and compare it with the performance of others.
provide any feedback! btw this works on any set of data, not just websites. We are applying it to a seperate project of ours and its doing incredible results. Thank you for your interest everyone. (helps motivate us) :)
I have a web directory. Many links would be suitable for more than one category. Can your software Check all links and put them in all other categories that they relate to or only one category will be suggested.
Well it works as if a human reviewer were attempting to place each link (semi-artifical logic at play). If a link had a apparent and common sense placement into one of your categories or sub-categories it would very likely be placed. However if it had no relation to any of your topics it would fail to place (as there would be no relevant category for it).
Giga
1. Have you done any analysis to estimate the accuracy of your automated sorting process? It doesn't need to be 99% accurate, but the more accurate the better. Have you done any methodical testing to compare how its' sorting compares to what humans would do? (note: you could set up a simple test by asking your software to sort 30 web pages or websites into a classification schema; separately, ask at least 3 different humans to sort the same 30 web pages or websites into the same classification schema. Then compare results.
2. How much computing power does it take to accomplish the sorting? Stated differently, if you devoted a decent quality PC with an adequate amount of RAM to the task of running your software full blast around the clock, how many individual pages within a website could it classify in a 24 hour period, or per hour? How many entire websites (of average size and complexity) could it classify per 24 hours or per hour?
Depending on the answers to these questions, it might or might not be useful (both for my project, and for other applications).
1. Have you done any analysis to estimate the accuracy of your automated sorting process? It doesn't need to be 99% accurate, but the more accurate the better. Have you done any methodical testing to compare how its' sorting compares to what humans would do? (note: you could set up a simple test by asking your software to sort 30 web pages or websites into a classification schema; separately, ask at least 3 different humans to sort the same 30 web pages or websites into the same classification schema. Then compare results.2. How much computing power does it take to accomplish the sorting? Stated differently, if you devoted a decent quality PC with an adequate amount of RAM to the task of running your software full blast around the clock, how many individual pages within a website could it classify in a 24 hour period, or per hour? How many entire websites (of average size and complexity) could it classify per 24 hours or per hour?
1. We have never attempted to go head on head with a human as to accuracy. This would be an EXCELLENT methodology from which to compare the two. In regards to accuracy, we believe we are generally satisfied with the results it is producing. I think the next step may be to follow your advice and attempt a test. I would suggest someone, anyone either sticky me (or post, not sure about the tos though..) 30 random websites. We then take an idependant possibly even DMOZ editor and he attempts to "quickly" categorize all 30 sites. We record the time it takes, and the final results of his attempts for each site. Then run our categorizer on the same pool of data. IF you would like to assist us (or anyone) in this academic endeavor to ensure it is 100% unbiased and objective in nature, we gladly welcome the challenge (i'm curious how it compares as well!).
2. The computing power is an issue. We have found a way to build a clustering of computers to help assist in the task of categorizing websites. IN other words several servers on our network can each work on multiple tasks at 1 time, and all contribute their final results to the total tally. As we have stated from the beginning, many many many probability scenarios, and calculations are being made for each and every site that we attempt to categorize as the computer tries not only to categorize the website, but also to understand the general theme of what the site is about. I am not the programmer, so I do not know specific details regarding computing power, however i do know that there is alot of applications going on in the background as it attempts to find relevancy, also a clustering of servers greatly increases computation time.
-Giga
Honestly your right, i'm talking to myself and hear only my own echo reply. We'll move on from this forum since it feels like no one here really "gets it".
[65.87.196.76...]