Forum Moderators: bakedjake
From what I've seen, most of the alternative search engines tend to be either niche or country based since the resources to do more than one country tend to be expensive. By taking the problem of site acquisition out of the equation, it would make running a small country specific search engine somewhat more economically feasible. But what I am trying to figure out is if it is worth building a service for these country based search engines that provides search lists of country specific URLs on a monthly basis or is it just one of those zero dark thirty mad ideas?
Regards...jmcc
Site submission is only necessary when you want people to "stuff" your engine if your spiders aren't keeping up with the pro-active spidering of sites as they come up.
What this service would do is provide country lists of all active com/net/org/biz/info websites in addition to identified $cctld websites to the search engine operators on a monthly (or periodic) basis. This way any new websites appearing in country X would get into the small search engine index without blind spidering. The difference between this and dmoz seeding for the UK would be about 163k sites versus, at guess, 1M sites. For IE (Ireland) it is the difference between about 9000 sites and about 65000 sites. The number of new Irish domains/sites tends to run at about 2500 per month. Submission based directories here in .ie tend to get about 5 new subs a day if they are lucky. The other important factor is that this would be a pre-checked list of active websites rather than a mix of active and coming soon sites.
The small search engine would still need to spider the sites but it would no longer be dependent on blind spidering to detect new sites or on user submissions to survive.
Regards...jmcc
It is of course the question of cost. How do you plan to finance this operation? Generally small SE’s in the start up phase don’t have much money, sow they may be reluctant to pay for something like this, even is it saves them many by have to crawl less.
[edited by: Woz at 9:00 am (utc) on April 16, 2005]
[edit reason] No self promos please, see TOS#13 [/edit]
It is not a case of setting up something new as much as finding a use for something that already works. The hard part of course is working out a model that startups and small SEs can afford.
Regards...jmcc
For example the project i'm working on, mozdex.com uses nutch and that software allows me to seed from dmoz.org. EVERY time i run a fetch i process only the -topN 1 million urls. The fetcher goes out and grabs the contentm, i update the link db, run an analyze and do the process over and over. This way the "top" sites get linked & spidered quickly and i can grab the bottom feeders when i have time (or not at all).
Ofcourse you can also use time stamps to get update data so you can intelligently refresh your data as well as use ranking to set update frequencies and such.
Eventually you can scale to billions of documents.. not sure its worthwhile if you can smartly index the urls for scoring but only fetch the first x amount of pages. (does anyone click to pay 229384793824234 of the results on google?)
I think you under estimate some of the search technologies out there. "blindly" following links would be dumb, but anything that has an algorithm or process to build, seed & rank the links and process fetch lists from that will grow nicely.
Blind spidering is a pretty accurate way of describing how search engines detect sites by following links and it is exactly what Nutch is doing with the Dmoz seed. This process is well above a simple Dmoz seed in that it is far more highly targeted and fresher.
By creating a clearly defined set of URLs for specific countries, this cuts down the time wasted by small SE operators (often up to six months of blind spidering before coming up with a viable index) and allows them to concentrate on providing a good service for their users.
The other important factor in this is that blind spidering only works for sites that have inbound links. Many new sites do not. Getting to these new sites before Google does would give a country based SE an advantage over the big search engines like Google and Yahoo. If a small country based search engine could have a new site in its index within a day or so of the site going live, it would be quite a marketing advantage.
Regards...jmcc
My approach is more akin to codebreaking (something that I used to do in another life). The standard site acquisition method therefore looks like a brute force attack that tries all combinations in order to derive a key. The approach I am using is to break down the sets of URLs into sets of country specific URLs. In this way it allows small country based (micro rather than macro) search engines to concentrate on their local market. Taking Ireland as an example, Dmoz only has about 10000 Irish related sites. Yet the number of Irish sites that my SE spiders is probably about 80,000. Many of these sites are not linked to other sites. The lack of Dmoz editors in specific country categories means that sites can wait months to appear in Dmoz.
The country based search market is one area where the large players such as Google and Yahoo have to compete on a near level playing field. It is a micro orientated search engine environment rather than a macro orientated search engine environment.
Regards...jmcc
Blind spidering is a pretty accurate way of describing how search engines detect sites by following links and it is exactly what Nutch is doing with the Dmoz seed. This process is well above a simple Dmoz seed in that it is far more highly targeted and fresher.
I'm sure the scalability of a 8 billion page index will be vastly different than mine, but the way i see it is as follows.
Your first "batch" of seed sites will be fetched blindly as you have no score or rank or link count to analyze from. Once you "seed" your database (either from dmoz.org) you have 6-8 million urls you can use to do basic calculations and ranking from. From that data you then build fetchlists of sites you are going to crawl. Your only interested in gathering content for the sites you are fetching but following and gathering links from the outbound links from each site you crawl. When you import & analyze this data you then create a score for those links you gathered and then you can fetch them based upon the calculations you have made for what percentage of your pages will be fetched (or fetch all if you have the resources).
Think of spidering as this.
1. Empty database
2. Seed with dmoz.org data
3. Analyze your seeded database for link scoring
4. Fetch all or -topN links
5. Update DB with fetched segments and NEW links
6. Analyze Score of database
7. Fetch -topN links or un fetched documents
8. Repeate 4-6 million of times...
Your database can have a static "freshness" of 30 days for typical content and as the doc score is increased you can spider merely looking for content update or forced freshness on busy or highly ranked sites.
Typically most spiders will only fetch content for which they have direct initiative to do so, and they will grab all outbound links from that content in order to add that to a links database from which a score will be computated and the spiders next fetch list will be built from.
Not really a dumb process at all, however it could be forced to do so if you are indeed looking to scale for whole internet resources and have the hardware & bandwidth capability to do so.
I'm sure the scalability of a 8 billion page index will be vastly different than mine,The idea of the country target URL list is that the 8 billion pages of mud does not have to sifted to get to that gold of the country URLs list. To give you an idea of the problem - there are approximately 45 Million gtld domains and at least another 20 Million cctlds. Now even with a 50% utilisation that's probably about 32.5 Million potential sites. Now the number of Irish sites is probably about 120K. Extracting 120K or so toplevel Irish sites from that 32.5 Million requires a bit more work than using Dmoz as a seed. The UK would probably have about 8 Million domains out of the 65 Million. Germany would have 10 Million at least. Deriving country URL target lists is a highly complex problem.
Country based search engine operators generally don't have years to play around with Dmoz in order to get a viable product and few hobbyists survive in this highly competitive market. The operational lifetime of a country search engine can be about 18 months. Some, if they are lucky, convert into some hybrid web directory but the number of true country level (spidering) search engines is small.
Your first "batch" of seed sites will be fetched blindly as you have no score or rank or link count to analyze from.This does not distinguish the site acquisition phase from the actual spidering phase. Applying this model to a country based (micro) search engine is a disaster in terms of resources and time to market. The idea here is to create highly targeted country search engines not a macro-search Google clone. The target list will be country relevant URLs that have active content (rather than holding pages/coming soon sites etc).
Typically most spiders will only fetch content for which they have direct initiative to do so, and they will grab all outbound links from that content in order to add that to a links database from which a score will be computated and the spiders next fetch list will be built from.Not exactly. The spiders can be restricted to index a specific page or they can have an unrestricted directive which means that they follow every link. The unrestricted range is a great way to end up with an index full of rubbish. And restricting to IP ranges does not really work well due to a percentage of any country's websites being hosted outside of the country's IP ranges.
Not really a dumb process at all, however it could be forced to do so if you are indeed looking to scale for whole internet resources and have the hardware & bandwidth capability to do so.It is a blind spidering process (especially the idea of following every link). It is a blind process of discovery rather than being dumb. Country based search engine operators do things somewhat differently and most that I know do not consider Dmoz's data quality to be good enough as a seed data set. Sticking the Dmoz data set in as a seed is a very inefficient way of doing things for the obvious reason of there being more sites in country X than in Dmoz's country X sub directory.
Google, Yahoo and the morons in MSN are the dinosaurs. Country level operators win by being smarter at exploiting niche markets and by being better targeted. For us, losing is not an option.
Regards...jmcc
The idea of the country target URL list is that the 8 billion pages of mud does not have to sifted to get to that gold of the country URLs list. To give you an idea of the problem - there are approximately 45 Million gtld domains and at least another 20 Million cctlds. Now even with a 50% utilisation that's probably about 32.5 Million potential sites. Now the number of Irish sites is probably about 120K. Extracting 120K or so toplevel Irish sites from that 32.5 Million requires a bit more work than using Dmoz as a seed. The UK would probably have about 8 Million domains out of the 65 Million. Germany would have 10 Million at least. Deriving country URL target lists is a highly complex problem.
Not as complex as you make it. Seeding by country level TLD's and using regex expressions on your fetch lists and crawlers will do the job nicely. Infact setup your own DNS servers that are only aware of tld's that you wish to spider and force your spiders to use them or proxy everything through a proxy server that only filters on the TLD's you wish to use. Your adding complexity where it doesn't need to be.
Even by spidering "the whole internet" if you build your database correctly you should be able to setup your links, scoring and ranking by TLD and use that as part of your query to tie into the specific countries you wish. There is also ways to indentify content by the character encoding, magic strings, mime types and other methods that you can use on your spider to enhance this method.
Country based search engine operators generally don't have years to play around with Dmoz in order to get a viable product and few hobbyists survive in this highly competitive market. The operational lifetime of a country search engine can be about 18 months. Some, if they are lucky, convert into some hybrid web directory but the number of true country level (spidering) search engines is small.
I'm not sure I follow. It is a simple XML/XSLT conversion to pull out the data you need from Dmoz and using the above steps it would be fairly simple to target even more.
Search engines don't last very long since many of them aren't really true "search engines" but a script or meta directory of the big players. How many of these smaller niche markets actually spider, process, rank and manage there own corpus?
This does not distinguish the site acquisition phase from the actual spidering phase. Applying this model to a country based (micro) search engine is a disaster in terms of resources and time to market. The idea here is to create highly targeted country search engines not a macro-search Google clone. The target list will be country relevant URLs that have active content (rather than holding pages/coming soon sites etc).
You referenced Nutch, and that is what i use. There is a distinguishment between fetchlists and URL/Link databases and it is used internally as a process of calculating which sites to fetch next. You would need the same ranking of results and search distribution layers on a Country by Country basis as you would in a whole internet approach. Infact applying what Nutch, Google, Yahoo et all do to your local region would only add value to your market share.
Search has very little to do with the acquisition of content or targeting of content but ranking, sorting, processing and ease of access.
It is a blind spidering process (especially the idea of following every link). It is a blind process of discovery rather than being dumb. Country based search engine operators do things somewhat differently and most that I know do not consider Dmoz's data quality to be good enough as a seed data set. Sticking the Dmoz data set in as a seed is a very inefficient way of doing things for the obvious reason of there being more sites in country X than in Dmoz's country X sub directory.
As for dmoz.org data, i will stand by it is the best & most easily/readily available data to start from. YOUR process should determine what is/isn't relevent. It's very easy to parse out what you need from the Dmoz.org and apply the techniques of "whole internet search" to build your niche directory.
Discovery through link analysis is about the *ONLY* way to build a map of the sites that compose your internet as you choose to index/search it. Without that no matter how good your data is coming in, you have no relevency to relate that to unless you come up with some fancy shmancy process that the yahoo's and google's haven't been able to think up. Datamining is what search is about. HOw you choose to make a snapshot of what your searching relates directly on how good your results will be.
Google, Yahoo and the morons in MSN are the dinosaurs. Country level operators win by being smarter at exploiting niche markets and by being better targeted. For us, losing is not an option.
Not as complex as you make it. Seeding by country level TLD's and using regex expressions on your fetch lists and crawlers will do the job nicely.Building a good country level search engine is complex. It is not as simple as seeding by country level cctlds because the registries most cctlds, especially in Europe, no longer release lists of their domains. So what you see in directories is only a small fraction of what really exists. And there is a growing trend against interlinking in many business websites. Many of them have links to authority sites and these are not reciprocal links.
Search engines don't last very long since many of them aren't really true "search engines" but a script or meta directory of the big players.What I am talking about are real search engines that actually spider but due to the problem of site acquisition and having to rely on user submissions have no way of competing with the bigger players such as Google. They cease spidering and become hybrid directories. Most of them never gave any serious thought to the site acquisition aspect of running a search engine.
How many of these smaller niche markets actually spider, process, rank and manage there own corpus?The ones that survive - all of them.
Search has very little to do with the acquisition of content or targeting of content but ranking, sorting, processing and ease of access.Good. Now we are getting somewhere. The aquisition phase is separate from the search phase. This is the point I was making.
Discovery through link analysis is about the *ONLY* way to build a map of the sites that compose your internet as you choose to index/search it.No. There are other methods that allow you to build a target list. These other methods are often far more useful when it comes to building a good country level search engine. Just to outline the situation again - this is building a country level search engine with limited time and resources not a well funded macro-search engine. And of course link analysis only works when there are links to analyse.
MSN are far from morons. MSN did a good job for its first release and it's a pretty darned good effort.Read some of the threads here about how its badly written spider chewed up bandwith allowances and was searching for non-existent URLs due to corruption in its target list db. Treating webmasters like that is evidence of the MSN people being morons. The result was that webmasters started banning MSNbot.
Regards...jmcc
Building a good country level search engine is complex. It is not as simple as seeding by country level cctlds because the registries most cctlds, especially in Europe, no longer release lists of their domains. So what you see in directories is only a small fraction of what really exists. And there is a growing trend against interlinking in many business websites. Many of them have links to authority sites and these are not reciprocal links.
The probability of hitting a "dead end" is nill when it comes to linking. You can garnish enough links from the tld's through the methods i have explained to cover the entire spectrum easily.
You could also use the google api and yahoo api to pull down 1k links a day to add to your engine in you wanted to seed from more existing materials.
Good. Now we are getting somewhere. The aquisition phase is separate from the search phase. This is the point I was making.
I think you misunderstood me. Search has nothing to do with seeding your database but everything to do with how you process the information you have. If you can't seed with dmoz.org and use that to branch out then you have a flaw in your design.
Don't forget while some authorities may be dead ends there are many many many inbound links that point to them so even finding sites that don't link out to others is simply a process of link analysis and mapping.
No. There are other methods that allow you to build a target list. These other methods are often far more useful when it comes to building a good country level search engine. Just to outline the situation again - this is building a country level search engine with limited time and resources not a well funded macro-search engine. And of course link analysis only works when there are links to analyse.
We are talking about a search engine that searches and provides results that return a URL right? That in itself is a link. Can you please explain other methods of sorting and providing results and such that don't use link analysis? either from incoming/outgoing links to link length, link keyword density, link age, link depth and such?
An authority site is an authority site because of the links granted to that site, and for no other reason.
Are you talking about building a human edited search engine and manually ranking your sites based upon how you determine what is "authority"?
Read some of the threads here about how its badly written spider chewed up bandwith allowances and was searching for non-existent URLs due to corruption in its target list db. Treating webmasters like that is evidence of the MSN people being morons. The result was that webmasters started banning MSNbot.
Hardly moronic, it happens with everyone. They obey all the robots.txt's i have created and it's easy enough to block the ip's if you wish.
I'm not disputing what your effort is, i just think there is a much simpler way of doing it. Bandwidth, hardware and facilities are cheap these days. I pay flat rate for 20mbps and i have a half dozen servers and i keepup with 100million pages just running updates while i work out the bugs, nothing really 24/7 or fully scripted yet.
Using DNS servers that return only results for your tld, using proxy servers to filter out stuff you may/may not need all help alleviate from pre/post processing that can be heavy on the "backend" that does your ranking & calculations.
Infact i would go as much to say that if you provide me with a country level TLD, i'll setup a mini search, spider those tld's and write up a process that describes the trials and tribulations of such.
I like to think of internet searching as intelligent design. You can use clustering, ontology, mapping, geo-targeting, ip's, hostnames, tld's, language detection and much much more to build a pretty nice niche search engine as you are looking to do.
The probability of hitting a "dead end" is nill when it comes to linking. You can garnish enough links from the tld's through the methods i have explained to cover the entire spectrum easily.Link analysis only works where there are links inward to a site. Trying to map the web based purely on Search Results link analysis will not allow you to cover the entire spectrum. Building a good country level index requires other techniques as well.
We are talking about a search engine that searches and provides results that return a URL right? That in itself is a link.What we've got here is a chicken and egg problem. The URL has to be indexed by the search engine for the link to exist. I am talking about building a country level search engine and how to go about creating a target URL list cheaply, effectively and quickly.
Can you please explain other methods of sorting and providing results and such that don't use link analysis? either from incoming/outgoing links to link length, link keyword density, link age, link depth and such?
Hardly moronic, it happens with everyone. They obey all the robots.txt's i have created and it's easy enough to block the ip's if you wish.Tell that to the thousands of webmasters who were affected by the MSN idiocy and had to pay for it.
I'm not disputing what your effort is, i just think there is a much simpler way of doing it.
The more effective way is to split the site acquisition phase from the search phase leaving the country level search engine operator with a continually updated target list that does not depend on blind spidering. In doing so, the initial seed would save months of blind spidering and give the country level search engine a chance of surviving. Think of it as a trickle down system.
The way of the big search engines is to index everything and hope that the algorithms and search results link analysis will provide context. This is the way that you seem to be approaching the problem of building a target list as well.
Country level search engine operators have a clarity of objective that is denied to the macro search engines. The country level operators have to make everything in their country searchable whereas the macro search engines have to make everything searchable. The original idea here was to provide a ready to spider URL target list for these search engine operators that was far more comprehensive than a Dmoz seed. Perhaps survival has made me cynical but I've seen too many country level search engines fail over the years.
Regards...jmcc
Link analysis only works where there are links inward to a site. Trying to map the web based purely on Search Results link analysis will not allow you to cover the entire spectrum. Building a good country level index requires other techniques as well.
It would be safe to assume 99.9999999999999999999999 percent of all sites have atleast one link to them.
My utilization of link analysis is both for doc scoring but also for spidering. Link analysis helps build an intelligent and smart spider solution to grab documents that match a score based upon the methods you have used to score them. Thus going back to my theory that bots/spiders aren't merely dumb and fetching everything.
What we've got here is a chicken and egg problem. The URL has to be indexed by the search engine for the link to exist. I am talking about building a country level search engine and how to go about creating a target URL list cheaply, effectively and quickly.
Again, use filtering methods to limit your fetch process to the topology/networks you wish to fetch. I don't see any way around using links for url discovery since it *IS* a critical component of search engine design.
You seem to be considering link analysis in terms of search results and PR rather than in terms of discovery here.
The difference is that I know which way works more effectively for a country level operator having tried both. Survival as a search engine operator is all about finding easy wins and innovative ways to dominate a niche.
I don't see a difference in local search vs internet search at all other than an imposed limit on the content your are indexing (whatever you choose that to be). For local search to suceed it's not a matter of a good list of urls to build from but the quality of service you provide.
Search engine survival has little to do with the process you use to build your fetch list but how you handle the data you have already fetched and what you provide as a value add. When people search they're looking for something. They don't judge what is in your index as much as they judge the results that lead them to answer the question they are searching for.
Use of ontology, clustering, stemming, mapping, geo-targeting on top of all the standard index processes is what builds a nice niche engine. If you build a search engine for Pennsylvania and you use ontology to describe much of the common pennsylvania particulars and tie that to your search then you are providing a value add that enhances the results for what people are specifically looking for. If you map out PA tourism, culture, cities, counties, boroughs and all that jazz and use that to correlate to the web THAT is what niche search is about.
The way of the big search engines is to index everything and hope that the algorithms and search results link analysis will provide context. This is the way that you seem to be approaching the problem of building a target list as well.
Not at all. Using processes described above as part of your index process and fetch process you can build a highly targeted fetch list and you can most certainly use dmoz.org data to seed this process and grab most, if not all of the content you need.
Country level search engine operators have a clarity of objective that is denied to the macro search engines. The country level operators have to make everything in their country searchable whereas the macro search engines have to make everything searchable. The original idea here was to provide a ready to spider URL target list for these search engine operators that was far more comprehensive than a Dmoz seed. Perhaps survival has made me cynical but I've seen too many country level search engines fail over the years.
Again, how you use the data and the tools you develop to mine the data is what makes a search engine survive, especially in a niche market. The data is out there and there are intelligent ways to spider it and use the links and paths that already exist. Reinventing the wheel on URL discovery does NOTHING to solve the niche market needs of search result capabilities & features.
That is where i'm coming from. Intelligent design is a process that works from the ground up to solve your business needs. Whole internet searching is different in scale to your niche needs, but otherwise very similar :)
I think link discovery and url seeding is the smallest of all (if any) advantages any search engine has. Your niche market isn't a TLD search but how you relate searching to what the TLD is. (use of linguistics, locality, culture, heritage and all that jazz)
i think it's got potential.
i don't want to spider the whole internet to find every site in the uk. if i had your list i could simply spider the X million active websites in the list and job done. no need to follow all external links on every site.
It would be safe to assume 99.9999999999999999999999 percent of all sites have atleast one link to them.You do know the old saying about the word "assume". :) It would not be safe to assume that all or as good as all sites have one inbound link. What a country level search engine operator has got to do is to move beyond that kind of thinking in order to build a superior index to Google and the big players.
Thus going back to my theory that bots/spiders aren't merely dumb and fetching everything.The reality is that bots and spiders are working blind with directions and restrictions that set the depth to which they spider a site and the distance away that they follow links.
I don't see any way around using links for url discovery since it *IS* a critical component of search engine design.But you've already found the way around it. :) By using Dmoz as the seed, you have in effect bootstrapped the acquisition phase.
The original idea, (at the top of the thread), was to provide country level search engine operators with a far more effective and highly targeted seed of relevant country URLs. The difference is between using a bicycle and using a ramjet. The targeted country URL list/seed will allow a country level search engine operator to get up to speed in a far shorter time than using Dmoz.
Using processes described above as part of your index process and fetch process you can build a highly targeted fetch list and you can most certainly use dmoz.org data to seed this process and grab most, if not all of the content you need.Building a target list in this manner is a highly recursive process. The idea is to remove most of that from the process for a country level search engine operator and allow them to concentrate on building a good index.
I think link discovery and url seeding is the smallest of all (if any) advantages any search engine has. Your niche market isn't a TLD search but how you relate searching to what the TLD is.I'd disagree with this because the advantage of a country level search engine has is that it is local and able to update quicker. Thus a fresher local index is going to give it an advantage. When dealing with a country, it is a cctld that makes the difference. A country code top level domain where the registry does not provide open access to the zonefiles or domain lists means that the country level search engine and the big players like Google and Yahoo are equally in the dark. If a small player can get an advantage here then it will provide a unique selling point. Taking that idea further, if newer country level sites can be picked up before they become properly linked, that takes the link factor out of detecting new sites. And that is one hell of an advantage to have over Google and Yahoo.
Regards...jmcc
You do know the old saying about the word "assume". :) It would not be safe to assume that all or as good as all sites have one inbound link. What a country level search engine operator has got to do is to move beyond that kind of thinking in order to build a superior index to Google and the big players.
Assuming because of years of experience, history and speaking with people on a subjet matter is different than assuming otherwise :)
Since the internet is never a single entity that you can easily comprehend or scale for any long period of time, assuming is all you have; you can be educated in what you assume.
The reality is that bots and spiders are working blind with directions and restrictions that set the depth to which they spider a site and the distance away that they follow links.
I think you need to re-read the process used to build fetch lists and spider sites and try it yourself. It's hardly blind and very configurable for whatever process you want to put behind it.
It's not just "depth" or "distance" that spiders follow, its the list they're told to consume and how you use the NEW data that they find to process your consumed list you developed and forced down the spiders throat.
But you've already found the way around it. :) By using Dmoz as the seed, you have in effect bootstrapped the acquisition phase.
Dmoz gives me a few million urls that i start from. I compoud the list almost 10 fold for every 1 new site i find (because of the outbound links from that site that aren't seeded directly from dmoz.org data).
Dmoz boot straps my database, just as it could yours. If your just interested in your .co.uk or .au or whatever tld you wish, you simply regex the import to only grab those domains and have regex rules on your spider, and fetchlists so you only consume & process hosts related to your niche or in your case tld. Really not rocket science here :)
The original idea, (at the top of the thread), was to provide country level search engine operators with a far more effective and highly targeted seed of relevant country URLs. The difference is between using a bicycle and using a ramjet. The targeted country URL list/seed will allow a country level search engine operator to get up to speed in a far shorter time than using Dmoz.
I don't think we are talking on the same issue here. Search has 0.. ZERO to do with what you feed your index but HOW YOU USE the data you have. Google isn't popular because they have billions of pages, but because you can find what your looking for in those billions of pages. It is the process of finding what your looking for that makes ANY search engine relevent and lasting because that is what it is about.
If your just a searchable directory of websites limited to a tld then that is NOT the same search i am speaking of.
Building a target list in this manner is a highly recursive process. The idea is to remove most of that from the process for a country level search engine operator and allow them to concentrate on building a good index.
Recursive processing is how you get the data to be relevent. Your seeeded links have no value until you recursively process them (even multiple times) to find the topology and relevency of the linked subjects.
I'd disagree with this because the advantage of a country level search engine has is that it is local and able to update quicker. Thus a fresher local index is going to give it an advantage. When dealing with a country, it is a cctld that makes the difference.
Google has no problems keeping 8 billion pages fresh almost daily. Your freshness processing goes directly into your link processing and your web page rank. The higher your rank, the more you get fetched. Doesn't have much to do with the locality or scale of the project as the same would hold true for a country level search engine.
A country code top level domain where the registry does not provide open access to the zonefiles or domain lists means that the country level search engine and the big players like Google and Yahoo are equally in the dark. If a small player can get an advantage here then it will provide a unique selling point. Taking that idea further, if newer country level sites can be picked up before they become properly linked, that takes the link factor out of detecting new sites. And that is one hell of an advantage to have over Google and Yahoo.
This goes back to my smart fetching process. You only seed from dmoz what is valid for your seeding and if your dns server or proxy server is only aware of your countries TLD's or that of which your searching it will ignore all non relevent links and only process what it is aware of. A quick and simple fix to expensive rules base processing on your fetching/indexing cycle and a lot cheaper than manual intervention.
I don't doubt country based search at all. more power to you for looking at it that way, but i will stand by my beliefs that you can't ignore the semantics of the web just because your focus is a single tld system. You can weight the semantics differently and tweak your search results but the metods behind that are largely the same.
Again, i think what will make a great country wide search engine is someone who takes the semantics of the web and mixes that with the semantics of the country they are focusing on with intelligent use of clustering, ontology, smart weighting and a good common words and dictionary list. Hardly related to seeding your system or basing your relevency upon what you feed it instead of what you process from it.
Assuming because of years of experience, history and speaking with people on a subjet matter is different than assuming otherwise :)But without any apparent experience of building a country level index you are in danger of going from assumption to pure guesswork. :)
Since the internet is never a single entity that you can easily comprehend or scale for any long period of time, assuming is all you have; you can be educated in what you assume.Despite what some people think, the internet is a very simple model. It is just the magnitude of the model that scares people. It is actually quite structured - otherwise the network would be a notwork.
The problem with country level indices is that you have the cctld index, which is easily handled by simple regexps, and the parallel com/net/org/biz/info country level index.
Dmoz gives me a few million urls that i start from. I compoud the list almost 10 fold for every 1 new site i find (because of the outbound links from that site that aren't seeded directly from dmoz.org data).No matter how you present it, it is still the equivalent of a brute force attack on the problem, trying all links to detect others. In market and business terms, speed is everything and the sooner a search engine goes operational, the sooner it begins to start making money. This is the business of search.
Dmoz boot straps my database, just as it could yours. If your just interested in your .co.uk or .au or whatever tld you wish, you simply regex the import to only grab those domains and have regex rules on your spider, and fetchlists so you only consume & process hosts related to your niche or in your case tld. Really not rocket science here.Yeah - bicycle science really. It is slow but it eventually gets there if it ever really knew where there was. :)
Going back to a country level index being two parallel indices, the simplistic .cctld regexp ignores a pile of com/net/org/biz/info domains. In some cases that is a very significant number of sites. Often then combined number of com/net/org/biz/info sites for a country is more than the number of cctld sites. Now do you see where your blind spidering strategy develops problems? How does it distinguish between a US hosted $country .com site and a US owned site?
I don't think we are talking on the same issue here.I am talking about getting a far better country level index to a search engine operator that would allow him to get his search engine operational quicker and start making money. Search is the end product but in order to provide a good search engine, the seed index has to be good. As a macro search seed, Dmoz is good. As a country level seed, it is very low quality. Country level search is not just about a single cctld. Applying your brute force attack strategy to country level search runs into the GIGO problem very quickly. That's why the acquistion phase is very important when it comes to building a good country level index.
Regards...jmcc
But without any apparent experience of building a country level index you are in danger of going from assumption to pure guesswork. :)
I beg to differ. If this is what you still interpret from our huge discussion then i don't think your hearing me.
Despite what some people think, the internet is a very simple model. It is just the magnitude of the model that scares people. It is actually quite structured - otherwise the network would be a notwork.
Exactly, it's structured. That is where semantics come in. The structer that a big search engine uses over a little is only in scale. My issue was that even country level based searches could scale to be the size of google and your point of the "big boys" is moot since in 5-10 years your local search may have all the "Crap" of a big search as the internet "explodes" into something you or I can't control. Thus the incomprehensable scale and agility.
No matter how you present it, it is still the equivalent of a brute force attack on the problem, trying all links to detect others. In market and business terms, speed is everything and the sooner a search engine goes operational, the sooner it begins to start making money. This is the business of search.
No, your just failing to notice or look into an automated method. It isn't brute force when you have a methodology and process that works. I could build a country level tld and have 99.99999% of the URLS/Sites pulled for lass than 2 grand and be operational in 2 weeks and i could incorporate the features i spoke of before to tune it to that market.
I am talking about getting a far better country level index to a search engine operator that would allow him to get his search engine operational quicker and start making money. Search is the end product but in order to provide a good search engine, the seed index has to be good. As a macro search seed, Dmoz is good. As a country level seed, it is very low quality. Country level search is not just about a single cctld. Applying your brute force attack strategy to country level search runs into the GIGO problem very quickly. That's why the acquistion phase is very important when it comes to building a good country level index.
I still don't think we are talking on the same page here. I've given you the methods to building a great search engine and i still don't know what your leaning to other than a human edited directory because otherwise is "too expensive" "brute force" and "dumb".
I personally think your failing to read my messages and ignoring what really sells a search engine. The cost of market entry for your idea is no less or greater than the cost of any other method i have proposed/suggested. Infact you could use one of the MANY free/opensource/apache foundation programs to get seeded and expirment from.
Give me a tld or country code and i'll have a search engine up in two weeks and you can see if it's up to snuff.
beg to differ. If this is what you still interpret from our huge discussion then i don't think your hearing me.Have you actually build a country level index? Not a simple country code tld (eg .uk or .pl or .cn) restricted index but an actual country level index consisting of $cctld,com,net,org,biz,info sites associated with that country? Or are you looking at this as being a simple case of restricting sites to the cctld extension of the particular country?
My issue was that even country level based searches could scale to be the size of google and your point of the "big boys" is moot since in 5-10 years your local search may have all the "Crap" of a big search as the internet "explodes" into something you or I can't control. Thus the incomprehensable scale and agility.I disagree. Local search will lead to an ordered fragmentation of search where macro search engines will still exist but more localised micro search engines will carry out local search. At a higher level these will be aggregated into large, macro search engines. Over the last five years in the search business, this pattern is playing out quite well.
No, your just failing to notice or look into an automated method. It isn't brute force when you have a methodology and process that works.I have not excluded the automated method. I just know from experience that there are simpler solutions to building a country level search index.
A Brute Force Attack is when you try all possible keys to decrypt a message. In this case, it is following all links and hoping that the search technology and algorithms will give the resulting pile of data context.
Simply restricting the index to $cctld is basically doing what Google, Yahoo and Altavista were doing about five years ago when it came to searching country code top level domains. A country level search index is more than the cctld range of sites.
I've given you the methods to building a great search engineThe methods you are talking about are common knowledge and nothing new. It is not like you've discovered something that other search engine operators do not know or have not being doing for the last n years.
i still don't know what your leaning to other than a human edited directory because otherwise is "too expensive" "brute force" and "dumb".
1: Getting a good, clean, country level search index to search engine operators that can be used immediately.
2: Building that index rapidly with techniques that are more efficient than simply spidering blind.
3: To detect and include new country related sites in this country level index before the large search engines do.
4: To allow small search engine operators to compete with the big players on a more level playing field.
5: A mug of coffee that I forgot to drink when typing this. :)
The country level operators using this index will still have to spider and will build up more sites on the basis of this seed index. It is a more highly targeted seed index than Dmoz.
Regards...jmcc
Have you actually build a country level index? Not a simple country code tld (eg .uk or .pl or .cn) restricted index but an actual country level index consisting of $cctld,com,net,org,biz,info sites associated with that country? Or are you looking at this as being a simple case of restricting sites to the cctld extension of the particular country?
Either way, there is easy methods to getting the information you need. I can buy a geo-ip database and mine the urls that through smart spidering processes for under 300.00 us.
TLD based mining is an easy start, if you want to incorporate all other tld's within your locality it would just be another data mining process on top of the standard processes any search engine already uses.
I disagree. Local search will lead to an ordered fragmentation of search where macro search engines will still exist but more localised micro search engines will carry out local search. At a higher level these will be aggregated into large, macro search engines. Over the last five years in the search business, this pattern is playing out quite well.
In a sense this is how big search engines build are built. I think locality is 1% of what your people are searching for. If i'm looking for "shoe repair" and i live in philadelphia i can get on google and search for "shoe repair philadelphia pa" and have a response in less than a 10th of a second. The beauty of the internet is it makes the world closer. Because i'm looking for shoe repair i can find someone who is cheaper or better at it no matter where they are and refine that to something that is better suited to what i'm searching for.
I have not excluded the automated method. I just know from experience that there are simpler solutions to building a country level search index.
Please explain.. thats why i'm responding to your threads. If you know of a simpler solution to smart automated processes that don't cost more then a few bucks of bandwidth and a 350.00 pc than let me know.
A Brute Force Attack is when you try all possible keys to decrypt a message. In this case, it is following all links and hoping that the search technology and algorithms will give the resulting pile of data context.
That is a complete mis understanding of the process and a false statement and that is what i have been disputing with you. The process of link following and link discovery is a process of search and fundamental to the semantics of the web no matter if your global, local or municipal. The web is something that morphs and changes and you have to have a process that can give weights and balances to this morphing and score the documents accordingly.
The web isn't a location for you to make a billboard to and to provide a constant fixed methodology to. Your entry will change as the world changes and you have to have a process that can update and quantify the information. Link analysis and discovery is a small process of the entire concept of search.
Simply restricting the index to $cctld is basically doing what Google, Yahoo and Altavista were doing about five years ago when it came to searching country code top level domains. A country level search index is more than the cctld range of sites.
I don't doubt that, but it's a start. You could even go to your country level pages in the dmoz.org data and start from there. My point is that you have to impose rules on the data that you are collecting, but the data itself isn't nearly as important as what you do with it. You seem to give importance on something i don't grasp and that is what i'm asking for your feedback on. I don't understand what your scoring process is - that is what i'm trying to dig into.
I also believe you continue to under estimate the power of the "big boys" out there and the processes used to build the indexes they use. Brute force in whole web searching is cheap, but i know for sure they have very intelligent methods and processes in place to build & maintain such complex systems. I know this from speaking to engineers, going to tradeshows and trying to learn and run my own index as well.
1: Getting a good, clean, country level search index to search engine operators that can be used immediately.2: Building that index rapidly with techniques that are more efficient than simply spidering blind.
3: To detect and include new country related sites in this country level index before the large search engines do.
4: To allow small search engine operators to compete with the big players on a more level playing field.
5: A mug of coffee that I forgot to drink when typing this. :)
I can use some coffee thats for sure. What your asking for is something i've already described and can be done today using stuff like Nutch with a combination of other systems & processes.
Buy some data from netcraft, use a geo-ip database, start with a tld that is local, limit your search through smart proxies & dns systems tied to the above data and use the standard processes that search engines use so you can focus on your service instead of re-inventing the semantics of the web as it is to be viewed but creating your own semantics of how it is to be represented to the audience that you are targeting.
You can't market your index is clean, people don't care.
You can't market that you are faster at getting new sites, people don't care.
You can't market your idea on the principal of competing with the big boys when you try and compare your niche market to a market that has no relevency.
You have to sell your service, by providing relevent and timely results as well as giving value to those results that they don't get anywhere else.
Google WILL have everything and then some that your niche Search will have, same with Yahoo and MSN and a few others simply because they can and they have the processes and know how to do so.
Local search won't survive because they define the value they offer as a marketing strategy, but because of the services they offer and the value they add to the industry that differs them from the others.
You have to remember Yahoo and Google already have the internet indexed as a foundation and it's easy for them to process the information they aready use and gather to create sub-search indexes based upon locality, language and other demographics.
I can buy a geo-ip database and mine the urls that through smart spidering processes for under 300.00 us.Many of those geo-ip databases are based on the delegated ranges and often do not have the granularity for serious geo-ip work. A good geo-ip database needs to be updated continually.
TLD based mining is an easy startThe TLD mining aspect is more complicated than ordinary search engine work though it takes a long time to get it right.
Because i'm looking for shoe repair i can find someone who is cheaper or better at it no matter where they are and refine that to something that is better suited to what i'm searching for.Yeah but if you are looking for shoe repair in Philadelphia because your shoe has sprung a leak in the middle of the street, are you going to go to New York for the cheaper repair? :) Good local search is about giving the user the local information they need when they want it.
Please explain..pm sent.
That is a complete mis understanding of the process and a false statement and that is what i have been disputing with you.Ok then we'll have to agree to disagree. I just think that purely relying on link analysis is a brute force attack method of discovery.
The web isn't a location for you to make a billboard to and to provide a constant fixed methodology to.Do a historical analysis on the rate of change of content and you'll find that the majority of the web changes on a yearly or near yearly basis with only a minority of sites having active content that regularly changes.
Link analysis and discovery is a small process of the entire concept of search.But search is the end result of that link analysis and discovery.
Buy some data from netcraft,I don't quite trust Netcraft's methodology when it comes to websites.
the audience that you are targeting.The audience I am targeting consists of country level search engine operators that need a good seed index for their country. I'm not targeting the end user.
I'll end with what i've been saying. You build your system in much the same way any other search would, but the power you give on top of the processes already in place is what defines your niche market.
If your defining yourself because of the limited subset of data, google could rock your world in a matter of days if you were a bleep on there radar by incorporating some of the processsing and intelligence on top of the dataset they already have to blow you out of the water.
Clustering, stemming, ontology, and relating your data to the people you are servicing is what makes a niche search engine exist.
The semantics of link analysis, ranking, processing and summarizing is all tweakable but hardly a defigning feature to base yourself on. Just a fundemental process of building a search engine and utilizing the fetch & repeat process that builds up the data you use to intelligently map what you are trying to achieve.
Best of luck! :)
I'm sure we're putting on a great show here :)
Yes, you 2 win the all time award for loving to write!
Would the proposed list service work for a niche search engine rather than local? Say for Tennis, or Plumbing?
That is, can you find/track/update sites in one of these categories rather than places?
Would the proposed list service work for a niche search engine rather than local? Say for Tennis, or Plumbing?
That is, can you find/track/update sites in one of these categories rather than places?
Regards...jmcc