A web search engine is a software system that is designed to search for information on the World Wide Web. The search results are generally presented in a line of results often referred to as search engine results pages (SERPs). The information may be a mix of web pages, images, and other types of files. Some search engines also mine data available in databases or open directories. Unlike web directories, which are maintained only by human editors, search engines also maintain real-time information by running an algorithm on a web crawler.
Internet search engines themselves predate the debut of the Web in December 1990. The Whois user search dates back to 1982 [1] and the Knowbot Information Service multi-network user search was first implemented in 1989.[2] The first well documented search engine that searched content files, namely FTP files was Archie, which debuted on 10 September 1990.[citation needed]
Prior to September 1993 the World Wide Web was entirely indexed by hand. There was a list of webservers edited by Tim Berners-Lee and hosted on the CERN webserver. One historical snapshot of the list in 1992 remains,[3] but as more and more web servers went online the central list could no longer keep up. On the NCSA site, new servers were announced under the title "What's New!"[4]
The first tool used for searching content (as opposed to users) on the Internet was Archie.[5] The name stands for "archive" without the "v". It was created by Alan Emtage, Bill Heelan and J. Peter Deutsch, computer science students at McGill University in Montreal. The program downloaded the directory listings of all the files located on public anonymous FTP (File Transfer Protocol)
sites, creating a searchable database of file names; however, Archie
did not index the contents of these sites since the amount of data was
so limited it could be readily searched manually.
The rise of Gopher (created in 1991 by Mark McCahill at the University of Minnesota) led to two new search programs, Veronica and Jughead. Like Archie, they searched the file names and titles stored in Gopher index systems. Veronica (Very Easy Rodent-Oriented Net-wide Index to Computerized Archives) provided a keyword search of most Gopher menu titles in the entire Gopher listings. Jughead (Jonzy's Universal Gopher Hierarchy Excavation And Display)
was a tool for obtaining menu information from specific Gopher servers.
While the name of the search engine "Archie" was not a reference to the
Archie comic book series, "Veronica" and "Jughead" are characters in the series, thus referencing their predecessor.
In the summer of 1993, no search engine existed for the web, though numerous specialized catalogues were maintained by hand. Oscar Nierstrasz at the University of Geneva wrote a series of Perl scripts that periodically mirrored these pages and rewrote them into a standard format. This formed the basis for W3Catalog, the web's first primitive search engine, released on September 2, 1993.[6]
In June 1993, Matthew Gray, then at MIT, produced what was probably the first web robot, the Perl-based World Wide Web Wanderer,
and used it to generate an index called 'Wandex'. The purpose of the
Wanderer was to measure the size of the World Wide Web, which it did
until late 1995. The web's second search engine Aliweb appeared in November 1993. Aliweb did not use a web robot,
but instead depended on being notified by website administrators of the
existence at each site of an index file in a particular format.
JumpStation (created in December 1993[7] by Jonathon Fletcher) used a web robot to find web pages and to build its index, and used a web form
as the interface to its query program. It was thus the first WWW
resource-discovery tool to combine the three essential features of a web
search engine (crawling, indexing, and searching) as described below.
Because of the limited resources available on the platform it ran on,
its indexing and hence searching were limited to the titles and headings
found in the web pages the crawler encountered.
One of the first "all text" crawler-based search engines was WebCrawler,
which came out in 1994. Unlike its predecessors, it allowed users to
search for any word in any webpage, which has become the standard for
all major search engines since. It was also the first one widely known
by the public. Also in 1994, Lycos (which started at Carnegie Mellon University) was launched and became a major commercial endeavor.
Soon after, many search engines appeared and vied for popularity. These included Magellan, Excite, Infoseek, Inktomi, Northern Light, and AltaVista. Yahoo! was among the most popular ways for people to find web pages of interest, but its search function operated on its web directory,
rather than its full-text copies of web pages. Information seekers
could also browse the directory instead of doing a keyword-based search.
A search engine maintains the following processes in near real time:
Web search engines get their information by web crawling from site to site. The "spider" checks for the standard filename robots.txt, addressed to it, before sending certain information back to be indexed depending on many factors, such as the titles, page content, JavaScript, Cascading Style Sheets (CSS), headings, as evidenced by the standard HTML markup of the informational content, or its metadata in HTML meta tags.
Indexing means associating words and other definable tokens found on
web pages to their domain names and HTML-based fields. The associations
are made in a public database, made available for web search queries. A
query from a user can be a single word. The index helps find information
relating to the query as quickly as possible.[14]
Some of the techniques for indexing, and cacheing are trade secrets, whereas web crawling is a straightforward process of visiting all sites on a systematic basis.
Between visits by the spider, the cached version of page (some
or all the content needed to render it) stored in the search engine
working memory is quickly sent to an inquirer. If a visit is overdue,
the search engine can just act as a web proxy instead. In this case the page may differ from the search terms indexed.[14]
The cached page holds the appearance of the version whose words were
indexed, so a cached version of a page can be useful to the web site
when the actual page has been lost, but this problem is also considered a
mild form of linkrot.
Typically when a user enters a query into a search engine it is a few keywords.[15] The index
already has the names of the sites containing the keywords, and these
are instantly obtained from the index. The real processing load is in
generating the web pages that are the search results list: Every page in
the entire list must be weighted according to information in the indexes.[14] Then the top search result item requires the lookup, reconstruction, and markup of the snippets
showing the context of the keywords matched. These are only part of the
processing each search results web page requires, and further pages
(next to the top) require more of this post processing.
Monday, 2 May 2016
Thursday, 28 April 2016
How did the Internet start?
Mention the history of the Internet
to a group of people, and chances are someone will make a snarky
comment about Al Gore claiming to have invented it. Gore actually said
that he "took the initiative in creating the Internet" [source: CNN]. He promoted the Internet's development both as a senator and as vice president of the United States. So how did the Internet really get started? Believe it or not, it all began with a satellite.
It was 1957 when the then Soviet Union launched Sputnik, the first man-made satellite. Americans were shocked by the news. The Cold War was at its peak, and the United States and the Soviet Union considered each other enemies. If the Soviet Union could launch a satellite into space, it was possible it could launch a missile at North America.
It was 1957 when the then Soviet Union launched Sputnik, the first man-made satellite. Americans were shocked by the news. The Cold War was at its peak, and the United States and the Soviet Union considered each other enemies. If the Soviet Union could launch a satellite into space, it was possible it could launch a missile at North America.
Why did Internet start ?
The
Internet started out as an american military project its original
purpose was to enable different military bases and later academic
centers to communicate with each other, and share information without
the need to travel. It was started in the 1960's and at first used very
complex protocols , which were later standarized by the introduction of
the WWW we use today, and the Http protocols along with the e-mail
services, at later points Internet introduced data sharing and the
ability of video streaming, which along side the high speed internet
that is used in more countries each day starts to provide a form of
entertainment,another ground breaking invention for the internet was the
ability to instantly communicate with anyone around the world, which
was developed by first creating forums and later on instant chat
messangers ,culminating in the now popular social networks that are a
hybrid between a chat,mail and forum services , providing also streamed
videos. The internet is sharing information and its INFORMATION its a
storage of all kinda information. Internet is also communication its a
line for sharing information and creating new information that can be
shared with everybody around the globe its the bibble of the new era.
Internet has information to offer and everything that was either banned,
removed remains un-used. Everything produced outside the so called
entertainment industry ( like anime or Internet animation) and its
also a new form of society that is connected to sharing different kinda
of data ...its a second spiritual,introverted and matriarchal society
that takes it upon itself to archivize the human civilisation and its
own existence storing it for future generations.
Where did Internet start ?
The history of the Internet begins with the development of electronic computers in the 1950s. Initial concepts of packet networking originated in several computer science laboratories in the United States, United Kingdom, and France.[1]
The US Department of Defense awarded contracts as early as the 1960s
for packet network systems, including the development of the ARPANET (which would become the first network to use the Internet Protocol). The first message was sent over the ARPANET from computer science Professor Leonard Kleinrock's laboratory at University of California, Los Angeles (UCLA) to the second network node at Stanford Research Institute (SRI).
Packet switching networks such as ARPANET, NPL network, CYCLADES, Merit Network, Tymnet, and Telenet, were developed in the late 1960s and early 1970s using a variety of communications protocols.[2] Donald Davies was the first to put theory into practice by designing a packet-switched network at the National Physics Laboratory in the UK, the first of its kind in the world and the cornerstone for UK research for almost two decades.[3][4] Following, ARPANET further led to the development of protocols for internetworking, in which multiple separate networks could be joined into a network of networks.
Access to the ARPANET was expanded in 1981 when the National Science Foundation (NSF) funded the Computer Science Network (CSNET). In 1982, the Internet protocol suite (TCP/IP) was introduced as the standard networking protocol on the ARPANET. In the early 1980s the NSF funded the establishment for national supercomputing centers at several universities, and provided interconnectivity in 1986 with the NSFNET project, which also created network access to the supercomputer sites in the United States from research and education organizations. Commercial Internet service providers (ISPs) began to emerge in the very late 1980s. The ARPANET was decommissioned in 1990. Limited private connections to parts of the Internet by officially commercial entities emerged in several American cities by late 1989 and 1990,[5] and the NSFNET was decommissioned in 1995, removing the last restrictions on the use of the Internet to carry commercial traffic.
In the 1980s, the work of British computer scientist Tim Berners-Lee on the World Wide Web theorised protocols linking hypertext documents into a working system, marking the beginning of the modern Internet.[6] Since the mid-1990s, the Internet has had a revolutionary impact on culture and commerce, including the rise of near-instant communication by electronic mail, instant messaging, voice over Internet Protocol (VoIP) telephone calls, two-way interactive video calls, and the World Wide Web with its discussion forums, blogs, social networking, and online shopping sites. The research and education community continues to develop and use advanced networks such as NSF's very high speed Backbone Network Service (vBNS), Internet2, and National LambdaRail. Increasing amounts of data are transmitted at higher and higher speeds over fiber optic networks operating at 1-Gbit/s, 10-Gbit/s, or more. The Internet's takeover of the global communication landscape was almost instant in historical terms: it only communicated 1% of the information flowing through two-way telecommunications networks in the year 1993, already 51% by 2000, and more than 97% of the telecommunicated information by 2007.[7] Today the Internet continues to grow, driven by ever greater amounts of online information, commerce, entertainment, and social networking.
Packet switching networks such as ARPANET, NPL network, CYCLADES, Merit Network, Tymnet, and Telenet, were developed in the late 1960s and early 1970s using a variety of communications protocols.[2] Donald Davies was the first to put theory into practice by designing a packet-switched network at the National Physics Laboratory in the UK, the first of its kind in the world and the cornerstone for UK research for almost two decades.[3][4] Following, ARPANET further led to the development of protocols for internetworking, in which multiple separate networks could be joined into a network of networks.
Access to the ARPANET was expanded in 1981 when the National Science Foundation (NSF) funded the Computer Science Network (CSNET). In 1982, the Internet protocol suite (TCP/IP) was introduced as the standard networking protocol on the ARPANET. In the early 1980s the NSF funded the establishment for national supercomputing centers at several universities, and provided interconnectivity in 1986 with the NSFNET project, which also created network access to the supercomputer sites in the United States from research and education organizations. Commercial Internet service providers (ISPs) began to emerge in the very late 1980s. The ARPANET was decommissioned in 1990. Limited private connections to parts of the Internet by officially commercial entities emerged in several American cities by late 1989 and 1990,[5] and the NSFNET was decommissioned in 1995, removing the last restrictions on the use of the Internet to carry commercial traffic.
In the 1980s, the work of British computer scientist Tim Berners-Lee on the World Wide Web theorised protocols linking hypertext documents into a working system, marking the beginning of the modern Internet.[6] Since the mid-1990s, the Internet has had a revolutionary impact on culture and commerce, including the rise of near-instant communication by electronic mail, instant messaging, voice over Internet Protocol (VoIP) telephone calls, two-way interactive video calls, and the World Wide Web with its discussion forums, blogs, social networking, and online shopping sites. The research and education community continues to develop and use advanced networks such as NSF's very high speed Backbone Network Service (vBNS), Internet2, and National LambdaRail. Increasing amounts of data are transmitted at higher and higher speeds over fiber optic networks operating at 1-Gbit/s, 10-Gbit/s, or more. The Internet's takeover of the global communication landscape was almost instant in historical terms: it only communicated 1% of the information flowing through two-way telecommunications networks in the year 1993, already 51% by 2000, and more than 97% of the telecommunicated information by 2007.[7] Today the Internet continues to grow, driven by ever greater amounts of online information, commerce, entertainment, and social networking.
When did internet happen ?
20 years ago today, the World Wide Web opened to the public
Today is a significant day in the history of the Internet. On 6 August 1991, exactly twenty years ago, the World Wide Web became publicly available. Its creator, the now internationally known Tim Berners-Lee, posted a short summary of the project on the alt.hypertext newsgroup and gave birth to a new technology which would fundamentally change the world as we knew it.The World Wide Web has its foundation in work that Berners-Lee did in the 1980s at CERN, the European Organization for Nuclear Research. He had been looking for a way for physicists to share information around the world without all using the same types of hardware and software. This culminated in his 1989 paper proposing ‘A large hypertext database with typed links’.
While the initial proposal failed to gain much momentum within CERN, it was later expanded into a more concrete document proposing a World Wide Web of documents, connected via hypertext links. World Wide Web was adopted as the project’s name following rejected possibilities such as ‘The Mine of Information’ and ‘The Information Mesh‘. The May 1990 proposal described the concept of the Web as thus:
HyperText is a way to link and access information of various kinds as a web of nodes in which the user can browse at will. Potentially, HyperText provides a single user-interface to many large classes of stored information such as reports, notes, data-bases, computer documentation and on-line systems help. We propose the implementation of a simple scheme to incorporate several different servers of machine-stored information already available at CERN, including an analysis of the requirements for information access needs by experiments.The document envisaged the Web as being used for a variety of purposes, such as “document registration, on-line help, project documentation, news schemes and so on.” However, British Berners-Lee and his collaborator Robert Cailliau, a Belgian engineer and computer scientist, had the foresight to avoid being too specific about its potential uses.
In 1990, working on a computer built by NeXT, the firm Steve Jobs launched after being pushed out of Apple in the mid-80s, Berners-Lee developed the first Web browser software called, fittingly, WorldWideWeb. By the end of that year he had a working prototype of the Web running on a server at CERN.
On 6 August 1991, the World Wide Web went live to the world. There was no fanfare in the global press. In fact, most people around the world didn’t even know what the Internet was. Even if they did, the revolution the Web ushered in was still but a twinkle in Tim Berners-Lee’s eye. Instead, the launch was marked by way of a short post from Berners-Lee on the alt.hypertext newsgroup, which is archived to this day on Google Groups.
The WWW project merges the techniques of information retrieval and hypertext to make an easy but powerful global information system.The post explained how to download the browser and suggested users begin by trying Berners-Lee’s first public Web page, at http://info.cern.ch/hypertext/WWW/TheProject.html.
The project started with the philosophy that much academic information should be freely available to anyone. It aims to allow information sharing within internationally dispersed teams, and the dissemination of information by support groups.
Although that page is no longer available, a later version from the following year is archived here. It acted as a beginner’s guide to this new technology.
The evolution of the Web
From here on, things began developing rapidly for the Web. The first image was uploaded in 1992, with Berners-Lee choosing a picture of French parodic rock group Les Horribles Cernettes.In 1993, it was announced by CERN that the World Wide Web was free for everyone to use and develop, with no fees payable – a key factor in the transformational impact it would soon have on the world.
While a number of browser applications were developed during the first two years of the Web, it was Mosaic which arguably had the most impact. It was launched in 1993 and by the end of that year was available for Unix, the Commodore Amiga, Windows and Mac OS. The first browser to be freely available and accessible to the public, it inspired the birth of the first commercial browser, Netscape Navigator, while Mosaic’s technology went on to form the basis of Microsoft’s Internet Explorer.
The growth of easy-to-use Web browsers coincided with the growth of the commercial ISP business, with companies like Compuserve bringing increasing numbers of people from outside the scientific community on to the Web – and that was the start of the Web we know today.
What was initially a network of static HTML documents has become a constantly changing and evolving information organism, powered by a wide range of technologies, from database systems like PHP and ASP that can display data dynamically, to streaming media and pages that can be updated in real-time. Plugins like Flash have expanded our expectations of what the Web can offer, while HTML itself has evolved to the point where its latest version can handle video natively.
The Web has become a part of our everyday lives – something we access at home, on the move, on our phones and on TV. It’s changed the way we communicate and has been a key factor in the way the Internet has transformed the global economy and societies around the world. Sir Tim Berners-Lee has earned his knighthood a thousand times over, and the decision of CERN to make the Web completely open has been perhaps its greatest gift to the world.
The future of the Web
So, where does the Web go from here? Where will it be in twenty more years? The Semantic Web will see metadata, designed to be read by machines rather than humans, become a more important part of the online experience. Tim Berners-Lee coined this term, describing it as “A web of data that can be processed directly and indirectly by machines,” – a ‘giant global graph’ of linked data which will allow apps to automatically create new meaning from all the information out there.
Subscribe to:
Posts (Atom)



