Friday, November 14, 2008

Muddiest Point week #10

People manipulate a web crawler by increasing links in their work. How do web crawlers recognize this and fight against it?

Readings Week #11

Two of this week's readings were about digital libraries and one was about institutional repositories.
Digital Libraries Challenges and Influential Work
Federal support funded the Digital Library Initiative, DLI-1 in 1994 and DLI-2 in 1998. The purpose of the DLI projects were to make large scale digital resources and collections accessible and interoperable. University led teams worked with commercial vendors and software companies to define and identitfy important document, data, and metadata standards and protocols for Web based searching. The DLI program contributed to the development of best practices. Significant technology was transferred from this program. A spin off from DLI program resulted in Google.
Dewey Meets Turing Librarians, Computer Scientists, and the Digital Libraries Initiative
This article discussed the association of the National Science Foundation with the Digital Libraries Initiative in 1994 . It also mentioned that Google emerged from funded work. This article dealt mainly with the relationships between Librarians and Computer Scientists as a result of their working together on Digital Libray projects. Publishers are also mentioned as interested parties to Digital Library development. According to this article the DLI project is seen as broadening opportunities for library science since the core functions of librarianship, organizing, collating and presenting information still need to be preformed.
Institutional Depositories: Essential Infrastructure for Scholarship in the Digital Age
This article concerned the need for universities to move beyond their historically passive roles of supporting publishers and into the development and maintenance of digital repositories. Recent technological trends and developments, such as the drop in online storage costs and standards development, have made this possible. The development of an institutional repository requires collaboration among the institution's community as well as a commitment to organization, access, distribution, and long term preservation of digital materials produced by its members. In this article an institutional repository is seen as "complement and a supplement rather than a substitute for traditional scholarly publication venues." This author of this article warns of possible problems with institutional repositories, among which are the problems associated with material loss due to technical failure. Little built in system redundancy is also mentioned.

Saturday, November 8, 2008

Thursday, November 6, 2008

Muddiest point Class #9

I am still alittle uncertain about XML. When it is said that XML does not use predefined tags, that you define your own, wouldn't that lead to alot of computer confusion? Or is it the use of the DTD or XML Schema that tells the computer what your tags mean?

Readings week #10

The article by David Hawking was all about search engines. Search engines index and answer billions of queries per day. They provide high quality answers and reject low value content. The major search engines named in this article are Google, Yahoo, and Microsoft. A large, geographically distributed infrastructure is neccessary in order to support a search engine. Search engines use crawling algorithms to compile lists of URLs. Crawlers use links in documents to find high quality websites. Documents without links are often not searched by crawlers. Crawlers can be prone to system problems and failures, and spammers in addition to a failure to consider unlinled documents. The second part of this article concerned the methods used by crawlers to index documents, usually by creating an inverted file that is stored, often compressed, in memory. Search engines often maintain lists of common queries in order to return search results quickly.
The next two articles concerned the deep, hidden, or invisible web as opposed to the surface web. Search engines usually do a poor job of accessing quality content from the deep web since article in the deep web are often html documents without links. The article about the Open Archives Initiative Protocol for Metadata Harvesting discusses the open access method for gaining federated access to eprint archives through metadata harvesting and aggregation. OAI's goal is to develop and promote interoperability standards and efficient dissemination of content. The OAIPMH protocal is based on common standards and was funded by grants. OAIPMH attempts to provide better communication between data providers who build repositories and collections with important content and services providers who are harvesters that build services for collections and contents. The final article, which was a bit dated concerned the use of a for cost product called "Brightlight" which claimed to be a search engine capable of searching the entire web, the surface and the deep web.