Monday, November 24, 2008

Wednesday, November 19, 2008

Week 12 Reading Notes

Social Bookmarking: One thing that really interested me was the possibilities for social bookmarking in libraries or archives. The article discusses the university of Pennsylvania's use of tagging to incorporate websites into the catalog. I wonder about the use of social bookmarking for the items within the collection. Obviously there would be some of the pitfalls associated with untrained cataloging of collection, but I wonder whether it might produce interesting alternative results to LoC subject headings.

Wikipedia: Wale's discussion of the vandalism systems was especially interesting to me. This week I was reading the Pieter Bruegel the Elder page on Wikipedia. It was clearly vandalized, as one of Bruegel's paintings was listed as being held at the "What the Frick? Collection in New Pork City (ha ha ha.). Sure enough, the page was fixed within days.

Monday, November 17, 2008

Friday, November 14, 2008

Week 10 Muddiest point

Is there any search engine for scholarly materials that creates a ranking system based on citation instead of links?

Also, I know you've gone over the diagram of the internet with the tentacles a number of times, but I still don't get the "IN" and "OUT" parts of the diagram. what do those refer to?

Saturday, November 8, 2008

Assignment 6: My website

Here is my website. It is best viewed in Firefox.

Week 9 Muddiest Point

When an xml document links to another website for its schema, which I understand is an XML document, can that document be automatically interpreted by the parser? I assume it can, but I wanted to make sure.

Friday, November 7, 2008

Week 10 Reading notes

The readings covered the structure of web search services. Modern search engines use a sophisticated combination of crawling algorithms, parallelism, and filtering to crawl web pages. Indexing algorithms then collect significant terms on the web page, as well as additional information such as the frequency the term appears and its position. A query processing algorithm then returns documents from the index that contain all the search terms where possible. Search engines use strategies to make the results more accurate and to speed queries up.

Most search engines do not crawl the "deep web," that is, the temporary web pages that are produced as a result of searches on commercial websites, etc. There is many times more information in the deep web than on the surface web, but most browsers do not search it because it would take to much time to create searches of these sites then index them. Some search engines, such as BrightPlanet attempt to search the deep web with the belief that there are more quality results to be found there than on the surface web.