Wednesday, November 12, 2008

Week 11 readings/muddiest point

Deep Web: Michael Bergman

This is an interesting article on 'deep web."  I'm still not sure how it gets such a fancy name.  I'm still not sure exactly what it is.  For example, they listed ebay as a deep-web page.  I can get to that via search engine very easily.  In the web, it seems that deep web will not look any different than surface web.  In fact, all web was deep web until maybe the advent of search engines.
Maybe it is differentiated because search engines can reach only about 16% of the web.  It is a shame, because there is 400-550 more times the public information on the deep web.  To conclude, I surmise that the deep web is not inaccessible, just not randomly searched.  People who use the deep web know where they're going and so don't google it.  Though I could be wrong.

Web Search Engines: part 1 /David Hawking

The premise here is that web search engines provide high quality information quickly. They cannot and should not attempt to index the web in its entirety. Indexing begins with a "seed" Url.   The search engine can then search inside the seed (ex: topics within wikipedia).
Different search computers search different areas, and forward search requests to the machines that are assigned it.  They also make sure that web browsers are not overwhelmed with requests by adding a politeness delay to make sure each request goes in 1 at a time. 
Robots do not-recrawl over all the web.


Web search Engines part 2/David Hawking

The vocabulary of the web includes many languages, including new words specific to internet culture, and also includes misspellings and grammatical errors.
Most queries are 2 words long. All query searches include all query words.
Search engines have strategies to speed up searches: they can skip, make lists of decreasing value, assign number scores according to their decreasing value.  They can cache: pre-store anticipate search answers.

OAI Protocol for Metadata harvesting: 
I think this about metadata and steps taken to be able to comprehensively search it?  Seriously, I'm lost.

Muddiest Point: Is there something I'm missing about the deep web?  Why is it not linked to search engines?
Could someone explain OAI to me in simple language?

Friday, November 7, 2008

Week 10 readings

Digital Libraries: challenges and Influential Work by William Mischo

This is an interesting article explaining the beginnings of research into digital libraries, i.e. having some sort of methodology programmed into how one searches  the web (putting the books back on the bookshelf if you will).
From government grants for a few schools with different needs, other institutions have adopted what was done in these studies, and google was born.  
Now, the need is to build something akin to Google scholar in digital library form.

Dewey Meets Turing: Libraries, computer Scientists, and the DLI

Hmmm.  So, computer scientists and librarians could have gotten along rather well together, systematically organizing a digital library, but the pesky web was born and grew up so fast and stressed the relationship.  Enter publishers and computer scientists can no longer make their discoveries open and free, but can only tease their colleagues with their new programs.
Meanwhile, librarians needs are not met and librarians think that computer scientists don't understand them and have forgotten about them, and computer scientists just wish librarians were more like computer scientists!

Institutional Repositories

I agree with the author that it does not make sense to have authors, especially of academic institutions to be in charge of posting their work online and archiving it.  This is not their job, and if forced to do it, will often be sub-par to what a centralized, specialized department could do with these works.
I also agree that they can be useful as collectors of ephemera.
But: they cannot claim ownership over a faculty or students work.
And cumbersome gate-keeping policies will be, in the words of the author, counterproductive. The author believes in keeping it simple.  Too much policy would undermine its effectiveness.
Also, as important as they are, let us not be too hasty in their implementation, but let us be thoughtful and end up with a useable, sane, repository.

Tuesday, October 28, 2008

Week 9 (: Blogs where I have posted

http://pittmlis.blogspot.com/2008/10/reading-assignment-9.html

http://iit2600.blogspot.com/2008/10/week-9-muddiest-point.html

http://kirstenbell.blogspot.com/2008/10/week-9-reading-comments.html

Week 9 readings & muddy point



XML: Martin Bryan
Ok. XML is a formal language which can be used to 'pass information about
component parts of a
document to another computer system.'  It is object-oriented, hierarchical, and
 is a clearly defined format.  
Luckily, the other readings for this week cleared this all up for me.

XML Standards: Uchi Ogbuji
I sense that XML is a deeper language than HTML.  I looked at the ZVON
XML
tutorial (
http://www.zvon.org/xxl/XMLTutorial/General/book.html
posted on this page, and began to see what XML is about.  I think I can
do it.

Tutorial: Andre Bergholz: I was unable to locate this at the URL provided,
 and in the magazine
 it was published.  Through the U Pitt Libraries, I found the IEEE Internet
 Computing Journal for
 July 2000, but the pages 74-79 ( this article) were missing.  As proof, I've
included the two
sandwiching articles.  The ZVON Tutorial helped a lot however so I feel ok
about missing this one.

From July 2000 IEEE Internet Computing: U Pitt E journals:


Building an IP network quality-of-service testbed

McWherter, D.T.; Sevy, J.; Regli, W.C.
Page(s): 65-73
Digital Object Identifier 10.1109/4236.865089
Abstract  | Full Text: PDF (124 KB) 
Rights and Permissions
Dreams of a unified information space W3C activities at WWW9
Woods, S.
Page(s): 81-83
Digital Object Identifier 10.1109/MIC.2000.865090
Abstract  | Full Text: PDF (84 KB) 
Rights and Permissions

________________________________________________
XML Schema Tutorial:
Deeper than XML, better than DTD, it enhances both using coding
 language already in existence, and boasts standardized formations
 to avoid confusion in dates or quantities.  Largely way over my head.

Muddiest Point: Not so much for class, but as in how these readings
could apply to me.  Could I actually use these in a web page?
 It's worth a try!








Tuesday, October 21, 2008

Koha Library

http://pitt4.kohawc.liblime.com/cgi-bin/koha/bookshelves/shelves.pl?viewshelf=59