Wednesday, November 17, 2010

Week 11 Reading Notes

Web Search Engines
Thanks to Sarah Denzer, I found the article. Apparently, using the full citation of IEEE Computer threw the entire system of citation linker off.

Reading through these articles gives me a different kind of appreciation for the systems used for crawlers. I was familiar with the concept from a book for LIS2000 (Laszlo's Linked), but I didn't know about the degree of equipment involved. Knowing that it would take some of our better net connections 10 days to do a crawl caused me to do a double take, and seeing the numbers made my head swim.

This is a MUCH different idea than what I've seen from the simple "find it" programs I wrote as an undergrad, as this will find the information, index it, and in a way, learn from it. As I said, new appreciation for what is done and how it is done.

Did anyone else have a similar feeling?

Current Development and Future Trends for the OAI Protocol for Metadata Harvesting
Once again, we’re referencing back to the Dublin Core and other metadata standards to sift through and organize data. While the article does give a few inspiring notes as to how this could come to be with examples of organizations/consortia trying this, I still have to wonder the same as I did before: can this really be done?

I mean, honestly: even with a “standard” set of metadata, how viable will this be? Will we actually have a comprehensive set of usable search terms to actively search the “deep web” (including databases), or are we just going to add more clutter to the already vast amount of data hidden on the Internet?

The Deep Web
Some parts of this article made me think of Laszlo’s book Linked, especially with the early section regarding how sites would often be connected or a crawler would find the data.

Thankfully, the article covered more than that, by offering statistics (which gives me a new found respect for the amount of digital data on the internet, as it is measured in thousands of terabytes) and a comparison of “surface” and “deep” web searches, which also explains why the general “sweep” done by the standard search engines just doesn’t cut it for finding what you really need.

There isn’t much to say about the article beyond definitions and numbers (and the feeling that it was a plug for certain technologies), but it does make me interested to learn what is really out there hidden away in the depths of the web.

Friday, November 12, 2010

Week 10 Muddiest Point

I don't have a muddiest point this week.

Either I'm understanding enough to get by, or I'm missing something. I'm sure I'll find out soon enough!

Monday, November 8, 2010

Week 10 Reading Assignments

Alright folks, now that the dust is settling after the fiasco of the past few weeks, I should be able to get back on track and get back to my usual writing.
At least, I’m hoping so.

Digital Libraries: challenges and influential work.

The article is basically a history of where we’ve come about for digital library resources, including the DLI projects and which universities/institutions took part to get us where we are today.

Personally, I liked the reference to aggregators and how they are impacting how we do our jobs. Even more interesting is the note on how Google is still the team to beat, as Google Scholar is a product commercial companies are trying to replicate. I still don’t think Google is the end-all-be-all that people make it out to be; there are some great ideas there (such as standard combined with full-text metadata), but with my experiences with Google as a whole, I’m not entirely comfortable with the idea of putting this product on the pedestal.

One thing I will agree with: aggregators are a great idea, especially when coupled with the concept of full-text metadata. Maybe I’m living in a dream world and have been blown away by some of the commercial products I’ve had presented at work, but hey, a guy can hope.

Dewey Meets Turing: Librarians, Computer Scientists, and the Digital Libraries Initiative

To begin, the title gave me a chuckle, and the introduction gave me some hope as to what was to be discussed. Comparing the expectations of Librarians and Computer Scientists, and tossing in Publishers to complete the trifecta? Move over soap operas, library science has you beat!

But now to be serious for a few moments: when you consider library science and computer science merging together to work on something, you’d only assume it could be a match made in heaven. Libraries need ways to sort and sift information, computer scientists need better and more efficient ways to do their own research. Sounds like a good idea in general.

The article explains the complications these two groups faced, especially with the Web connecting machines (and therefore, data) in unexpected ways and improvements to technology and the way computers “think.”

This article does bring up some other food for thought that has come about due to the changes in technology and its integration into library services, the biggest one referencing the acceptance of online-only publications. With all of the debate regarding copyright law and open access publishing, I do have to wonder if this medium will come to be the primary method of doing things, and if so, how long until print materials and other “traditional” library resources and services are phased out for digital materials and “capable” computers?

Thankfully, we get to see some glimmer of hope at the end of the article, showing that we librarians aren’t entirely phased out just yet. . .

Institutional Repositories: Essential Infrastructure for Scholarship in the Digital Age
http://www.arl.org/resources/pubs/br/br226/br226ir.shtml

And now we have an article on Institutional Repositories. With this article, I was walking in blind, as I haven't exactly heard the term utilized in the workplace before. The author defines an institutional repository as "a set of services that a university offers to the members of its community for the management and dissemination of digital materials created by the institution and its community members." To me, this basically states it is an archive of things created by the members of the institution (in this case, a college), and allows access to this information by members of the set community. Have a missed something in this?

It would seem this is a step toward institution-sponsored open access, in that the creator (in this case, a faulty member) can update a previously written work or continue the work in that same topic without the time consuming steps of scholarly publications. Correct me if I'm wrong, but don't we need more of this in the academic community, where faculty members who want to write about a topic can do so without the hassle?

Maybe I'm living in a dream world, but I would like to see that come to be.

Sunday, November 7, 2010

Week 9 Muddiest Point

A bit late, I know. I do not have a muddiest point this week.

I also do not have any comments this week. I won't go into the details of what has occurred in my life to cause me to trail so far behind. Hopefully I can get back on track without something else happening sooner rather than later.