Showing posts with label digital books. Show all posts
Showing posts with label digital books. Show all posts

Friday, October 14, 2011

A New Mode of Full-text Case Retrieval - a work in progress

[This past academic year, John Palfrey, Professor of Law, Vice Dean, Library and Information Resources, Faculty Co-Director, Berkman Center For Internet and Society at Harvard Law School, was intrigued by an idea that I’ve been kicking around for several years and invited me to come to Harvard to work on it.

With the support of an incredibly talented staff in my home library, I felt comfortable taking a semester off. And so, I am visiting, as an Academic Fellow at Harvard Law School Library’s Innovation Lab for the 2011 fall semester.]

Designing a Solution to a Problem
The project that I’ve been asked to explore has to do with the inherent challenges of conducting case law research using full text online databases. The working title of the project is “Leading Case Service” and is designed to make online case law research more productive and more efficient. There are three factors that make online case law research very difficult.

First is the size of the database. It is estimated that there are approximately ten million published cases in the American legal system. The size of the database alone poses very serious difficulties for designers of search engines and indexing systems, both digital and analog.

The size of the corpus of case law in the American legal system isn’t merely the result of our society’s litigious nature. Prior to the mid-nineteenth century, the publication of cases was done very judiciously. Most cases were published in selective case reporters that only published leading cases. In fact, the most influential American case reporter in the nineteenth century was the predecessor to what we know today as American Law Reports, or ALR, and it only published cases of some particular significance, either because the opinion made a ruling on a novel aspect of the law or clarified an issue that had been dealt with by many courts with varying outcomes. In the late nineteenth century, the West Publishing Company entered the case law publishing market and effectively turned cases into a commodity. The method by which West published cases was virtually indiscriminate because it published any and all cases submitted to it by the courts. Its business model was founded on the premise that the more cases it could publish, the better; the more cases it could publish, the more volumes it could sell.

As the volume of cases it published grew, West developed an elaborate subject indexing system to help researchers. We know the indexing system as the Key Number System, and the index as the West Digest System. Today, the index alone numbers several thousand volumes! Coupled with Shepard’s citations, the digest and case-verification systems helped researchers identify both cases that were useful and those that were “still good law,” in the sense that they hadn’t been specifically overruled by another court. This system was extremely accurate, thorough, and objective but still left the researcher with a serious problem of having to wade through a substantial mass of material. The comprehensiveness of the West National Reporter System, its Digests and Shepard’s meant that the cases discovered on any one particular topic could number in the thousands.

The enormous volume of case law poses difficulties for researchers for another reason. Important research in the field of information science that explains that, due to the vagaries of language and other empirical laws of linguistics, full text database searching is by definition inefficient, even in databases filled with documents of a professional nature and highly specialized vocabulary, such as law. Studies have shown that the best a full text search engine is capable of retrieving amounts of only about twenty percent of the relevant documents on a topic. With 10 million published cases, even a 20% efficiency yields far more cases than any person can reasonably be expected to read.

Second, full text databases are objective search tools. This makes full text case law databases very difficult places for researchers to go to find answers about the law. For instance, let’s say you want to know what the law is on the rights of grandparents to intervene in custody proceedings in dissolution cases. A search for cases on this topic, if done with absolute precision, may yield dozens, if not hundreds, of cases, not what you really want or need. In this instance, a more useful approach would be to consult secondary sources, such as handbooks or treatises, that not only discuss the leading cases in the field but also summarize and analyze what these cases mean to the practitioner. Full text case law databases themselves are only part of what researchers need to complete their research.

Third, in order for online databases to be efficiently used, each document, as well as its sections, parts, words, letters, etc., should be indexed and tagged with what’s known in the computing world as meta-data. Indexing on this scale is massive and extremely complex but can make the development of search engines designed to work with these huge databases much more efficient. This is why Westlaw’s and Lexis’s search engines are so useful. Each company runs the full text of each case contained in their databases through extensive indexing and tagging. Part of this process eliminates repetitive words that have no legal meaning, such as articles, conjunctions, etc. Further, the content of the cases is divided into sections, such as majority and minority opinions, jurisdictions, etc., that dramatically helps narrow the search results. Indexing and tagging on this scale is a very costly venture, leaving only Lexis and Westlaw dominating the field. The process also is so complex that each company’s processes are highly guarded trade secrets. The Lexis and Westlaw case law databases are comprised entirely of public domain materials, but they still are extremely expensive to use. Each company cites the high cost of thorough indexing, tagging and sorting as a rationale to charge high prices for access.

A Solution to a Problem
The “Leading Case Service” may be a means of leveling the playing field for newcomers to the online database market, or for existing services that offer access to case law for free. The theory behind the project is that among the ten million cases in the American legal system, there is a relatively small percentage of cases that are considered more significant and interesting than the rest. If these cases can be identified and a search tool developed to exploit them, it may make searching case law more efficient by helping researchers focus on the most important cases first, before moving into the vast body of case law to find newer cases, cases with variant facts or those more specific to a specific jurisdiction.

The first step is to determine if there is such a thing as a group of “leading cases” and, if there is, to figure out how to find them and use them. The theory at present is that this group of cases can be found in the body of secondary materials. Initially, I thought that we could find leading cases in footnotes and body of treatises, presuming that treatise writers would discuss or cite to only the most important cases in their fields. To gather this group of leading cases, we could “mine” treatises and discover the cases cited in them. However, there are significant challenges to mining treatises for the cases the authors have cited, not the least of which is that treatises are published in many different formats and by enough publishers to make it difficult to use a single system to acquire the desired information. I’ve been convinced to put this step on hold; at least for the moment.

Our thinking at present is that we may have better luck focusing our efforts on cases cited in law review articles. There are two reasons that we think that law review articles may be better sources with which to discover this body of leading cases. First, we presume that the writers of law review articles as experts in their fields, are vigilant in identifying important cases in those fields and that overall, these scholars will discuss all the important cases in American law. I realize that this is a strong presumption, but over the last century, virtually all significant developments in law have been discussed and debated at length in law reviews and law journals. It follows, then, that the cases cited by the writers should be the ones most important or significant for one reason or another and can be identified as “leading cases,” those that researchers should read or at least be aware of when researching case law in that field.

A second advantage of using law review articles to identify leading cases is that the body of scholarship is continually expanding. If my theory is correct and we can discover this body of case law, we may be able to create an automated process that will continually add to the corpus of leading cases.

Questions
There are many questions to be answered. The most interesting question, and the one that I’ll be spending my time exploring initially is, exactly how many cases are cited by law review articles?

We know that there are around ten million cases published in total, but we don’t know what percentage of those cases found their way into the footnotes and text of law review articles. We are very close to obtaining the tools to answer this question conclusively. Hunches about the percentage of cases discussed in law reviews ranges from 5% to less than 1%, between 100,000 and 500,000 cases. If this is true, then full text case law database searching should be greatly improved by the mere fact that the researcher would be searching in a database of two or three hundred thousand cases instead of ten million!

Assuming that our initial tests reveal that there is, indeed, a body of leading cases that we can identify, many interesting possibilities emerge. The cases themselves may be ranked based upon the numbers of law reviews or journals that have cited them. (This is sort of a twist on Shepard’s service for law review articles. Instead of Shepardizing articles to find cases that cite to the articles, we’re looking for cases cited in the articles themselves.) Other information that may help rank the value of the cases includes the standing of the journal itself in which the article is published, or the reputation, publishing record or school of residence of the author.

Even if we are successful in identifying this corpus of leading cases, we have yet to determine how they should be used. The options are to create a separate database or to use meta-data to tag, or identify the cases so that search engines will be able identify the leading cases from among the rest of the millions of cases in the corpus of American law. Depending on the tags used, the researcher can use this information to sort search results in interesting and valuable ways. For example, a researcher desiring to know what is the law in an area novel to him, could begin with a full text case law database and immediately identify the most important cases in the field. After perusing these cases, links and metadata could then be used to immediately find articles, blogs and other pertinent online materials.

The goal of the project is to create a new way of using online digital legal materials. New technologies have allowed us to think of combining information in ways that were unheard of, even unthinkable, before today. “Leading Case Service” is essentially a ‘mash-up’ of online case law databases and online databases of law review articles. To this mash-up, colleagues have suggested that we may be able to add blogs, digital commons, wire-services, websites, legal periodical indexes and possibly treatises. The use of this information is not merely academic. It may also prove to be a way to power new search engines or discover new ways that various parts of the conceptual, scholarly world of the law influence each other.

Friday, January 14, 2011

Waiting for the Other Shoe to Drop

I'm baffled by publishers' arrogance these days. Two recent events made me whack my head with the palm of my hand….

Law Journal Seminars Press is now rolling out a "fantastic" new program for their books. Instead of merely paying for the looseleaf supplements for their books (for the most part reasonably priced, by the way), we can now either opt to receive them in print and online, or online only. Print and online, of course, costs more than the print supplements alone. Online only costs about the same.

I'm in an academic law library and online has absolutely no interest for me - or my patrons. Apparently, each title would have to have a separate login. So, if I did opt for either option, I'd need to keep track of the various passwords for each title. Good grief. I can't imagine a more inconvenient process.

We're canceling all Law Journal Seminars Press titles.

The other situation is even more annoying. And it always has been. We subscribe to the Economist ($138/year) and route it among the faculty and put the routed one in the faculty lounge. They've got a pretty nice online service with email alerts, etc., and we looked into getting an online subscription. After six months, they finally got back to us with a "fantastic" deal: $1500 per year for online access for our library. Are they mad? Do they really think that only one person reads our single subscription to the print version?

Come to think of it, I wonder why they don't charge volume rates for the print version any way?

Actually, I'm pretty sure that that's coming….

Wednesday, December 08, 2010

Reflections on the End of the World Wide Web and the Future of the Internet as an Information/Service Resource

[This post is in an essay written in preparation for the December 10, 2010, Episode 16 of "Law Librarian Conversations," a podcast about all things law library.... This week's podcast with guests Tom Boone, Reference Librarian, Loyola Law School; Jason Wilson, Vice President Jones McClure Publishing; Ed Walters, CEO, Fastcase.

If you are reading this before Friday, 12/10, you can join us by clicking on this link:
https://www2.gotomeeting.com/register/537047386

Title: LawLibCon 16 - Future of Interface Design (12/10/2010)
Date: Friday, December 10, 2010
Time: 2:00 PM - 3:00 PM CST
After registering you will receive a confirmation email containing information about joining the Webinar. Follow the conversation in the chat room during the live broadcast athttp://lawlibcon.classcaster.com/chat.

Subscribe to LawLibCon on iTunes here: http://u.cali.org/2jwf . Enough shameless self-promotion.... RL]

==================================================

I've been fascinated recently by a new trend in consumer use of the internet.

The internet itself has remained pretty stable as a continuing backbone for the electronic exchange of information. It has been surprisingly robust and scalable. Embarrassingly, I was one of those people who, in the early 1990's was predicting that the internet would soon break from the volume of usage as it spread from academic to commercial users. At the time, it seemed that as usage - especially, graphical intense usage - would fill the capacity of routers and cables of our national infrastructure. Surprise, surprise. Hardware manufacturers and ISPs have somehow figured out to meet the demand. (And have they ever….)

But while the internet backbone has scaled up and provided one and all with (potential) capacity for the mammoth amounts of bandwidth. A well wired home could (does?) have a wireless access point that supports at least ten simultaneous devices on the same router at speed and capacity to allow all ten users to stream music while surfing the web with several tabs open.

So the internet itself supports surprisingly intense usage. But the people who make money on the sale of this usage are the ISPs. What about the information providers who provide the information or services that consumers use? In order for content and service providers to make money for their information or services, they need two things: Unique, high quality information or services, and, eyeballs. Service or information providers either make money on the information that they sell to users, or they give their information or services away for free and sell advertising to others. This much is obvious. It's also the great challenge of being in business on the internet.

Either way, the vendor has a vested interest in holding the users' attention as long as possible. One way that they are doing that is by creating new "platforms" for access to, or usage of information or services available on the internet. If you think about it, this basic concept underlies nearly all recent developments in cyber-business. The creation of the iOS and its use of apps to access the internet was one of the very first examples of a way to lure users away from the wide-open world wide web and into a world where use of the internet was now carried out completely from within an application. This provided us with excellent, robust access to and usage of the information or services, while at the same time, keeping us in that very location for focused, discreet periods of time.

But as Mobile Safari began to become a major source of internet activity, other service and information providers began to see the potential for customized, focused internet experience. Google keeps developing more and more products and services and makes them available for free to all-comers - and yet makes billions. And they make that money without ever sending you a bill! The strategy is simple, get users to click into the world of Google for search, mail, documents, RSS feeds, phone, etc.; And keep them there! Google is striving to extend their reach even further by developing its mobile platform, Android and its forthcoming operating system, ChromeOS. Once in the Google world, a user will be able to stay in that world. FaceBook, too, making plays to become the one-stop source of all your internet life and activity.

Developments such as these may ultimately serve to make the intent a series of walled gardens where users can't easily move from one application or platform at will, at least not easily. Examples of how fine these walled gardens have become can be seen in two recent announcements of publishing ventures which have begun entirely as applications on the iPad. For example, Rupert Murdoch's announcement of the creation of an iPad based newspaper called "The Daily," and Richard Branson's new magazine called, "Project." The Daily isn't yet released (as of this writing) but Project is. And it is stunning in nearly every respect. It's beautiful, packed with features and utility. But it is limited in one important respect.

Even if I wanted to share with you the wonderful cover story in Project, I couldn't. First off, there is no URL. Second, even if I could, somehow send you a link to the article, because Project was developed on an app built especially for the iPad, you need one in order to view it. And there's no end in sight. Such applications are bound to appear in all major platforms.

The irony is that well-executed applications provide outstanding experience for the user and many people prefer the experience of browsing Twitter, Facebook, RSS feeds and databases through available apps over accessing the same information with a browser.

What's a developer to do? The hope, from a user's point of view is that developers will focus their efforts on building good services and databases and make them available on every available platform. It is also important for a new system of link locators be developed so that links from within a ChromeOS application will be able to find the same article or information in an iOS application or from within a browser.

Such challenges are so subtle and nuanced as to be nearly invisible today. Tomorrow they may well be extreme obstacles for cross-platform use and may make today's successful platform the preferred one for distribution of tomorrow's information and services….

Monday, November 16, 2009

New Concept in Database Search Engines

I have been thinking about this concept for about a year, and I can't get it out of my head. It's time to share it. I hope that Google, CCH or BNA reads it, exploits it and sends me a hefty check....

Why online haven't legal database providers figured out that online databases are a new breed of legal research tool and developed something completely different? To date, all online databases are not much more than online versions of their old-fashioned print tools. There are differences, of course: Online searching allows users to find particular cases and documents quickly, sort rapidly and print more cleanly, but in reality, online tools do no more than allow users to skate around through masses of undifferentiated primary law, using cite-verification tools to sift through the mass of material fairly quickly. But without much help or guidance.

I propose development of a new kind of online search engine. First, let's establish a few assumptions. First, let's presume that cases cited by treatises, law review, blog writers and commentators are cases that are most important than cases that are not cited by these writers. Second, let's presume that cases cited more frequently are more important than less cited cases. Third, it is possible to make assumptions about the relative value of a case based upon the kinds of works a case is cited in, as well as the kind of treatment that a case receives in that work.

Based upon these three assumptions, I think that it is possible to develop a database(s) that is comprised of only cited cases. What's more, meta-data can be created that will note where it was cited, and the level of treatment.

There are at least six great sources from which you can build such databases. West has, perhaps the greatest library from which to build such a database. It's collection of secondary materials is tremendous. Lexis is also well-positioned to accomplish something like this with its Matthew Bender titles. But, perhaps the two companies best equipped to build such a high performance database are CCH and BNA. These companies own some of the very best specialized law treatises. It's nice for these companies to put their newsletters and looseleafs in electronic format, but, to paraphrase early library automation consultants, "an electronic version of a good looseleaf only creates a good electronic looseleaf." In other words, it doesn't make a good thing better; it only makes it electronic. In order to make a good thing great, it must be different. (That should be obvious, but somehow it's not….)

But what if you're not West, Lexis, BNA or CCH? Are you out of luck? I don't think so. There are two resources left. First, Hein Online is now comprised of an unprecedented collection of law reviews. This is a vast gold mine of notable cases. Hein itself could develop a search engine that sifts through the very best cases based on citation frequency among law review writers.

A newly emerging resource that may accomplish roughly the same thing, are digital commons and blogs. Looking forward, a crawler could be designed that will crawl through digital commons, legal blogs and law review websites looking for cited cases. Here, the presumption is that cases that are discussed by more writers are more significant.

Finally, it is possible that such as database could be made simply from cases cited by other cases. It can be presumed that cases that are cited by other cases most frequently are those cases that are more significant legal precedents.

Tuesday, March 24, 2009

U of Michigan Biting the Dust (?), Poised to Turn into Blog....?

The Great Lakes IT Report reports that the U of Michigan Press is following the trends and will revamp their publishing operation and expand into "3D animation and video", as well as publish it's scholarship in digital format so it can provide hot links and graphics. The announcement says that it "will be "restructured" to focus primarily on digital monographs, not the printed version."

It seems to me that ceasing publication is quitting publishing, and selling scholarship as pdf's and web pages won't enhance it's prestige, but will dilute it. It also seems odd to brush off concerns about customers who want to "hold something" can simply print them off on their own. Most scholars that I know would rather publish with a publisher who can actually capture the scholarship and sell it as an item. Blogs and hot links still don't have the cache of a printed book.

That's not to say that blogs don't have their place, or that bloggers aren't thinkers. It's just that their material is inherently different. It's a new format that's gaining respect and notoriety all it's own. Witness, Obama has even called on Politico correspondents in his first two press conferences. If that act alone hasn't given bloggers credibility, then nothing has. But does this mean that blogs are equivalent to University Presses?

UM's announcement, I think, is short-sighted. If anything, they should go slow, and start a blog, perhaps, and use it to promote it's catalog.

Fortunately, the announcement doesn't say that it is going to completely cease it's print publishing, but, spokes-people quoted in the article seem to indicate that it is going in that direction. I predict that ten years from now, it will largely be the same as it is now. But with the addition of a digital division; it will have higher overhead and will probably be selling more books.

Finally, I'd like to know how many libraries, or customers, for that matter, actually buy digital books. When I see adverts for e-books, I usually pass them up. What's a library to do with e-books, any way? To me, it seems that delivery of e-books is too personal for libraries to be involved with. I can provide links to the material, or direct patrons to useful titles, but I can't be responsible for how they actually obtain use, or fuss with setting up their equipment or software to guarantee their ability to use it.

If a patron has a Kindle (or the new e-book reader/web-book from Apple that's coming in the summer) how can a library lend it out? There's a missing link in this business model.

I wish the U of Michigan Press well, and hope that they are able to complete their misguided experiment before too many others go down the same road.

OK, a final thought: If a publisher publishes a title in a format that no one can read, have they still published a title? The thing that's neat and tidy about publishing a book is that the end user needs only two things to read it: light and the ability to read. (OK, knit-pickers, they do need access, but that's theoretical....) But look what's required to read an e-book: power, equipment of a particular variety, connection to the internet, software and the ability to make it all work together - plus the ability to read.