Showing posts with label Google Books. Show all posts
Showing posts with label Google Books. Show all posts

Thursday, November 14, 2013

Google Books: Scanning is Fair Use per Judge Denny Chin


Judge Denny Chin granted Google's motion for summary judgement today, dismissing the lawsuit by the Authors' Guild against their Google Books project.  Here is the ruling on Scribd.  Judge Chin found that Google's scanning and adding a search function to the print books was "highly transformative" lifting the Google Book Project out of the strictures of copyright into Fair Use.

Judge Chin referred to a number of amicus briefs filed by various library and scholarly groups to recognize the many benefits generated by the book and library projects in his opinion.  From preservation to data mining to increasing access, the Google projects are recognized as providing huge new benefits that were previously not reached by either print or e-book presence.

Judge Chin assumes that the plaintiff has established a prima facie case that Google has violated copyright through the scanning projects.  But because he finds that it falls under Fair Use doctrine


(§107 of  the Copyright Act), the case is dismissed. Because Google Books uses the words in the books for a different purpose - creating snippets for readers to sample, for instance, or to search with , or for datamining, Judge Chin finds Googles use "highly transformative."

Words in books are being used in a way they have not been used before. Google Books as created something new in the use of book text -- the frequency of words and trends in their usage provide substantive information.
     Google Books does not supersede or supplant booksbecause it is not a tool to be used to read books. Instead, it"adds value to the original" and allows for "the creation of new information, new aesthetics, new insights and understandings." Leval, Toward a Fair Use Standard, 103 Harv. L. Rev. at 1111. Hence, the use is transformative.
(from Chin,  Author's Guild, et al., v. Google, 05 Civ. 8136 (S.D. N.Y, Nov. 14, 2013) , at pp. 20 - 21). 
 Image is Judge Denny Chin, taken at the time of his confirmation as a federal judge.

Tuesday, March 29, 2011

Darnton on Digital Libraries

The New York Times Opinion pages of March 23, 2011 carried an essay from Robert Darnton, the Harvard Librarian, "A Digital Library Better Than Google's." In response to Judge Denny Chin's rejection of the Google Book Settlement, Darnton revisits the dream of Google Books - a vast digital library, which makes available freely to anyone with a computer and internet connection, the literature of the world.

Darnton gives a nice precis of the action up til now: The Authors' Guild, representing a mere 8,000 members, proposed to sue Google for infringement of copyright. Google, which could have defended its actions as fair use, elected instead to negotiate a settlement agreement. This was what eventually went before Judge Chin, in a much-amended version, which...

divided up the pie. Google would sell access to its digitized database, and it would share the profits with the plaintiffs, who would now become its partners. The company would take 37 percent; the authors would get 63 percent. That solution amounted to changing copyright by means of a private lawsuit, and it gave Google legal protection that would be denied to its competitors. This was what Judge Chin found most objectionable.
Other objections were that other authors (and illustrators as well), who were not represented by the Author's Guild, did not like the agreement's terms, and wished to negotiate their own terms. But the Agreement covered everybody, unless they specifically contacted Google to opt out. Instead of correcting past problems, the Agreement set into stone the development of digital books in the future, according to many critics. For example,
the question of orphan books — that is, copyrighted books whose rightsholders have not been identified. The settlement gives Google the exclusive right to digitize and sell access to those books without being subject to suits for infringement of copyright. According to Judge Chin, that provision would give Google “a de facto monopoly over unclaimed works,” raising serious antitrust concerns.
Judge Chin invited the parties to rewrite the Settlement Agreement. Darnton, at Harvard, is part of a group that is thinking of ways to create a noncommercial alternative to the Google commercial model that is currently proposed. How to fund such a thing is a puzzle, but Darnton suggests, for example, a coalition of foundations working with a coalition of research libraries. He hopes that Congress would pass a bill exempting orphan works from copyrights for noncommercial purposes in a truly public library.

As examples of public digitization efforts that have succeeded, Darnton offers the Knowledge Commons and the Internet Archive which between them have digitized several million books. He also cites efforts in several countries to digitize completely the national library. This includes France, the Netherlands, Australia, Finland and Norway. He hopes that Google itself might donate its digitized trove to such an effort. It cannot hurt to ask, and the idea is an exciting one. While the current management at Google is good-hearted and public-spirited, there is certainly no guarantee that this will always be the case.

Monday, March 28, 2011

Alternative to Google Books: Hathi Trust signs Summon as search engine -


The Chronicle of Higher Education reports that HathiTrust has made its large library of digital texts searchable by signing an agreement with Summon, a library-specific search engine produced by Serials Solutions. Indeed, if you click on the HathiTrust link, you will see that they now offer catalog and full text searches, as well as browsing on their site.

Looking at the Search tips, they offer phrase searches in both the catalog and full text searches. They offer wild card searches, for both single and multiple characters, ONLY in the catalog search. They allow Boolean searching, with AND, OR, and NOT searches in both catalog and full text searches. You can also search subsets of books in full text searching. So you can select to search a private collection of books, for instance, in the full text search. There is advanced searching in the catalog search only, so far. They say they are "exploring" advanced search for the full text option. You cannot search non-text material, such as graphs or illustrations. There are much more details on searching, especially surrounding foreign language works.

If you are affiliated with a library that is part of the HathiTrust, you can get the entire text of a book that is in the public domain by searching HathiTrust. If you are not affiliated with such a library, what you get is a single page with your terms. If, however, the book is in the Google Books collection, you can go from the HathiTrust search on into GoogleBooks, and get it there. This information is provided at the HathiTrust Search Tips site under Printing/Downloading. They also remind users that they can locate many of the items in physical form in the libraries near them or by inter-library loan through those libraries. HathiTrust would like feedback from users who find problem pages, or otherwise have comments on the service.

The GoogleBooks Settlement which recently was rejected by the judge in the case, would have allowed the HathiTrust to post snippets of text along with search results, according to the Chronicle article. But the HathiTrust search results now will only display the page or pages of the document on which the search term appears. You can select the result, and retrieve the page or pages, and the term will be highlighted in a certain format. You can see either a "page view" which is like a PDF of the page, or a text view, where they will add the highlighting. (from HathiTrust SearchTips, under Search Tips. Readers who use screen readers and OCR software will want the text view, of course.

The decoration is the Hathi logo, an elephant, from their own website.

Friday, December 17, 2010

New Study Using Google Books

The Boston Globe reports on a fascinating cooperative effort where GoogleBooks has created a new tool, the Google Books Ngram Viewer which allows a researcher to sift through the materials scanned into the Google Books project, and automatically calculate the frequency of a word and watch it change over time. You can then compare the changing frequency of different words across the decades or centuries.

It can be very interesting. The link above demonstrates at Google Labs with "Atlantis" and "El Dorado." But perhaps meatier questions (ha, ha) are raised by the examples in the Globe article. The online article reproduces what I saw in my print paper, and you can see it better online. They looked at changing frequencies of appearances of food terms: sausage, ice cream, hamburger, steak, pizza, pasta, and sushi. You can imagine that in English language publications, instances of pizza and sushi in particular, and pasta, a bit, have really only begun appearing since their popularization by returning World War II veterans. Increasing acceptance of ground beef, improved food inspection perhaps as well, and certainly the rise of fast food chains have increased the frequency of "hamburger."

They also tested the changing terms for types of influenzas. They looked at the frequency of the use of the word "God." This last in particular, allows the reporter to explain that this new tool is merely that. It is a new addition to the scholar's tool chest. It does not take the place of the scholar. The scholar eventually will have to sit down and read at least a portion of the literature. It makes a great difference if the appearance of "God" is in a prayer or an ejaculation or a discussion of theology. So the graphs carry a certain amount of meaning, but to really understand WHAT it means, the scholar still needs to visit the literature.

It's tempting to play with the data, but you really need to download a whole bunch of tiny files to begin. You have to be dedicated to this.

Tuesday, October 26, 2010

National Digital Library Proposal

The Chronicle of Higher Education has an exciting article on a proposal for a National Digital Library. There are a cluster of related articles, but you may not be able to read them without a Chronicle password. Fortunately, Robert Darnton of Harvard, who is spearheading the proposal, has put a lengthy essay about his arguments, in a more accessible place, at the New York Review of Books. In the essay, A Library Without Walls, he sets forth his utopian vision for a national digital library. I am not sure quite how this either builds upon or competes with either Hathi Trust Digital Library, which already exists, or the Google Book Project, and Google Scholar. Darnton's essay is lovely to read, full of stirring language, referring to Enlightenment ideals and actual projects in a number of other countries to digitize their national libraries.

At Harvard, we have conducted a preliminary survey of the projects underway in other nations. We have even located an incipient NDL in Mongolia. The Dutch are now digitizing every Dutch book, pamphlet, and newspaper produced from 1470 to the present. President Sarkozy of France announced last November that he would make €750 million available to digitize the nation’s cultural “patrimony.” And the Japanese Diet voted for a two-year, 12.6 billion yen crash program to digitize their entire national library. If the Netherlands, France, and Japan can do it, why can’t the United States?

I propose that we dismiss the notion that a National Digital Library of America is far-fetched, and that we concentrate on the general goal of providing the American people with the kind of library they deserve, the kind that meets the needs of the twenty-first century. We can equip the smallest junior college in Alabama and the remotest high school in North Dakota with the greatest library the world has ever known. We can open that library to the rest of the world, exercising a kind of “soft power” that will increase respect for the United States worldwide. By creating a National Digital Library, we can make our fellow citizens active members of an international Republic of Letters, and we can strengthen the bonds of citizenship at home.
Quite stirring and very exciting, but I am a bit confused about how his project relates to the afore-mentioned existing projects. If it manages to pull them together into something that will definitely become and remain open to the public, that would, indeed, be something worth cheering about.

Sunday, October 04, 2009

Google Books = Happy Publishers

The Google Books Settlement remains controversial, especially because of the "orphan books" provisions governing titles that are probably still under copyright but for which the copyright holders are unknown. Concerns about these provisions led United States District Court Judge Denny Chin to postpone an upcoming hearing so that the agreement could renegotiated. One group is happy about the Google Books Settlement, however, and that is publishers, according to this story in the October 3 edition of the Boston Globe. Google has made good on its promise to "bring to light books that otherwise would be nearly forgotten." This is happening thanks to the Google Books Partner Program through which publishers agree to Google's scanning copyrighted works and making them searchable through Google. "The majority of publishers in the program allow Google to preview 20 percent of each book. Google monitors Internet traffic patterns to prevent users from logging in multiple times and cobbling together the entire contents of a book." In addition, Google links to the publishers' websites and to online book retailers. At least one publisher now "uploads a digital file of every new title to Google as soon as the print version is published, [making] Google Books ...a fundamental part of the firm's book marketing plans." However lucrative this may be for the publishers, it is worth remembering that the publisher program solidifies Google's monopoly and makes it difficult for smaller organizations to compete.

Wednesday, April 22, 2009

U.N. opens it Digital Library

The Chronicle of Higher Education online has a brief but eye-catching note that the United Nations has opened a World Digital Library. Well, when you get there, it turns out to be more like a digital archive. It's pretty cool. It has lots of map images and photo images, a few scanned pages. It is really international. But it's not a library as I understand it, and it's not the answer to Google Book Project or any other digital library project. hmmm.

And speaking of Google Books Project, the same issue of the Chronicle has an article stating that Brewster Kahle's Internet Archive project has petitioned the judge in the Google Books Project Settlement case to be part of the settlement. I don't find the update at Justia.com, but it may just be too soon for it to have been entered there.