Tuesday, May 05, 2009

Check out the new look at GPO

The U.S. Government Printing Office has launched its new website with a slick new look. I have not explored it thoroughly yet, but it looks very nice. At a quick glance, though, it appears that it's no longer the one stop for government documents online that it once was. I guess it hasn't been that for a while. There are still a number of useful documents there, but it's not the place to go for Congressional materials, for instance. See http://www.gpo.gov.

New Search Engines: Wolfram Alpha & Google Public Data

The Boston Globe today has a story by Hiawatha Bray, "A hungry little number cruncher," featuring mainly Wolfram Alpha.

Computers can't think, but they can count.

That's the idea behind Wolfram Alpha, a new search service that could be as much of a game-changer as Wikipedia or Google. Alpha, created by renowned mathematician, author, and entrepreneur Stephen Wolfram, uses fast computers and vast statistical databases to answer questions just as a human would - a human with advanced degrees in math.

"The goal is to see just how much of the world's knowledge can be made computable," Wolfram said.

Quite a lot, it seems. At lightning speed, Alpha serves up detailed, number-based answers about everything from algebra problems to poker odds.

"It's a compelling new service," said search engine industry analyst Danny Sullivan. "It's very much one to watch."

A few Internet mavens have even suggested Alpha could threaten Google's dominance in Internet search.

Wolfram scoffs at such talk, saying that Alpha is a supplement to traditional search, not a substitute. He's interested in partnering with other search companies, though he wouldn't name names.

Google is exploring similar ideas. On the same day last week that Wolfram demonstrated Alpha at Harvard Law School, Google launched its new Public Data service, which automatically responds to questions about US unemployment or population with graphs generated from the government's statistical databases.

Google Public Data isn't nearly as versatile or powerful as Wolfram Alpha. And unlike Alpha, it didn't grow out of a theoretical quest to make machines think. Google project manager Ola Rosling had a much simpler goal: "Democratizing the access and usefulness of public data, by making it easy to find and use."

Rosling said Google's launch wasn't an attempt to step on Wolfram Alpha's debut. But he said Google plans to make Public Data more like Alpha, enabling it to crunch numbers into practical information for everyday use. "We're definitely on the same side of the street," Rosling said. "But the street is pretty long and pretty broad."

For now, Google Public Search is rudimentary. Users who type queries like "unemployment in Massachusetts" see a graph at the top of the page. Clicking on it provides a more detailed view of the data, as well as links to let you compare Massachusetts to other states and break out the numbers by county.

But Wolfram Alpha, which will be available for public use later this month at wolframalpha.com, offers detailed, math-based responses to a huge variety of questions. To Wolfram, anything that can be expressed as a set of numbers is a computable problem. So he and his colleagues have filed away vast amounts of numerical data on a host of subjects. Alpha uses these numbers to generate practical information for scientists, business people, or anybody else.

Type "International Space Station" into Google, and you will be directed to millions of pages of information about the orbiting science lab. Type the same query into Wolfram Alpha, and you will get a map of the station's precise location over Earth at the moment you typed your query. It will also tell you the next time the station will pass over Boston. It's all calculated instantly from statistics provided by the National Aeronautics and Space Administration.

Apart from Alpha's knack with numbers, the service is good at analyzing a user's questions and surmising what he is looking for. (snip)

But Alpha is a long way from perfect. Vast amounts of useful data are nowhere to be found. (snip)

Wolfram admits Alpha is still relatively ignorant. His team of developers is constantly adding statistical databases to expand the range of the service.

He's also still trying to figure out how to make money from Alpha.

Wolfram said a basic version of the service will always be free, perhaps with advertising support or a corporate sponsor.

In addition, he's planning a paid subscription version that will offer advanced features, such as the ability to download statistical data directly into a desktop spreadsheet program.

Wolfram has no interest in selling the business to Google or any other company that might use Alpha technology to enhance its own search offerings.

"My goal is to make this an authoritative source for knowledge on the Web," he said. "I fully expect to be working on developing this for decades to come."

The Wolfram Alpha search engine should definitely be on librarians's bookmarks as a tool when it becomes available to the public. And here is a nice video and blog post explaining how to use the Google Public Data feature in the Google search engine.
We just launched a new search feature that makes it easy to find and compare public data. So for example, when comparing Santa Clara county data to the national unemployment rate, it becomes clear not only that Santa Clara's peak during 2002-2003 was really dramatic, but also that the recent increase is a bit more drastic than the national rate. (there are images on the page -- go and see it; I'm extracting some text to make it clearer that this is a widget on a regular Google search)

If you go to Google.com and type in [unemployment rate] or [population] followed by a U.S. state or county, you will see the most recent estimates.

Once you click the link, you'll go to an interactive chart that lets you add and remove data for different geographical areas.
In both search engines, they plan to keep adding more. They are searching existing data sets that have been gathered over decades (or longer) by lots of organizations, government agencies and more. These are tools to make the mountains of data more easily accessible to searchers on the web and more easily digestible through visual aids such as graphs and charts. Pretty wonderful aids!

Monday, May 04, 2009

A Different Take on Supreme Court Appointments


Rumors are swirling about possible picks for the Supreme Court seat about to be vacated by Justice David Souter. Conventional wisdom says that the pick will probably be a woman, given the Court's blatant lack of gender diversity. The pick will probably also have served as a federal appellate judge since that has been the career path of most recent appointees. University of Maryland School of Law Professor Sherrilynn Ifill, however, advocates a different approach. In her article, Professor Ifill argues that the current justices are "representative of a very narrow slice of the [legal] profession." Missing from the ranks of possible appointees are "[s]tate court judges, full-time law professors, former criminal defense attorneys, even civil practice trial lawyers ..." She believes that the unique perspective of those who understand what it takes to prepare and put on a case are lacking among those who sit on the Court today. While racial and gender diversity is important, so is professional diversity. Professor Ifill makes an interesting argument that I have not heard anyone else make. It is useful to remember that some of our greatest justices had not been federal appellate court judges before their appointment to the Supreme Court. For example, Justice William Brennan was serving on the New Jersey Supreme Court at the time of his appointment, and Justice Felix Frankfurter was teaching at Harvard Law School.

Google Book Settlement - A terrific analysis

Tip of the OOTJ hat to Kathleen Vanden Heuvel for pointing us to a wonderful and powerful analysis of what is wrong with the Google Book Settlement. Professor Pamela Samuelson (the Richard M. Sherman Distinguished Professor of Law and Information at the University of California, Berkeley, as well as a Director of the Berkeley Center for Law & Technology and an advisor to the Samuelson High Technology Law & Public Policy Clinic at Boalt Hall), writes as a guest blogger at Radar O'Reilly.com today, "Legally Speaking: The Dead Souls of the Google Book Settlement."

I love the literary twist she gives her title, playing off the similarity of the name, Google, to the Russian author Gogol. And she ties in the plot line of Gogol's novel, Dead Souls, wherein a schemer tries to leverage the value of dead serf's titles which he bought up since the last census to make himself a wealthy man. Serfs were called souls, hence the title. The character's plot comes apart as word gets out that all his stock of souls are dead... But Prof. Samuelson sees a parallel between Gogol's story and Google's scheme.

A huge component of the Google Book Settlement is the orphan books component. Where the author is dead or cannot be located, and no current owner of the copyrights can be identified, the book in question is an "orphan work." The Settlement creates a Book Rights Registry to collect money from ads viewed for each use of any work, and to distribute the money according to the Settlement agreement. Google keeps some, the BRR keeps some for overhead, and then a bit goes to the copyright owner if they can be identified. In the case of orphan works, however, things get, well, strange. Furthermore, Prof. Samuelson raises the question, how was the Settlement negotiated? There were only a handful of authors or their representatives involved. And those tended not to be the academic sorts who were largely on the shelves of the research libraries scanned for the Google Book Project. These were mystery writers, for instance, and publishers of textbooks. Hmm.

Even more upsetting is the fact that the folks running the Book Rights Registry are Copyright extremists. The head of the Authors Guild, as Prof. Samuelson points out, was the one who led the charge to prevent Amazon from leaving the Kindle set so it would read books aloud. This would have let visually impaired users use the Kindle to access books in a simpler way. But the Authors Guild claimed that it impaired authors rights in their secondary market of books on tape (now on CDs I guess, but I think it may be a Trademarked name). You can see the mindset at work.

I recommend you read this short and accessible article in full. It is terrific and thought provoking. Here is a snippet to lure you in:

If asked, the authors of orphan books in major research libraries might well prefer for their books to be available under Creative Commons licenses or put in the public domain so that fellow researchers could have greater access to them. The BRR will have an institutional bias against encouraging this or considering what terms of access most authors of books in the corpus would want.

In reviewing the settlement, the judge who is supposed to consider whether the settlement is “fair” to the classes on whose behalf the lawsuits were brought. He may assume the settlement is fair because money will flow to authors and publishers. But importantly absent from the courtroom will be the orphan book authors who might have qualms about the Authors Guild and AAP as their representatives.

(snip)

In the short run, the Google Book Search settlement will unquestionably bring about greater access to books collected by major research libraries over the years. But it is very worrisome that this agreement, which was negotiated in secret by Google and a few lawyers working for the Authors Guild and AAP (who will, by the way, get up to $45.5 million in fees for their work on the settlement—more than all of the authors combined!), will create two complementary monopolies with exclusive rights over a research corpus of this magnitude. Monopolies are prone to engage in many abuses.

The Book Search agreement is not really a settlement of a dispute over whether scanning books to index them is fair use. It is a major restructuring of the book industry’s future without meaningful government oversight. The market for digitized orphan books could be competitive, but will not be if this settlement is approved as is.
I only hope Judge Denny Chin reads the post!

RSS Feeds from the Library Congress: Thomas & More

You can now keep up with federal legislation through the Congressional Record Daily Digest on an RSS feed from Thomas! See the Library of Congress list of RSS feeds available here. Besides the Thomas feed here are some that will probably be faves for the OOTJ crowd:

Copyright Office:

Legislative Developments
Federal Register Notices
NewsNet (deadlines for comments, hearings, etc.)
What's New (alerts on Copyright Office website postings)

Law Library of Congress:
Global Legal Monitor
Legal Research Reports
Webcasts

For Librarians
L.O.C. Classification Weekly Lists
L.O.C.S.H. Weekly Lists

But scroll around the site, too. There are lots of fun things, as always at the Library of Congress! There is Poetry and Science and History and Folklife and Photographs. So visit the RSS list and stroll around for yourself. Even if you don't want the RSS feed, at least feast your eyes on all the cool stuff at our national library!

Handy Bookmarklet!

Just a cool little tool to make the web nicer and more readable:

Readability from Arc90 Lab.

A bookmarklet is a little piece of script that tells the computer how to do something additional. In this case, you can control how the font appears and how large; you control the size of the margins as well. These functions add a huge amount to the readability of the web pages you visit.

The name, bookmarklet, tells you how you make it work. You either save it as a bookmark, or drag it into your web browser's toolbar. Then, when you visit a web page where you want to use it, click on the Readability bookmark up there, and choose how you want to force to page to appear. Yay!

For other bookmarklets of extreme usefulness, visit http://www.bookmarklets.com/. You can browse through their ever-growing list, for Windows, Mac, Unix, and more. More than 150 are available. One of my favorites:

Page Freshness (there is a version for frames and without)- shows how recently the page was updated. This is very useful since the major browsers no longer have a function that displays this information.

Pay attention to the browser information. If you use Firefox, Bookmarklets may not be so helpful as I did not find any that were compatible with the current Firefox version.

Sunday, May 03, 2009

Google Maps' Historic Maps of Japan Stir Up Trouble

The Boston Globe has an A.P. story by Jay Alabaster here about Google Maps. They posted a set of wood cut maps of feudal Japan, already available on other websites. But somehow, being on Google maps, as a layer feature, seems different.

The maps date back to the country's feudal era, when shoguns ruled and a strict caste system was in place. At the bottom of the hierarchy was a class called the "burakumin," ethnically identical to other Japanese but forced to live in isolation because they did jobs associated with death, such as working with leather, butchering animals, and digging graves.

Castes have long since been abolished, and the old buraku villages have largely faded away or been swallowed by Japan's sprawling metropolises. Today, rights groups say the descendants of burakumin make up about 3 million of the country's 127 million people.

But they still face prejudice, based almost entirely on where they live or their ancestors lived. Moving is little help because employers or parents of potential spouses can hire agencies to check for buraku ancestry through Japan's elaborate family records, which can span 100 years.

An employee at a large, well-known Japanese company who works in personnel and has direct knowledge of its hiring practices, said the company actively screens out burakumin job seekers. "If we suspect that an applicant is a burakumin, we always do a background check to find out," she said. She agreed to discuss the practice only on the condition that neither she nor her company be identified.

Lists of "dirty" addresses circulate on Internet bulletin boards. Some surveys have shown that such neighborhoods have lower property values than surrounding areas, and residents have been the target of racial taunts and graffiti. But the modern locations of the old villages are largely unknown to the general public, and many burakumin prefer it that way.

Google Earth's maps pinpointed several such areas. One village in Tokyo was clearly labeled "eta," a now strongly derogatory word for burakumin that literally means "filthy mass." A single click showed the streets and buildings that are currently in the same area.

Google posted the maps as one of many "layers" available via its mapping software, each of which can be easily matched up with modern satellite imagery. The company provided no explanation or historical context, as is common practice in Japan. Its basic stance is that its actions are acceptable because they are legal, one that has angered burakumin leaders. (snip)

Printing such maps is legal in Japan. But it is an area where publishers and museums tread carefully, as the burakumin leadership is highly organized and has offices throughout the country. Public showings or publications are nearly always accompanied by a historical explanation, a step Google did not take.

Matsuoka, whose Osaka office borders one of the areas shown, also serves as secretary general of the Buraku Liberation League, Japan's largest such group. After discovering the maps last month, he raised the issue to Justice Minister Eisuke Mori at a public legal affairs meeting on March 17.

Two weeks later, after the public comments and at least one reporter contacted Google, the old Japanese maps were suddenly changed, wiped clean of any references to the buraku villages. There was no note made of the changes, and they were seen by some as an attempt to dodge the issue quietly.

"This is like saying those people didn't exist. There are people for whom this is their hometown, who are still living there now," said Takashi Uchino from Buraku Liberation League headquarters in Tokyo.
However, erasing the burakumin designations of the neighborhoods disturbs me as well. It is not my fight, but it smacks of the New Speak in George Orwell's dystopian novel, 1984, in which the government simply changes history to match its current needs. I wrote earlier about my discomfort over my home state of Kentucky and its beautiful state song, "My Old Kentucky Home." This song by Stephen Foster, was published in 1853, when slavery was legal, and Kentucky was a slave-holding state, at least in the central portion where plantations made economic sense. It was made the state song of Kentucky in 1928, many years before our sense of racial equality began to be more widely held, or at least reflected in public speech. The song is sort of cemented as the state song because it is so tradition-bound to the Kentucky Derby (which just ran). But, some of the lyrics make modern people cringe, so many people have changed some of the more politically incorrect verses. We may be removing the discomfort today, but we are also whitewashing the evidence of our racist past. We should not try to erase our history or re-write it to suit our convenience or sensibilities. This seems to me a very dangerous course, whether in Kentucky racial politics or in Japan caste history.

However, it seems that Google missed a major step that is taken by museums an others in exhibiting the historic maps by including some explanatory text along with the label. It is, I suppose, a bit like saying, "no offense intended," or holding up a sign that says, "for educational purposes" when you display Huckleberry Finn. This is an interesting thing to remember when any entity ventures into foreign territory -- we develop a tone deafness when we are operating in an unfamiliar culture. It would be wise to check carefully with some local consultants who work with similar material!

Saturday, May 02, 2009

Scalia and Privacy

Wow! Justice Scalia does not do humor, does he? Or irony, either. See a delicious collection of links and dish at Above the Law here. Justice Scalia had spoken at a privacy conference hosted by the Institute of American and Talmudic Law, and reported at Concurring Opinions here. The justice was reported as being untroubled by privacy issues:

Scalia said he was largely untroubled by such Internet tracking. "I don't find that particularly offensive," he said. "I don't find it a secret what I buy, unless it's shameful."

He added there's some information that's private, "but it doesn't include what groceries I buy."

Data such as drug prescriptions probably should be protected, he said, suggesting areas off-limits to data gatherers could simply be listed for legal purposes.
And that is why, according to Above the Law, Fordham Law Professor Joel Reidenberg assigned his class on information privacy to create a dossier on private information about Justice Scalia and his family. The students were amazingly resourceful and managed to produce quite a dossier (online, but password protected -- that's what university counsel are good for!).

Prof. Reidenberg, like any proud teacher, sent Justice Scalia a copy of the dossier (or maybe the interfering busybodies at Above the Law clued Scalia in on the dossier -- the post is not clear!). At any rate, the justice was not amused.

Friday, May 01, 2009

Justice Souter announces retirement


The A.P. article, The Best Job in the Worst City, by Michelle Salcedo, is a nice, short profile, and was one of the earliest filed reports on Justice Souter's announcement that he is retiring. From Boston.com, Political Intelligence column, a nice bit of speculation on the replacement by Foon Rhee here. A bonus, if you follow the link, is a video of NECN's (New England Cable News Network) report on Justic Souter's retirement.

What kind of jurist will President Obama look for to replace David Souter on the Supreme Court?

Based on what he said as a candidate, perhaps someone very much like Souter, at least in a more moderate, restrained judicial philosophy.

While the president was rather circumspect during the campaign, he did suggest he would like those with real world experience and empathy for the vulnerable, possibly expanding the pool of candidates beyond the usual farm team of federal appeals court justices. (All nine justices now are former federal appeals court judges.)

In an interview with the Detroit Free Press editorial board last October, he praised Souter and Justice Stephen Breyer as "very sensible judges. They take a look at the facts and they try to figure out: How does the Constitution apply to these facts? They believe in fidelity to the text of the Constitution, but they also think you have to look at what is going on around you and not just ignore real life.

"That's the kind of justice that I'm looking for," he continued on. "Somebody who respects the law, doesn't think that they should be making the law, but also has a sense of what's happening in the real world and recognizes that one of the roles of the courts is to protect people who don't have a voice."

He added that the "special role" of the court is to protect "the vulnerable, the minority, the outcast, the person with the unpopular idea."

"We need somebody who's got the heart, the empathy, to recognize what it's like to be a young teenage mom," Obama said at a Planned Parenthood conference in 2007. "The empathy to understand what it's like to be poor, or African-American, or gay, or disabled, or old. And that's the criterion by which I'm going to be selecting my judges."

In the Detroit interview, while he also praised the more liberal Earl Warren, William Brennan, and Thurgood Marshall as "heroes of mine," he added, "that doesn't necessarily mean that I think their judicial philosophy is appropriate for today."

Obama went on to say that while activist judges were needed to "break that logjam" on racial discrimination, he wasn't sure the same was needed today. "In fact, I would be troubled if you had that same kind of activism in circumstances today."
This UPI report speculates that Obama may select Judge Sonia Sotomayor, an experienced Hispanic judge from the Second Circuit Court of Appeals. Another possibility is Diane Wood, a federal judge in the Seventh Circuit Court of Appeals and former colleague of the President from Chicago. The UPI article also mentions former Harvard Law dean Elena Kagan as a possible Justice, but discounts her chances because she is doing such a good job where she is, as solicitor general (isn't that the way of the world?).

I will miss having Justice Souter at the Court. Despite being nominated by Republican President George H.W. Bush, he was a very solid, sensible voice and to me, seemed middle of the road (does that tell you something about my politics?) Good luck, Mr. Souter! I hope you don't have to wait too long.

This nice photo of Justice Souter is from http://usinfo.org/enus/government/branches/souter.html where they include a very nice brief c.v. of the justice as well.

Google Books Settlement May be Hitting Trouble

The New York Times reported here yesterday that the Google Books settlement has attracted interest from the Justice Department, who are starting to wonder whether it violates antitrust principles. (you think?)

It also reports that the judge, Denny Chin, has ordered the settlement will be delayed another four months at the request of authors, who want more time to consider the terms of the settlement. The original deadline of May 5 will now be extended into September.

Library Journal has a brief statement about it here, with a very helpful and wide-ranging set of links here.