Attention readers: This blog has moved to a new home at https://chenghlee.wordpress.com/.

Monday, January 16, 2012

Mozart's Requiem and the public domain

I had the opportunity this past weekend to see the director's cut of Amadeus at the Alamo Drafthouse, which was being hosted by the Golden Hornet Project (GHP) as a fundraiser for their project finishing Mozart's Requiem. Being an Austin-based organization, GHP is doing this work "in collaboration with composers from the rock, hip-hop, video game and avant garde music scenes" and producing a version far different from what Mozart ever imagined—by, for example, modifying the "Larcimosa" movement to use the musical structure of a modern movie trailer. During the live preview performance of the new "Larcimosa", I started thinking about how different projects like this would be had Mozart's Requiem been protected by modern copyright law.

Commissioned as a work for hire, the Requiem would, under the terms of the Mickey Mouse Protection Act (AKA the "Copyright Term Extension Act" or CTEA), be subject to a 95-year copyright. To get an idea of how long a span of time this is, look at the maps below. When Mozart died in December 1791, the United States had just established itself as a new nation, with a three year old Constitution (still missing the Bill of Rights) and 14 states. By the time a 95-year copyright term would have expired in 1886, the U.S. was a nation of 38 states with what we would recognize as its modern shape. By 1886, the American Civil War had been fought, the Transcontinental Railroad built, and the telephone invented.

The United States in 1791 (left) and in 1886 (right). Pink regions denote states, yellow U.S.-controlled territories, and gray territories controlled by other countries. Images by Golbez via Wikicommons.

In 1791 Europe, Louis XVI of France still had his head, Spain was still a dominant colonial power, and the Holy Roman Empire still existed. By 1886, France and Spain were dramatically weakened, the HRE had crumbled, and imperial power was now wielded by England and the newly unified nations of Germany and Italy.

Put another way, had Mozart's Requiem been a new work covered by the CTEA, it would be in the public domain today and freely available for GHP to compose their movie trailer version of "Larcimosa" only if the original had been written prior to 1917.

So, how exactly does such a long copyright term help "promote the progress of science and useful arts"? Is there really a work of art so wonderful that its creation can be incentivized only by offering a period of protection lasting longer than nearly all human lifetimes?

This is not to say that everything should be public domain should upon creation. Artists, writers, and yes, even scientists deserve compensation and legal protections for the intellectual property that they create. But I genuinely find it hard to believe that those artists, writers, and scientists I know would stop producing if tomorrow their copyrights went from the current 95-years (or lifetime + 70 years) back down to the original 14-year term (renewable once to 28-years).

More importantly, creative endeavors build upon what has come before—art and science thrive when we have a vibrant remix culture. I enjoyed last night's performance precisely because GHP didn't have to ask permission or pay a royalty to remixed eight bars of a requiem mass. I look forward to 2014 when GHP's fully completed work is supposed to be released. But I sincerely hope that the opportunity for someone to take that work and freely create some new vision of Mozart's Requiem comes along in 2028 or 2042—not in 2109 as current law dictates.

CTEA and its follow-on acts (DMCA, Eldred v. Ashcroft, and now SOPA/PIPA) have closed the doors on our public domain and our remix culture. It's time to push back at Congress (and, let's face the truth, its RIAA/MPAA underwriters) to open those doors again.

Wednesday, January 11, 2012

Another bad science headline

Bad, usually meaning "not quite accurate", science headlines really bother me. Today's example is this BBC article proclaiming "Exoplanets are around every star, study suggests". Except it's not entirely true, is it?

If you read the Nature letter [1], you'll find that what the authors did was to use a small set of gravitational microlensing events associated with known exoplanets to derive estimates of the number of exoplanets across a range of masses (Figure 2 from the paper). Using this plot and applying a bit more math, the authors estimated the expected number of planets (of all masses) in the galaxy, and by dividing that number by the number of stars, they discovered that "on average every star has [~1.6] planets…in an orbital-distance range of 0.5–10 AU" (emphasis mine).

Note, however, that contrary to what the BBC article suggests, this paper does not actually provide observational evidence that every star has at least one planet. The phrase "on average" is critically important and should not have been omitted from the headline—the authors' claims are no where as definitive as "exoplanet(s) around every star", and I'm entirely confident that if we pointed a telescope in just the right direction, we would find a star in our galaxy with zero planets.

I know all of this sounds overly pedantic, but when reporting about science, accuracy is paramount. I strongly dislike the seemingly pervasive perception in popular science reporting that grand claims, even when not entirely supported by the study being reported on, are necessary to make discoveries sound interesting and/or exciting. Such claims, especially when made in headlines, are a great disservice to science, as they leave the audience with a highly distorted view of what has been accomplished and often put scientists in the awkward position of appearing to abandon great discoveries.

In my personal conversations, I often refer to this as the "'broccoli cures cancer' effect", wherein a single experiment finding that compound A, found only in trace quantities in broccoli, kills tumor cells in vitro generates the headline "Broccoli cures cancer!!"

Rant about bad science headlines aside, this is a really cool study. This, along with a previous study [2] identifying planets not gravitationally bound to a star, leads us to conclude that "planets…in our Galaxy thus seem to be the rule rather than the exception."


References

[1] A. Cassan et al. "One or more bound planets per Milky Way star from microlensing observations". Nature 481: 2012. pp. 167–169.

[2] T Sumi et al. "Unbound or distant planetary mass population detected by gravitational microlensing". Nature 473: 2011. pp. 349–352.

Tuesday, December 27, 2011

Journal impact factors and the reliability of research results

Via researchblogging.org: A paper in Molecular Psychiatry (subscription required) finds that impact factor weakly correlated with unreliability of research papers, at least in gene association studies. However, there's no reason to think that this finding is not generally applicable to other areas.

It supports statements I've repeatedly made telling people to be careful about immediately believing stuff that appears in high profile/"sexy" journals like Nature Genetics. The nitty gritty part of science (read "replication studies") are usually published in less well-known and less sexy journals, and lord knows it's extremely difficult, if not impossible, to get negative results (i.e., "see, they were wrong—there's actually no signal there at all") published.

That said, I still want at least one of my papers to eventually appear in Nature Genetics (scientist street cred).

Tuesday, December 20, 2011

What should computational papers report?

After spending a part of the weekend helping a co-author figure out how we produced a table for a soon-to-be-published paper, I started thinking about what the minimum reporting requirements for a computational paper are or should be in order to guarantee reproducibility of the results.

Clearly, the minimum standard should be the name and version number of the software package(s) used directly in the analysis (e.g., in my case, DNAcopy 1.22.1). Version numbers are critical, as the algorithms in these packages occasionally change in non-obvious ways, leading to startlingly different results. Unfortunately, in my experience, a large number of papers fail to report version numbers (though this is slowly starting to change), and in the times I've been called upon to review a paper, I've always asked authors to include version numbers of all software they used in their edits prior to publication.

In his 2009 paper on repeatability of microarray analyses [1], John Ioannidis pointed out a major factor limiting reproducibility of results is the lack of enough details in how data were processed. I strongly suspect a key reason why such details aren't provided is that many, perhaps most, researchers don't keep particularly detailed notes about what software they used. Certainly, I've been guilty of this, which is why I had to spend my weekend tracking down this issue.

As far as details go though, it isn't clear to me how far up the tool chain we should go. Details about "parent" software systems (e.g., "Bioconductor 2.6/R 2.12.0") might be especially useful to reproduce results, especially if there have been bug fixes or tweaks in versions released since the original analyses were performed. In my analyses for this paper, I've noticed that the same version of DNAcopy gives slightly different numerical results when used with different releases of R, though not enough to affect the conclusions we drew.

Moreover, in cases where the researcher has built their own software (like I do with R releases), I wonder if details like compiler settings need to be provided or whether those are just minutia that only serve to clutter up the literature. Choices for architecture flags (e.g., "-mfpmath=sse -msse4.1" versus "-mfpmath=387" for gcc) or for aggressive optimizations (e.g., "-ffast-math" [2]) certainly have effects on numerical results, though I would argue that conclusions drawn from results that are especially sensitive to compiler flags are dubious at best.

Similar questions can be asked about whether versions for various linked libraries should be reported. Even "standardized" libraries can lead to different behaviors depending on how they are implemented. For example, R and GraphLab behave badly (i.e., segfault in certain circumstances) when linked to GotoBLAS but run fine when linked to ATLAS or Intel MKL.

A final consideration for reporting has to do with randomized algorithms, such as MCMC simulations. Of course, best practices say that the algorithm should be run several times (at least) to ensure that the results are reasonably stable and consistent. However, should the interests of "full disclosure" and reproducibility require us to report the random seeds used to generate our results? Again, I would be highly skeptical of any results that are that sensitive to the initial seed, and in any case, getting such values might be impossible to get if seeding of the random number generator happens deep in the bowels of code that we don't have access to.

My conclusions after thinking about this? Report the version of any packages you use and the version of the parent software; compiler flags, linked libraries, and random seeds are probably extraneous details that can be left out. Of course, there's also the question of how thoroughly code you've developed needs to be tested for numerical stability and other computational "artifacts", but that's an issue I'll discuss at another time.

[1] J Ioannidis, et al. "Repeatability of published microarray gene expression analyses". Nature Genetics 41(2):2009. 144–155.
[2] Though I would argue that you should never use -ffast-math for numerical software.

Monday, December 19, 2011

A bit of Christmas cheer

Those who know me personally know that three months of Christmas-themed ads make me more than a little cranky around the holidays. But seeing this "LGBT Welcoming" sign at the church near my house brought a smile to my face; I'm glad to be reminded that not every denomination is as strongly homophobic as the talking heads on TV would have us believe.