Monday, May 28, 2012

GATCACA

In many American states it is legal to screen and select on the basis of sex, for non-medical reasons. In fact, a 2006 study (see below) found that 9% of [preimplanation genetic diagnosis] procedures carried out in [in-vitro fertilization] clinics in the U.S. were performed for this reason. Other reasons include screening for an embryo with the same immune type (“HLA type”) as a current child who is ill and requires a transplant of some sort. Screening for these “savior siblings” was done in 1% of PGD procedures. And 3% used it for a reason I personally find jarring – to specifically select embryos with a mutation causing a genetic condition. This is usually in cases where both parents have either deafness or dwarfism and they want their child to be similarly affected. This gets into the political movement objecting to society labelling conditions as “disabilities”. I can sympathise with that to some degree – more for some conditions than others – but I think, if it were my child, I would still rather he or she could hear.
That's Kevin Mitchell, discussing GATTACA, an entertaining sci-fi movie with a respectable 7.8 imdb rating. Spoiler alert, the premise of the movie is that at some point in the future there will be strong stratification of people into two classes, the "valids" and the "invalids", based on whether they had healthy traits selected for via preimplanation genetic diagnosis.

It seems to me highly unlikely (<0.01%) that a nightmare scenario of this sort would actually occur. One of the main reasons is because of the large plurality of values among parents, as seen above. A prevailing reason people have kids is to propagate a form of themselves into the future, and in many ways it defeats the purpose when you select against certain traits or even perform some sort of genetic engineering.

The other reason is something we know now better than we did 15 years ago, when GATTACA was released. And that is that DNA doesn't actually explain all that much of physiology and behavior--there are also strong epigenetic effects as well as stochastic effects of gene expression. 

Sunday, May 27, 2012

No Darkness But Ignorance

Here's Nancy Kanwisher's suggestion on how to improve the field of neuroimaging:
NIH sets up a web lottery, for real money, in which neuroscientists place bets on the replicability of any published neuroimaging paper. NIH further assembles a consortium of respected neuroimagers to attempt to replicate either a random subset of published studies, or perhaps any studies that a lot of people are betting on. Importantly, the purchased bets are made public immediately (the amount and number of bets, not the name of the bettors), so you get to see the whole neuroimaging community’s collective bet on which results are replicable and which are not. Now of course most studies will never be subjected to the NIH replication test. But because they MIGHT be, the votes of the community are real.... 
First and foremost, it would serve as a deterrent against publishing nonreplicable crap: If your colleagues may vote publicly against the replicability of your results, you might think twice before you publish them. Second, because the bets are public, you can get an immediate read of the opinion of the field on whether a given paper will replicate or not.
This is very similar to Robin Hanson's suggestion, and since I assume she came up with the idea independently, it bodes well for its success. Both Hanson and Kanwisher are motivated to promote an honest consensus on scientific questions.

When John Ioannidis came to give a talk at the NIH (which was interesting), I asked him (skip to 101:30) for his thoughts on this idea. He laughed and said that he has proposed something similar.

Could this actually happen? Over the next ten years, I'd guess almost certainly not in this precise form; first, gambling is illegal in the US, and second, the markets seem unlikely to scale all that well.

However, the randomized replication portion of the idea seems doable in the near term. This is actually now being done for psychology, which is a laudable effort. It seems to me that randomized replications are likely precursors to any prediction markets, so this is what interested parties should be pushing now.

One objection is that these systems might encourage scientists to undertake more iterative research, as opposed to game-changing research. I have two responses. First, given the current incentives in science (i.e., the primacy of sexy publications), this might actually be a useful countervailing force.

Second, it seems possible (and useful) to set up long-standing prediction markets for a field, such as, "will the FDA approve an anti-amyloid antibody drug to treat Alzheimer's disease in the next ten years?". This would allow scientists to point to the impact that their work had on major questions, quantified by (log) changes in the time series of that market after a publication. 

Saturday, May 26, 2012

Evaluating The Regret Heuristic, Part II

In a comment to my post on how our regrets change over time, Eric Schwitzgebel asks, 
But why adopt regret minimization as a goal at all? Regret seems distorted by hindsight bias, status quo bias, and sunk cost bias, at least.
I've written before that projecting your future views about your present actions can be a good way to make decisions. So, Eric's prompting is a good occasion to re-evaluate that.

Given perfect information, the theoretically best way to make decisions is to 1) calculate the costs and benefits of each possible outcome, 2) estimate how your choice affects the relative probability of those outcomes, 3) use the costs and benefits as inputs to some sort of valuation function, and 4) make the decision with the highest probabilistic value. 

Cost-benefit analysis is a common way to implement this, with, say, QALYs as the value measure. If you have perfect information, this is just math. 

But as Ben Casnocha says, if you don't have enough information, that framework can break down. In particular, even when #2 is pretty straightforward, #1 can still be very tricky. For example, although studying for the LSAT makes it much more likely that I will earn a JD, it's still hard to quantify the precise costs and benefits of entering that earning that degree. 

Here is where the regret heuristic can be useful. Instead of explicitly tallying each cost and benefit, it asks: in total, which would you regret more: studying or not studying? 

This is in fact a simplifying measure, but there remains oodles of freedom in how you perform the regret estimation. For example, you can:
Ultimately, I still think that the regret heuristic can be a useful one. But tread carefully, as there are many crucial micro-decisions to make; it's not magic. 

Friday, May 25, 2012

Age's Stealing Steps

Michael Wolff has written a gripping but narrative-heavy article about the troubles he has experienced in addressing his mother's worsening dementia. It is hard not to feel for him and his family. Still, I think there are two perspectives which his piece underemphasized:

1) Many debilitated but cognitively intact individuals do have a good quality of life. For example, in a recent survey of 62 seniors with an average of 2.4 daily living dependencies and fairly good cognitive well-being (≥ 17/30 on the MMSE), 87% reported that they had a quality of life somewhere in the fair to very good range. I consider this to be a testimony to the resiliency of the human psyche. Also, it makes me worry that people will read the article and think that LTC insurance is only useful for those with dementia, which Wolff implies, when that is far from the case.

2) Why is it that many of the doctors depicted his story seem so unhelpful? There's little doubt that fear of litigation plays a role. For example, in a Mar '12 study, over half of the 600+ palliative care physicians surveyed reported being accused of euthanasia or murder within the past five years. In many respects this is a legislative issue, and I wish his article had discussed that angle more.

Many pointers in this post go to the excellent blog GeriPal. 

Thursday, May 24, 2012

Should Revenge Have Bounds?

I recently finished Steven Pinker's book attempting to explain the decline of human-on-human violence over the last twenty thousand years. All in all, I recommend it. It has noteworthy psychology nuggets on nearly every page, explained with good data and lucid metaphors. I especially enjoyed how he built up many cute explanations of various phenomena--like the Freakonomics-popularized abortion theory of the crime rate decrease in the 1990s--only to soundly and evenly debunk them. My two major points of disagreement:

1) As Tyler Cowen argues, it is possible that although the mean number of causalities from interstate conflicts has been falling, the variance has been increasing. Aside from WWII, we can't easily observe this variance, though we can see signs of it in events like the Cuban Missile Crisis. Pinker employs per-capita log-scales for many of his charts and on these WWII does not seem quite as bad, but still it sticks out indelibly.

Wisely, Pinker does not project the decrease in violence indefinitely into the future, rather seeking to explain what we have observed so far, so his thesis is technically immune to this critique. Still, I imagine that there have been some not-easily observed historical aberrations which, if they had gone differently, would have meant that this book would never had been written. The winner's curse comes to mind.

2) As one reads about the incredible violence that occurs in US prisons, it is difficult not to wonder whether the benefits to decreasing violence always outweigh the costs. I have previously written about the protection vs freedom trade-off. The laudable decrease in person-to-person violence comes at the cost of constraining the actions of individuals by probabilistically putting them in prison. This is an imperfect process and has negative externalities in that it further exacerbates the burden of those locked up for non-violent crimes.

So, I would have liked to see more discussion about the violence in modern-day prisons and whether it is more apt to say that violence has been displaced rather than decreased. In a provocative article, Christopher Glazek argues that the US should be more like the UK and have slightly looser violent crime convictions which would make the conditions in prison slightly less awful. In most cases I would probably come down in favor of protecting innocent bystanders, but it is a conversation that needs to happen and that I wish Pinker had addressed. 

Sunday, April 29, 2012

A Brief History Of Bioinformatics, 1996-2011


That's from an interesting article by Christos Ouzounis. Here he discusses the "adolescence" period:
One factor in policymakers' high expectations might have been a certain lack of milestones: due to the field's dual nature, that of science and engineering, computational biology rarely has the “eureka” moment of a scientist's discovery and is grounded in the laborious yet inspired process of an engineer's invention. 
And there's this bit, too:
The notion of computing in biology, virtually a religious argument just 10 years ago, is now enthroned as the pillar of new biology.
So why has "bioinformatics" become less discussed? In part, because it has been so successful. 

Friday, April 13, 2012

Schelling Points And Bioinformatics

A lot of what think about when I do bioinformatics is how to set parameters non-arbitrarily. Basically I am looking for Schelling points: round, clear numbers that are easy to justify. The classic case is setting a p-value threshold to 0.05, which has been around for over eighty years and is still going strong, despite the haters. Other examples are setting e-value thresholds to 0.01 and setting Bayes factor thresholds of 10 as the first to indicate "strong" evidence. Like any threshold, these are arbitrary, but following the paradigm of statistics as rhetoric, their staying power make sense insofar as scientists need to be able to resort to standard procedures to settle debates. Anyway, I have no profound point here, I just think it's cool that a seemingly esoteric topic affects what I actually do on a day-to-day basis.


Is It Possible, In Principle, To Do Methodologically Sound Research?

In a paper published 31 years ago, Joseph McGrath argues (html, pdf) that the answer is no. Specifically, he claims that any research design faces two trade-offs: 1) being obtrusive vs unobstrusive (which maps to my terminology as acquiring info vs altering subject), and 2) being generalizable vs context-cognizant (which maps to my terminology as precision vs simplicity).

In his terminology, these trade-offs allow for the optimization of three distinct values (generalizability of samples to populations; precision in measuring variables; and context realism for the participants). Initially, I disagreed with this. To me, intuition suggests that there should be four points which maximize certain qualities when you are considering the intersection of two-trade offs: one in each corner of the 2-d space.

One way to get around this is if you claim that, in the context of this decision (study design), the trade-offs are not independent. For example, it might be very difficult for a design to be both highly generalizable and highly obstrusive.

Below I've drawn an example. Think of the dots as realizations of actually feasible study designs sampled from someone's mental generation process; i.e., they are probably not at the absolute extremes of the theoretical distribution, but with enough realizations, would come close.


I'm not sure that I agree with this exact distribution, and it would need some justification, but it seems like the only way to justify his three-pronged rather than four-pronged set-up.

Saturday, March 31, 2012

The Valiant Never Taste Death But Once

After reading this interesting excerpted article from Dick Teresi's book The Undead, which discusses the difficulties in defining death by a single, consistent set of criteria and the social qualms that stirs, I decided to check out the Amazon reviews. The associated ratings were (and still are) quite shockingly bad! They follow the classic "so bad it's good" distribution, with 5 5-star ratings, 1 3-star rating, and 33 1-star ratings. So, given that I am always up for a good controversy, I decided to read and review it myself. Ultimately I mostly side with the critics, giving it two stars. If you are interested in the subject matter, I'd suggest instead Kenneth Iserson's Death to Dust, which is a bit older but much more level-headed and thorough treatment of similar issues. 

Friday, March 30, 2012

What Makes Phrases Memorable

For 1000 movies, this study compared lines included on imdb's memorable quotes page to those that were not. People who hadn't seen the movies were able to pick the correct one 78% of the time, although, caveat lector, that's with only n = 68.

What features allow this above chance classification? The authors suggest 1) distinctiveness (i.e., a lower likelihood of coming from samples of standard English text), 2) generality (fewer personal pronouns, more present tense), and 3) complexity (words with more syllables and fewer coordinating conjunctions like "for" and "and").

Interestingly, their best support vector machine only correctly classified examples 64% of the time, so either the human data is somehow biased, or there are plenty more subtleties for machines to learn before they can best us humans in recognizing literary wit. 

Thursday, March 29, 2012

Indexing Wikipedia Article Submissions On Pubmed

I have complained before about few academics writing Wikipedia pages and instead writing reviews that few people will read. So, I feel compelled to admit that this is really cool:
We suggest a principal reason for this limited breadth and depth of coverage of topics in computational biology is one that affects a number of disciplines: reward. Authors in the biomedical sciences get academic reward for publishing papers in reputable journals that are indexed in PubMed and have associated digital object identifiers (DOIs).... 
Topic Pages are the version of record of a page to be posted to (the English version of) Wikipedia. In other words, PLoS Computational Biology publishes a version that is static, includes author attributions, and is indexed in PubMed. In addition, we intend to make the reviews and reviewer identities of Topic Pages available to our readership. Our hope is that the Wikipedia pages subsequently become living documents that will be updated and enhanced by the Wikipedia community...
I continue to be impressed by the innovation from the PLoS suite. 

Monday, March 26, 2012

Why Does Speed Variability Create Congestion?


Above are the results from one trial of an experiment designed to answer this question. Participants were randomly assigned to one of two groups, each with its own walking direction and color.

The authors defined "clusters" as groups of people walking in basically the same path, with some leeway. They then did simulations to determine the average lifetime of a cluster as a function of the group's variability in walking speed. As you can see, the greater the variability, the shorter the lifetime of the clusters.

N = the number of pedestrians in the simulation
This trend fits with their experimental results. Here's how the authors explain it:
[T]hose moving faster catch up with those walking slower, leaving an empty zone in front of the slow walkers ... [P]edestrians who are willing to walk faster than others make use of density gaps to overtake the slow walkers in front of them. By doing so, faster pedestrians move away from their lane, and meet the opposite flow head-on a few seconds later. This initial perturbation often triggers a complex sequence of avoidance maneuvers that results in the observed global instabilities. 
So here's a situation where more diversity, defined as inter-individual variability, leads to worse outcomes. Of course, as the authors mention, there are many other situations, such as collective decision making, where inter-individual variability is actually quite helpful.

Perhaps more diversity generally serves the function of pushing a group out of local optima. So you can think of diversity as shifting a group more towards the "explore" side of the exploration-exploitation trade-off. This would hurt in situations with a clearly defined goal, such as pedestrians walking in a circle as quickly as possible. But it might help in more complex situations.

Sunday, March 25, 2012

Comp Exams For Each Course

The solution I propose is comprehensive exams at the end of each course, much like Advanced Placement exams, that thoroughly and objectively distinguish students on merit alone. The emphasis in each classroom would then shift from fighting the teacher for high grades to cooperating with the teacher to learn the material necessary to perform on the exam.
That's from Andrew Knight, in an essay discussing problems that will not be new to anyone who is or has recently been in school; more here. This is exactly what I wanted during most of my science and math courses. The alternative is to place a greater emphasis on big standardized tests like the SAT, but there can be so much variability in results from just one day.

One question is whether such exams could be a part of classes that are less fact-based, such as history and english. There is actually a machine learning competition for automated essay grading going on right now. I don't pretend to know the answer to this question, but even if it is currently infeasible, that shouldn't stop the tests from being used in math, science, and foreign language classes. 

Saturday, February 25, 2012

Snub City At The Oscars

Of the Best Picture nominees, The Artist is currently the highest rated on imdb, at 8.4, though it will drop. A good comparison is Avatar, because both movies are technically adventurous, and they both have terrifyingly trite plots.

The main difference between Avatar and The Artist is that the latter is about the past, triggering nostalgia, whereas the former is about one possible version of the future, and is thus discomforting. This is why The Artist will win Best Picture and Avatar didn't come close. (No movie set in at any point in the future has ever won the award.)

But of course, the best movie of the past year is A Separation. The fact that it wasn't even nominated just showcases the Academy's striking anti-foreign film bias.

####

It is obviously very fun to hate on the Academy, and there are many good reasons to do so, but as imdb user Fish_Beauty reminds us, this year is highly unlikely to go down as the biggest black mark of all time. Here are the lowest rated Best Picture winners:

Around the World in Eighty Days (1956) 6.8/10 9,106 (which won over the amazing The Killing)
The Greatest Show on Earth (1952) 6.7/10 5,177 (which won over High Noon)
The Broadway Melody (1929) 6.4/10 2,459
Cavalcade (1933) 6.3/10 1,421
Cimarron (1931) 6.1/10 1,739 (which won over the best silent film ever, City Lights)

These are truly embarrassingly bad films.