Saturday, June 11, 2011

Trade-Off #20: Quality vs Quantity


If the above photo had two flowers instead of just this one, would each seem slightly less pretty? I don't know for sure (we are talking about aesthetics, after all), but I imagine that most would say yes, because each flower would stand out less from the background and seem less unique. The only exception would be if the flowers somehow contrasted or augmented each other's beauty.

In scarce environments, and aside from such cases where individuals interact to produce effects greater than the sum of their parts, the average quality of an agent's choice will be inversely related to its quantity. Here are some examples:
  • People often wonder whether they should spend lots of time and energy pursuing one high-quality mate, or distribute those resources pursuing many lower-quality ones. This is a very general quandary, which many if not all reproductive species face. (see here)
  • In searching a text, an increase in the proportion of relevant results to total results typically comes at the cost of missing more of the possible relevant results from the whole search space. (see here)
  • For an individual using an online social network, adding more "friends" usually decreases the quality of his relationship with his average connection. (see here)
Trade-offs are found everywhere, even in making a list of the most important and widespread trade-offs. So given the quality vs quantity trade-off that we face in adding more trade-offs to this list, the first draft of the canon will end here.

(photo credit to domesticated diva)

Sunday, June 5, 2011

Validating The Next Revolutions

From my hopefully not overly-insular vantage point, the two books which have had the biggest impact in the first half of 2011 have been Richard Arum and Josipa Roksa's Academically Adrift and Tyler Cowen's The Great Stagnation.

There are many similarities between them. They both address broad, ongoing trends in American society, in higher ed and macroeconomics. They have both been read largely in e-book format, AA due to the prohibitive price of the print version, and TGS due to some brilliant/lucky marketing. And, oddly, their approaches both owe at least some homage to the views of entrepreneur and raconteur Peter Thiel, who has widely-discussed qualms with higher ed, and who won the dedication of TGS for his insight into the lack of innovation and growth in our economy.

This last point deserves some drawing out. For all the hype, the overarching theme of neither book is necessarily novel. Arum and Roksa make many points that were assumptions, not conclusions, of conversations at Vassar's cafeteria. No one ever wondered with incredulity, "wait, instructors are gaming their end of semester ratings?" The case is similar for Cowen's thesis. Individuals who decry the sluggish innovation and dormant middle class prospects in America are hardly in short supply. Just notice how many express fears that America doesn't "make anything" anymore.

But what both books do accomplish is to make these everyday arguments more rigorous. And, perhaps because the authors are academics, they also serve to validate what might otherwise be seen as merely mumbles and whimpers amongst the broader populace.

The impact of these books also speaks to the fact that we as a culture are still only barely embracing the brave new world of the internet. Content-wise, these both could easily have been "merely" articles. Publishing in some gated academic journal would obviously have reached few, but even if they had been published in the popular press, I doubt they would have had the same success. Regardless of the e-book format, we still love the idea of the book. For instance, I get expontentially more comments and questions IRL about my shelfari page than my del.icio.us page, although I've surely invested more time and energy into curating the latter.

Cowen himself says, on his blog, that, "we’ve yet to really organize our economy around the internet, as we someday will, and then the gains will be enormous." Perhaps the impact of these e-books could be considered Exhibit A of both our current lack of mobilization around the internet idea economy, as well as its potential once we do get our act together.

Saturday, June 4, 2011

Boredom Trade-Offs?

Aaron Haspel says that "an above-average capacity for boredom is optimal; a superior one is disastrous." Somewhat similarly, Mike Tully argues that becoming bored with a pursuit will inhibit artistic and athletic greatness.

Perhaps we can think of the capacity for boredom as a cognitive trait that pushes you towards the plasticity side of the plasticity vs specialization trade-off.

But really this idea seems a bit too vague. It's not clear whether one's capacity for boredom extends uniformly across all domains, and there are many other factors involved.

For example, you could argue that Ted Williams was able to specialize because he never grew bored of baseball, or you could argue that he specialized because he so quickly grew bored of everything else. With the former frame he has a below-average capacity for boredom, while with the latter it's above-average, but the end result is still the same.

Monday, May 30, 2011

Which Parts Of Crowds Are Wise?

Peter Freed has written a pretty ambitious critique of Jonah Lehrer's summary of this study (pdf) on the wisdom of the crowds. The crux is that:
But now that I realized he really meant median, and that maybe he didn’t know what median meant.  Because median guesses are not guesses by a crowd, as Lehrer states.  They are guesses by a single person... [Lehrer] is talking about that 0.7% single-person data point: one person, selected after giving their answer, got close to the correct answer on one of six questions.  One person guessed 10,000 when the answer was 10,067.  That’s one hit out of 144 x 6 = 864 attempts.  That seems about right to me, from a common sense perspective. Which is to say, that is a shitty batting average.
Scrolling through the comments, I was pleased to see Ian Sample point out the critique of Freed's critique that I was going to make:
In Wisdom of Crowds studies you can look at the mean and / or the median. The median usually gives the best result if the guesses *do not* follow a normal distribution. The mean, of course, exploits the error-cancelling advantage that WOC is known for, that is, as many people under-estimate as over-estimate the right answer, so averaging cancels all but systematic biases. But to my point. To dismiss the median answer – one guy’s response – misses the fact that without the crowd you have no median answer to dismiss. Without the crowd, you do not know which value to pick. That’s the whole point. The crowd steers you to the median value, which in many cases outperforms the mean.
The median is indeed generated by only one person, but it becomes interesting only in the context of all the other estimates. It is useful here because it offers resistance to outliers. For example, some less numerate soul might have guessed 1,000,000, which is way off from the true value of ~ 10,000, thus skewing the arithmetic mean. In that case you'd much prefer a more robust statistic like the trimmed mean or the median.

In prediction markets, the most recent price of a transaction doesn't always best represent the current beliefs of the market. There's more info if you look at the whole distribution of orders. Similarly, it is unfair of Freed to dismiss the whole data set just because one type of estimator is flawed. This is one of the coolest parts of statistics, using potentially counter-intuitive methods to extract useful info out of data, to find the wisdom in the crowds.

Sunday, May 15, 2011

Fighting The Lernaean Hydra Bias

I'll just only mention the heads I do cut off

In one Greek myth, Hercules takes on the task of killing a serpent-like, many-headed beast. This is made more difficult by the fact that its heads regenerate, so even if Hercules chops one off with his sword, another will simply sprout in its place. John Ioannidis uses this frustrating scenario as an analogy for a problem in the world of scientific publishing in his discussion of meta-analyses (doi:10.1002/jrsm.19).

The example Ioannidis employs to explain this problem is his experience doing a meta-analysis on the pharmacogenetics of certain polymorphisms for asthma treatment (doi:10.1097/01.fpc.0000236332.11304.8f). There were many studies that fit the criteria, but they each evaluated their own endpoints and genetic contrasts. That is, in most of the studies, the vast majority of possible correlations that could have tested with the data between phenotype and genotype were either not done or not reported.

So the surface problem, in so far as this case generalizes to others, is that published studies are not as exhaustive as they could be. But the central, troubling implication is that these studies do not fail to be exhaustive because of time or computational constraints, but because the researchers want to emphasize the usefulness and/or interestingness of their results. This is more insidious--this is why the hydra heads regenerate.

Now, one can use meta-analysis to retrospectively "chop off" findings that are truly insignificant by combining the results of many different data sets. But meta-analysis itself can be biased in many ways (e.g., during study selection), and moreover, later researchers can just come back to the issue and cherry pick more novel associations, thus "sprouting" more statistically significant findings.

When faced with the hydra, Hercules knew he couldn't go it alone, so he called on his nephew for help, who suggested that they cauterize the stumps with fire before the heads could regrow. An analogy to this strategy might be to post warnings on the electronic copy of papers that have been called into question by later studies. Such a warning would be much milder and hopefully less political than a retraction, which typically implies some sort of error. Publishing a potentially informative result that is eventually overturned is still laudable.

But instead of this type of patchwork fix, a more fundamental approach seems more fruitful. In the original myth, only one of the hydra's heads was truly immortal, and this was the one that Hercules needed to chop off to finally defeat the beast. The immortal head of the scientific publishing hydra is the incentive structure pushing researchers towards significance hunting in the first place.

Reworking these incentives is what Ioannidis is fundamentally arguing for, as the way to kill the Lernaean hydra bias once and for all: more standardization, more consortia, and more of a push towards openness and replicability. Every study might combine previous data with its own for estimating the posterior probability of the parameters it is examining, and all research might be seen as a continuous and cumulative meta-analysis. Maybe one day.

(photo credit to Frank Rafik)

Saturday, May 14, 2011

Color Me Old-Fashioned

A fantastic idea from Risto Saarelma on how to re-design the comment section of the website Less Wrong:
Provide an ambient visual cue on how old a comment is. First idea is to add a subtle color tint to the background of each comment, that goes by the logarithm of the comment's age from reddish ("hot", written in the last couple of hours) to bluish ("cold", written several months or more ago). Old threads occasionally get new comments and get readers in via them, and the date strings in the comments require some conscious parsing compared to being able to tell between "quite recent" and "very old" comments in the same thread by glance.
Too true. Who takes the time to read the actual date of a comment? This way you wouldn't have to.

This sort of subtle clue is something that people will appreciate and pick up on quickly. For example, on the blog Marginal Revolution you can always tell whether Tyler or Alex is posting because Tyler only capitalizes the first word in the title of his posts whereas Alex capitalizes all of the words in his titles. Knowing this, you won't have to waste time scanning the byline as you plow through your RSS feeds because you'll already know who wrote it from the title.

Sunday, April 24, 2011

Three Thoughts For Spring

1) If there are things that you can do to increase your perspective on your current problems, like taking a weekend off or writing a journal, then there also should be things you can do to decrease it. But I can't think of any. So is perspective the sort of thing where your typical state is a steady decrease unless you actively increase it through certain, discrete actions? You either have to agree with this model or describe specific ways you can lose perspective.

2) Who will systematically review the systematic reviews? Cochrane reviews, that's who.

3) One thing I wonder, as I try to get into Anki, is how we could make spaced repetition learning into a game. And I don't mean some boring game, like "how many flashcards can I get right today?", but a sweet game, with long-term goals and leveling up and side-missions and bad guys to defeat. I don't know if it could be done, but couldn't you imagine this as a big part of the future of education?

Friday, April 22, 2011

Sometimes Simple?

"Everything is more complicated than you think." - Synecdoche, New York

Is this true? No way. We can easily come up with counterexamples. Take any superstition, like some time in your youth that you were afraid of monsters in your closet and it turned out to just be a broom propped up at a weird angle. Which is more complicated--the angled broom or the hidden monster? That's a layup.

So the better question is: are things on average more complicated than you think? It sort of seems like it. But part of the problem is that we tend to simplify old beliefs to make our current ones look more intelligent in comparison. Consider the history of Dale's principle. Some authors understand this to mean that neurons can only release one type of neurotransmitter. If stated in this form, it's clearly wrong, and so newer researchers can claim credit for debunking it. But when you look at its inception, it turns out that "one neuron = one neurotransmitter" is probably not what the principle was actually meant to imply. So our intuitions about how our beliefs tend to change probably speak more to what we currently believe about the past than to what we will believe in the future.

If we could show conclusively that things in general are more complicated than we think they are, that'd be good to know, because if reality tends to deviate in some predictable way from your expectations, then you're doing something wrong. But I'm not sure that the answer will turn out to be so simple.

Monday, April 18, 2011

Milestones In...

I've recently discovered Nature's milestones index, which links to timelines of the major advances in the research of many fields: light microscopy, gene expression, development, etc. These were chosen by panels of many experts. For example, these 40 helped decide the milestones in cancer research. The timelines have links that explain why each milestone was important, like this one on the first methods of DNA sequencing. Awesome.

I wonder if there's some way that we could allow people to vote on these milestones in a similar way that others have set up for people to vote on milestones in computer science? If so, we could tap into what seems to me like the most productive form of crowdsourcing, where experts define the field, and then the masses rank the entries in that field.

Sunday, April 17, 2011

The Wisdom Of Whuber

That's William Huber, whuber for short, dispensed in his answers at the relatively new stats Q&A site, Cross Validated. His answers are the best on there, reputation normalized to the number of answers (with shrinkage). Here he writes about whether the median is a better summary stat than the mean:
Statistics does not provide a good answer to this question, IMO. A mean is ok to use, too, and is relevant in mortality studies for example. But ages are not as easy to measure as you might think: older people, illiterate people, and people in some third-world countries tend to round their ages to a multiple of 5 or 10, for instance. The median is more resistant to such errors than the mean....  Thus, for demographic, not statistical, reasons, a median appears more worthy of the role of an omnibus value for summarizing the ages of relatively large populations of people.
Here he writes about the biggest questions in statistics, from which I'll reproduce two (emphasis his):
  • Coping with scientific publication bias. Negative results are published much less simply because they just don't attain a magic p-value. All branches of science need to find better ways to bring scientifically important, not just statistically significant, results to light. (The multiple comparisons problem and coping with high-dimensional data are subcategories of this problem.)
  • Probing the limits of statistical methods and their interfaces with machine learning and machine cognition. Inevitable advances in computing technology will make true AI accessible in our lifetimes. How are we going to program artificial brains? What role might statistical thinking and statistical learning have in creating these advances? How can statisticians help in thinking about artificial cognition, artificial learning, in exploring their limitations, and making advances?
And here he writes about whether you should use a normal distribution to assign student grades:
I think that if any of those 800 students were to read this question, they might be offended. How well did they perform? How much learning was accomplished? That is what a grade should reflect, not some arbitrary statistical summary of their position in a group. IMHO this question should be recast in terms of teaching objectives, not statistical procedure, such as "what is a good way to convert raw scores to grades in a way that respects student accomplishments and advances the learning objectives of this class?" Statistics can help, but blind statistics--like standardization--will not.
Although they are often quite quantitative, his answers show how good stats rely on far more than just math. 

Saturday, April 16, 2011

Testing Robustness vs Fragility In Chemotaxis

receptor modulation, from doi:10.1371/journal.pone.0011224

Bacterial chemotaxis depends (like most biological functions) upon an intricate signaling network, in which all of the molecules (mostly enzymes) must work in unison. Oleksiuk et al have just published a paper (doi:10.1016/j.cell.2011.03.013) showing convincingly that ambient temperature would affect many of the molecular components of this pathway in E. coli, but the system is optimized to work despite variations in temp.

For example, they show that the pathway's receptor kinases have modification states (see above) with opposing temperature dependencies. So, when the temp changes, the activities of receptors with different mod states compensate for one another to allow the system to maintain the same function.

Since there is apparently a canonical trade-off between robustness and fragility, E. coli's robustness to variability in temperature should come with some costs. One form this cost could take is that it would make a mutation to a temp sensor gene more deleterious. Another form this cost could take is that it would make the bacteria more susceptible to viruses that mess with parts of the temperature regulation system. Maybe some group will show one such cost to be present, or even to be the dominant force? We'll see.

Monday, April 11, 2011

Ranking Ideas In Science

Last summer I bought and read The 100 Most Important Science Ideas after noticing it in a bookstore (my first mistake--I should have checked the ratings online first). I learned a fair amount from it, but I have to say it fares miserably in its attempt to actually rank ideas in science. First, it only covers three subjects: genetics, physics, and math. Second, even within those subjects, the topics are listed merely by date of discovery, not importance. Finally, there was little to no space devoted to methodology.

Subsequent attempts to find lists of the most important science ideas, via google searches and cold e-mails to potentially knowledgeable people, have also left me empty-handed. Uncool.

A good system to rank science ideas, both historically and as they are published, would be so money. The historical list would be really useful for educating the next generations and as outreach to the public. And dynamic, post-peer review ratings would help researchers use their precious time reading the best papers, instead of relying solely on the impact factor of the journal.

Given the above, you can imagine my immense pleasure to see Scott Aaronson's announcement today of a site that allows anyone to vote on milestones in computer science.

There are at least a couple of ways this voting could be done. The first way, as they currently have the site set up, is that users can pick and choose to vote any individual idea on the list up or down. The advantage of this is that users can choose to vote only on the ideas that they actually know something about.

The second way is that the site could present two options to users, the users would choose which of those two are better, and then an algorithm would use those preferences to rank all the ideas. The advantage of this is that it's more fun. Indeed, you might recall that a similar system was employed by the young Mark Zuckerberg in facemash. Wait, you haven't seen The Social Network? C'mon now, it's #190 on the top 250. Step your game up.

Anyway, bravo to Jason, Ammar, and Scott. Now we just need to create similar lists for all other scientific disciplines, incentivize people to vote on them, and aggregate the results. We'll also def need some kind of normalization to account for the fact that computational pursuits will have at least 10x the votes, because those people are on their computers like all day.

Saturday, April 9, 2011

When Can We Measure Grit?

Jonah Lehrer's interesting, 15,000 character article about measuring NFL quarterbacks concludes by saying that we have neglected grit in favor of IQ because "grit can't be evaluated in a single afternoon". But this is clearly not true, as earlier in the same article he notes that Angela Duckworth has developed a survey for grit that predicts (well) both Westpoint cadet graduation rates and spelling bee performance. Here (pdf) is Duckworth's rating system for grit, including "self-report and informant-report versions of the Grit Scale, which measures trait-level perseverance and passion for long-term goals." So what's the deal?

I suspect grit is shunned as an aptitude test not because it is un-measurable, but because it is game-able. That is, if NFL scouts started judging players on how they rated themselves 1-10 on perseverance and passion, the players would all give themselves 10's on everything, except maybe one or two 9's to maintain some semblance of honesty. With millions of dollars on the line, wouldn't you?

Still, it does seem to me that you could measure grit in an afternoon, if you wanted to. You'd just have to test it when the player doesn't suspect she is being tested.

Tuesday, April 5, 2011

Three Cool Active Ideation Innocentives

1) A Communication Platform to engage the “Hidden Community” of Family Caregivers. Link. Reward: $8,000, Deadline: 5/2/11, Description: There are more than half a billion people taking care of someone elderly at home worldwide and the number is growing. Most of these dedicated "at-home" caregivers are not professionally trained to deal with such things as dementia, personal hygiene, medical conditions and complications. Our investigations lead us to believe that this "Hidden Community" would benefit greatly from educational materials, product information /recommendations and established healthcare techniques. We are looking for a "communication platform" to reach out to these individuals to provide educational information and respond to feedback to meet their needs.

2) Educating About the Importance and Acceptance of Purifying Drinking Water. Link. Reward: $5,000, Deadline: 4/27/11, Description: [This org] strives to bring clean, safe water to people in developing countries. With this Challenge they would like suggestions for addressing one of the biggest problems they encounter in this process – namely, that of educating illiterate populations about the importance of purifying drinking water.

3) Humanitarian Air Drop. Link. Reward: $20,000, Deadline: 5/2/11, Description: Humanitarian food and water drops can only be done over an unpopulated drop zone because there is danger of falling debris to people below. We are looking for an alternative way to drop large amounts of Humanitarian food and water packages from an aircraft into populated areas such that there is no danger of falling objects (i.e. non-food items) causing harm to those on the ground.

If you have any ideas on these, write them up and make money!