Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Friday, July 20, 2012

Statistics Is Like Medicine, Not Software

Stats questions--even when they're pure cut-and-dried homework--require dialog. Medicine might be a better analogy than software: what competent doctor will prescribe a remedy immediately after hearing the patient's complaint? One of our problems is that the [Stack Exchange] mechanism is not ideally suited to the preliminary dialog that needs to go on.
That's from the ever erudite William Huber, in this chat about why the statistics Q&A site has problems that the software Q&A site does not. Some users argue that a high proportion of questions on the stats site should not be answered unless they are disambiguated further.

You might assume that "more answers are better," but answering an ill-posed question adds more noise to the internet. When searching to clarify an ambiguous term, somebody might find that question, read the answer, and end up even more confused. Recall that this is a field already stricken by diametric ideology and short-term incentives.

Here is my previous post on the wisdom of Huber

Friday, April 13, 2012

Schelling Points And Bioinformatics

A lot of what think about when I do bioinformatics is how to set parameters non-arbitrarily. Basically I am looking for Schelling points: round, clear numbers that are easy to justify. The classic case is setting a p-value threshold to 0.05, which has been around for over eighty years and is still going strong, despite the haters. Other examples are setting e-value thresholds to 0.01 and setting Bayes factor thresholds of 10 as the first to indicate "strong" evidence. Like any threshold, these are arbitrary, but following the paradigm of statistics as rhetoric, their staying power make sense insofar as scientists need to be able to resort to standard procedures to settle debates. Anyway, I have no profound point here, I just think it's cool that a seemingly esoteric topic affects what I actually do on a day-to-day basis.


Is It Possible, In Principle, To Do Methodologically Sound Research?

In a paper published 31 years ago, Joseph McGrath argues (html, pdf) that the answer is no. Specifically, he claims that any research design faces two trade-offs: 1) being obtrusive vs unobstrusive (which maps to my terminology as acquiring info vs altering subject), and 2) being generalizable vs context-cognizant (which maps to my terminology as precision vs simplicity).

In his terminology, these trade-offs allow for the optimization of three distinct values (generalizability of samples to populations; precision in measuring variables; and context realism for the participants). Initially, I disagreed with this. To me, intuition suggests that there should be four points which maximize certain qualities when you are considering the intersection of two-trade offs: one in each corner of the 2-d space.

One way to get around this is if you claim that, in the context of this decision (study design), the trade-offs are not independent. For example, it might be very difficult for a design to be both highly generalizable and highly obstrusive.

Below I've drawn an example. Think of the dots as realizations of actually feasible study designs sampled from someone's mental generation process; i.e., they are probably not at the absolute extremes of the theoretical distribution, but with enough realizations, would come close.


I'm not sure that I agree with this exact distribution, and it would need some justification, but it seems like the only way to justify his three-pronged rather than four-pronged set-up.

Tuesday, February 21, 2012

Tuesday Statisticz: Pick Your Poison At Starbucks

Every healthy drink is the same, every unhealthy drink is unhealthy in its own way. That's my conclusion upon analyzing this data set containing nutrition info of Starbucks drinks from the data set aggregator Factual

It turns out that there's an inverse, non-linear relationship between the sodium and sugar content of a drink. Drinks with very little sodium tend to have lots of sugar. These are mainly things like lemonade and iced tea. (Of course, it's controversial whether salt is bad for you, but the prevailing evidence points to yes.) 

normalized to the number of calories, filtered for >50 calories only; code

Surely this is not a universal trade-off, but it seems to indicate that in order to survive as a popular drink, you have to offer consumers something that they'll enjoy. 

Saturday, November 19, 2011

What Does "Statistical" Mean To You?

What exactly do [the authors] mean by a quantum state being “statistically interpretable”?... Basically, [the authors] call something “statistical” if two people, who live in the same universe but have different information, could rationally disagree about it.... As for what “rational” means, all we’ll need to know is that a rational person can never assign a probability of 0 to something that will actually happen. 
To illustrate, suppose a coin is flipped, and you (but not I) get a tip from a reliable source that the coin probably landed heads. Then you and I will describe the coin using different probability distributions, but neither of us will be “wrong” or “irrational”, given the information we have.
That's Scott Aaronson discussingpaper about the nature of quantum states. Googling "define statistical," I see, unsurprisingly, "of or relating to the use of statistics," and then googling "define statistics," I see "the practice or science of collecting and analyzing numerical data in large quantities." 

To me, the large quantities bit emphasizes that the role of statistics is to parse signal from noise, which is only possible with more than two data points (or, to be fair, some assumptions). So, I'd consider the authors' use of the word statistical to be sort of non-standard, because it seems to be able to be used for interpreting just one quantum state. 

Quite possibly this is actually standard use of the word statistical among certain physicists, which would make this yet another example of why you shouldn't assume that terminology is at all consistent across disciplines. 

Sunday, November 6, 2011

A New Solution To A Grid Coloring Challenge

It is here, as explained by Alexandre Thiery. The challenge is to find a four color schema such that a 17 x 17 grid has no rectangle with the same four colors at each corner. The best known solution, shown below, has three rectangles. They are denoted by the black lines.


Who can find a schema with no such rectangles? Does one exist?

####

The fact that I enjoy this so much indicates some sort of bias towards colorful things. Or maybe just pretty things, more generally. 

Monday, May 30, 2011

Which Parts Of Crowds Are Wise?

Peter Freed has written a pretty ambitious critique of Jonah Lehrer's summary of this study (pdf) on the wisdom of the crowds. The crux is that:
But now that I realized he really meant median, and that maybe he didn’t know what median meant.  Because median guesses are not guesses by a crowd, as Lehrer states.  They are guesses by a single person... [Lehrer] is talking about that 0.7% single-person data point: one person, selected after giving their answer, got close to the correct answer on one of six questions.  One person guessed 10,000 when the answer was 10,067.  That’s one hit out of 144 x 6 = 864 attempts.  That seems about right to me, from a common sense perspective. Which is to say, that is a shitty batting average.
Scrolling through the comments, I was pleased to see Ian Sample point out the critique of Freed's critique that I was going to make:
In Wisdom of Crowds studies you can look at the mean and / or the median. The median usually gives the best result if the guesses *do not* follow a normal distribution. The mean, of course, exploits the error-cancelling advantage that WOC is known for, that is, as many people under-estimate as over-estimate the right answer, so averaging cancels all but systematic biases. But to my point. To dismiss the median answer – one guy’s response – misses the fact that without the crowd you have no median answer to dismiss. Without the crowd, you do not know which value to pick. That’s the whole point. The crowd steers you to the median value, which in many cases outperforms the mean.
The median is indeed generated by only one person, but it becomes interesting only in the context of all the other estimates. It is useful here because it offers resistance to outliers. For example, some less numerate soul might have guessed 1,000,000, which is way off from the true value of ~ 10,000, thus skewing the arithmetic mean. In that case you'd much prefer a more robust statistic like the trimmed mean or the median.

In prediction markets, the most recent price of a transaction doesn't always best represent the current beliefs of the market. There's more info if you look at the whole distribution of orders. Similarly, it is unfair of Freed to dismiss the whole data set just because one type of estimator is flawed. This is one of the coolest parts of statistics, using potentially counter-intuitive methods to extract useful info out of data, to find the wisdom in the crowds.

Sunday, May 15, 2011

Fighting The Lernaean Hydra Bias

I'll just only mention the heads I do cut off

In one Greek myth, Hercules takes on the task of killing a serpent-like, many-headed beast. This is made more difficult by the fact that its heads regenerate, so even if Hercules chops one off with his sword, another will simply sprout in its place. John Ioannidis uses this frustrating scenario as an analogy for a problem in the world of scientific publishing in his discussion of meta-analyses (doi:10.1002/jrsm.19).

The example Ioannidis employs to explain this problem is his experience doing a meta-analysis on the pharmacogenetics of certain polymorphisms for asthma treatment (doi:10.1097/01.fpc.0000236332.11304.8f). There were many studies that fit the criteria, but they each evaluated their own endpoints and genetic contrasts. That is, in most of the studies, the vast majority of possible correlations that could have tested with the data between phenotype and genotype were either not done or not reported.

So the surface problem, in so far as this case generalizes to others, is that published studies are not as exhaustive as they could be. But the central, troubling implication is that these studies do not fail to be exhaustive because of time or computational constraints, but because the researchers want to emphasize the usefulness and/or interestingness of their results. This is more insidious--this is why the hydra heads regenerate.

Now, one can use meta-analysis to retrospectively "chop off" findings that are truly insignificant by combining the results of many different data sets. But meta-analysis itself can be biased in many ways (e.g., during study selection), and moreover, later researchers can just come back to the issue and cherry pick more novel associations, thus "sprouting" more statistically significant findings.

When faced with the hydra, Hercules knew he couldn't go it alone, so he called on his nephew for help, who suggested that they cauterize the stumps with fire before the heads could regrow. An analogy to this strategy might be to post warnings on the electronic copy of papers that have been called into question by later studies. Such a warning would be much milder and hopefully less political than a retraction, which typically implies some sort of error. Publishing a potentially informative result that is eventually overturned is still laudable.

But instead of this type of patchwork fix, a more fundamental approach seems more fruitful. In the original myth, only one of the hydra's heads was truly immortal, and this was the one that Hercules needed to chop off to finally defeat the beast. The immortal head of the scientific publishing hydra is the incentive structure pushing researchers towards significance hunting in the first place.

Reworking these incentives is what Ioannidis is fundamentally arguing for, as the way to kill the Lernaean hydra bias once and for all: more standardization, more consortia, and more of a push towards openness and replicability. Every study might combine previous data with its own for estimating the posterior probability of the parameters it is examining, and all research might be seen as a continuous and cumulative meta-analysis. Maybe one day.

(photo credit to Frank Rafik)

Sunday, April 17, 2011

The Wisdom Of Whuber

That's William Huber, whuber for short, dispensed in his answers at the relatively new stats Q&A site, Cross Validated. His answers are the best on there, reputation normalized to the number of answers (with shrinkage). Here he writes about whether the median is a better summary stat than the mean:
Statistics does not provide a good answer to this question, IMO. A mean is ok to use, too, and is relevant in mortality studies for example. But ages are not as easy to measure as you might think: older people, illiterate people, and people in some third-world countries tend to round their ages to a multiple of 5 or 10, for instance. The median is more resistant to such errors than the mean....  Thus, for demographic, not statistical, reasons, a median appears more worthy of the role of an omnibus value for summarizing the ages of relatively large populations of people.
Here he writes about the biggest questions in statistics, from which I'll reproduce two (emphasis his):
  • Coping with scientific publication bias. Negative results are published much less simply because they just don't attain a magic p-value. All branches of science need to find better ways to bring scientifically important, not just statistically significant, results to light. (The multiple comparisons problem and coping with high-dimensional data are subcategories of this problem.)
  • Probing the limits of statistical methods and their interfaces with machine learning and machine cognition. Inevitable advances in computing technology will make true AI accessible in our lifetimes. How are we going to program artificial brains? What role might statistical thinking and statistical learning have in creating these advances? How can statisticians help in thinking about artificial cognition, artificial learning, in exploring their limitations, and making advances?
And here he writes about whether you should use a normal distribution to assign student grades:
I think that if any of those 800 students were to read this question, they might be offended. How well did they perform? How much learning was accomplished? That is what a grade should reflect, not some arbitrary statistical summary of their position in a group. IMHO this question should be recast in terms of teaching objectives, not statistical procedure, such as "what is a good way to convert raw scores to grades in a way that respects student accomplishments and advances the learning objectives of this class?" Statistics can help, but blind statistics--like standardization--will not.
Although they are often quite quantitative, his answers show how good stats rely on far more than just math. 

Friday, November 19, 2010

P-Value Polemics

As I am always up for a good scholarly debate, I was quite pleased, after reading this '05 article calling for a replacement to the p-value called p-rep (cited 200+ times), to see a somewhat vitriolic '09 rebuttal (pdf). First, the abstract of the '05 paper by P. Killeen:
"The statistic Prep estimates the probability of replicating an effect. It captures traditional publication criteria for signal-to-noise ratio, while avoiding parametric inference and the resulting Bayesian dilemma. In concert with effect size and replication intervals, Prep provides all of the information now used in evaluating research, while avoiding many of the pitfalls of traditional statistical inference."
A rather bold claim! And, shortly after its publication, the journal Psychological Science (6th highest psyc impact factor) recommended that authors report p-rep instead of the traditional p-value. Which makes the rebuttal article by Iverson et al that much more tantalizing. They write:
"This probability of replication prep seems new, exciting, and extremely useful. Despite appearances however prep is misnamed, commonly miscalculated even by its progenitors, misapplied outside a common but otherwise very narrow scope, and its seductively large values can be seriously misleading. In short, Psychological Science has bet on the wrong horse, and nothing but mischief will follow from its continued promotion of prep as a scientifically informative predictive probability of replicability."
Now that is what I call a take down! These same authors calm down quite a bit in their '10 article and even make the level-headed suggestion that p-rep is a step in the right direction, but that is uncool so I won't quote from it.

####

Reading about p-values makes me want to start a blog about them (how does such a blog not already exist?!). A good subtitle could be "where one in every twenty posts will be significant by chance alone."

Wednesday, August 25, 2010

Mark Cuban's Non-Probabilistic Thinking

His investment advice today is to pay off high interest debt, save your money in cash, and try to cut personal spending. Fair enough. But then he makes the outrageous claim that "If you have under 100k dollars in liquid assets, your net worth will be higher in one year if you follow this advice than if you follow ANY other investment advice any broker or banker will give you this year."

The likelihood of this claim proving true is vanishingly small. Out of all of the other pieces of investment advice proffered, surely some of these will beat the null strategy of playing it safe. Now, Cuban might argue that you can't identify which advice will allow you to beat the null a priori, and so you're better off not trying, but that's a totally different claim.

Bottom line: Cuban's blog gets demoted from "medium" to "low" priority on the Google reader hierarchy, and is now teetering on the edge of unsubscribe territory.

Tuesday, July 20, 2010

Is Chris Nolan The Best Director Of All Time?

The short answer is, quite possibly. 

For the longer answer, you'll have to indulge me with a bit of stats. You see, there's this website called imdb, (you may have heard of it), and one interesting fact about it is that it has the largest depository of user-generated ratings in recorded history.

So, after aggregating all of the movies that somebody has directed, it is easy to calculate his average rating on imdb. In order to see whether Chris Nolan really is the best director of all time, I did this for anyone who had seemed to have a reasonable chance of winning.

To be fair, I didn't count early movies that the director probably didn't have much funding for, movies that he wasn't the main director of, and documentaries, because they tend to be uncool.

Once I did this, I realized that I had meandered into a dilemma. You see, the best directors had only directed one movie each! At the tops of the list were Florian Henckel von Donnersmarck, who directed The Lives of Others (an 8.5), and Tony Kaye, who directed American History X (an 8.6).

To appropriately punish these slackers for their limited sample sizes, I had an excuse to employ a fancy Bayesian estimator. This sounds much more complicated than it is.

Basically, I calculated the total number of movies each person had directed, inputted the average rating of a movie on imdb (6.9), and set an arbitrary variable, m, to be some value between like 0.001 and 1000. Then, I put each director's average rating and total number of movies through this equation, and viola: it spit out rankings that took into account the fact that von Donnersmarck and Kaye had only directed one movie each.

Now, determining which value to use for the variable m is an open and interesting question. It depends on your values: do you prefer a director that has made a whole lot of good movies, or one who has made just a few great movies? It'd be hard to answer this objectively.

If you prefer quality over quantity, then you should set your m low, so you don't punish low sample sizes as much. If you think that a director has to be somewhat prolific to be even included in the discussion, then you should set your m high. I set m to three different values to be fair to each of these reasonable positions.

When m = 20, the top 5 directors are:

1) Akira Kurosawa, score = 7.40 (weighted), directed 25 movies. Highlights: Seven Samurai, Yojimbo.
2) Stanley Kubrick, score = 7.36, directed 11 movies. Highlights: Paths of Glory, Dr. Strangelove.
3) William Wyler, score = 7.34, directed 26 movies. Highlights: Dodsworth, Ben-Hur.
4) Ingmar Bergmann, score = 7.34, directed 30 movies. Highlights: The Seventh Seal, Wild Strawberries.
5) Luis Buñuel, score = 7.29, directed 32 movies. Highlights: Viridiana, The Discreet Charm of the Bourgeoisie.

When m = 10, the top 5 directors are:

1) Stanley Kubrick, score = 7.58 (weighted).
2) Akira Kurosawa, score = 7.54.
3) Chris Nolan, score = 7.50, directed 7 movies. Highlights: Memento, The Prestige.
4) William Wyler, score = 7.46.
5) Hayao Miyazaki, score = 7.45, directed 9 movies. Highlights: Spirited Away, Princess Mononoke.

When m = 3, the top 5 directors are:

1) Chris Nolan, score = 7.93 (weighted).
2) Stanley Kubrick, score = 7.92.
3) Sergio Leone, score = 7.88, directed 6 movies. Highlights: The Good, The Bad, and The Ugly, Once Upon a Time in America.
4) Quentin Tarantino, score = 7.80, directed 7 movies. Highlights: Pulp Fiction, Inglourious Basterds.
5) Hayao Miyazaki, score = 7.78.

Another sort of difficult thing to choose is how to count Pixar's movies. Most of the movies list different directors, but really, who actually knows what goes on in that forsaken place? If you consider the Pixar movie making team as its own distinct entity, then that entity would end up at 5th, 3rd, and 5th on the above lists.

If you'd like to check my raw data, feel free to peruse this google document at your leisure.

So there you have it. Chris Nolan is the best director of all time... under certain assumptions. Finally, implicit in this post is the recommendation that if you haven't seen Nolan's Inception yet, you need to get your act together.

Tuesday, May 4, 2010

Alcohol And Vocab Aptitude

Vocabulary, which seems to be a pretty good proxy for general intelligence, shows a positive and dose-dependent correlation with being an alcohol drinker, among Americans: 

Woah. Razib first found this relationship (here; HT R Wiblin) and we both used UC Berkeley's awesome SDA to do the crunching. The error bars are 95% confidence intervals, and their non-overlap between groups means most of these differences are very unlikely (< 5%) to be due to random chance. The general trend holds for all kinds of different age groups (18-30, 30-50, 50-60, 60-70, 70-100, etc.).

Allow me a couple stabs in the dark as to the relationship here:

1) Alcohol reduces anxiety (see here) and, as Steven Pinker has speculated, "people with higher intelligence are better at overcoming their anxious temperament and more likely to see their own worry list as a problem to be solved, minimizing unnecessary anxiety while still being anxious enough to get things done.” So, people with higher intelligence may be more likely to be self-medicating their anxiety with alcohol.

2) People with higher vocabs are more likely to be more highly educated, and thus been introduced to the drinking culture that is commonplace in institutes of higher ed. It may be a part of the culture there because it is more impressive (see bottom here) to be able to succeed in school and party on the weekends. And most people seek to maximize their relative impressiveness.

Thursday, July 2, 2009

Human Number Generating Flaws

Chris Lloyd reviews some evidence that the Iranian elections may have been fraudulent. Here is one of his surprising points:
Humans are especially terrible at generating random numbers. And for a large voting count, for instance 325911 which was Ahmadinejad’s count in the region of Ardabil, the last few digits should be essentially random. On the other hand, if someone were making the numbers up and not concentrating too hard on the unimportant final digits, you might expect to see some tell-tale signs of non-randomness in the those final digits.

This idea is due to Alexandra Scacco and Bernd Baber who have suggested that there is indeed such evidence in the data. They claim that human generated random numbers tend to have too many 7’s and not enough 5’s. And looking at pairs of digits, they claim that human generated digits will have too many adjacent sequences such as 23 and 76.
Another cool numerical phenomenon is Benford's Law, which is that in lists of numbers from real life sources of data, the leading digit is the number one almost one-third of the time. Detecting crime and fraud in a quantifiable manner is really sweet.

Thursday, June 25, 2009

The Non-Decline of the Dark Knight

Last July I predicted that The Dark Knight would drop from #1 on imdb to somewhere between #s 15 and 22, although I conveniently didn't mention a specific time scale. Nevertheless, a year or so is a reasonable amount of time to consider, and since I am sometimes a reasonable person I will admit that yes, I was off. It remains at #7 today. Here are some charts with data from its relatively small decline over the past year. Most of my data points are from the first few weeks, because that was when it was most exciting.


I'm missing a bunch of data points past the first few weeks but you can see in this last one that the chart fits a power function well. The decay will stabilize at equilibrium at some point--I wouldn't expect the movie's rating to drop foreover.

The Dark Knight has set the bar for imdb chart domination in the modern era. Despite its impressive staying power, it, like every movie before it and like every movie after it, has dropped. This is the rule for imdb ratings. The question is whether the imdb rater's tendency to favor new stuff is specific to the movie industry or whether it generalizes to other domains as well. Too bad there are so few quantifiable rating systems.

Saturday, April 25, 2009

Fighting Confirmation Bias in Fish Oil

Ben Goldacre covers both one null result of fish oil's efficacy in children and explains how subgroup analysis in statistics can be used to mine for positive results when there are none in the main sample. He notes that in 1973, Lee et al randomly assigned patients to non-existent treatment groups. They were able to find a subgroup, characterized by odd disorders ("three-vennel disease" and "abnormal left ventricular contraction"), where Treatment 1 had a significantly higher survival rate than Treatment 2. So be wary when researchers report statistical significance in subgroups only, unless it is clearly a biologically relevant subgroup and/or the researchers explicitly hypothesize the differences in subgroups before the trial. I do have one question though. Can't you control for this selection bias with an ANOVA for main effects?

Wednesday, April 15, 2009

Vanishing Employment Since January '07

Slate has a cool interactive graph chronicling the job loss in local regions for the last 26 months. Click "play" and watch it progress on a month by month basis. It's like watching a zombie movie where people are initially infected one by one but the situation soon starts spiraling out of control. To my naked eye Texas seems to be relatively immune to the disease so far, but if we follow the zombie movie analogy then the state is merely due for an extremely gruesome death. Where is Bruce Campbell when we need him?

(HT Razib)

Friday, March 6, 2009

Why have free throw percentages remained constant?

Free throw percentages in the NBA over the past 50 years have remained remarkably constant, hovering at around 75 percent. Since most other tangible facets of the game have improved in that time span, it is puzzling that there has not been much change in free throw percentages.

My explanation of this phenomenon is that there has not been much selection pressure towards better free throw shooters. Most NBA teams emphasize size and athleticism, arguing that if you are outmatched in those areas you will be so dominated in other areas that free throw shooting will be irrelevant. Moreover, free throws are seen as something that you can teach, but you can't teach height, and you can't teach tomahawk dunks.

Finally, even if NBA teams do have a preference for slightly better free throw shooters, that effect will be counteracted by concurrent preference for slightly bigger players. Big men will always have a tougher time shooting free throws. This is partly due to simple biomechanics--image trying to throw a tennis ball into the hoop and you understand Shaquille O'Neal's daily struggle. It is also partly due to reduced incentives for those 7 feet and over. Even if Andrew Bynum cannot shoot free throws well, he will still have a job somewhere in the NBA because of his abnormal size, whereas a 6'2" player would not be able to survive with a 60% average.

Assuming the inflation-adjusted wage stays relatively constant or increases, and as the global talent pool increases, this model predicts that either players will continue to get taller and stronger or average free throw shooting will improve. The 2006-2007 NBA average height was 6'6.9" and the average free throw percentage was 75.2%. I am willing to bet that, assuming the inflation-adjusted wage of an NBA player stays the same and the league has not expanded drastically, one or both of those measures will have increased by 2026.

###

Seth Roberts reports that Cal Tech's had one of the top basketball teams in the nation during the 1950s! As Seth notes, in the 1950s you would look at that and think "well that's just how it should be," but now it looks so weird.

Tuesday, February 17, 2009

Predicting the 2009 Oscar Winners

I have lots of problems with the Best Picture Nominees this year, including the non-nomination of by far the most popular movie of the year, and the rampant recency bias. But it's still fun to predict the winners. Here are the IMdB votes broken down by demographic:


Last year I found that although There Will Be Blood had a higher overall rating than No Country for Old Men, No Country had higher ratings in two key categories: Males over 45 and Top 1000 Voters. These correspond roughly to what pretentious old men like, which is basically the same people who vote on the Oscars.

Although Slumdog Millionaire is far and away the favorite to win the Oscar, with over an 87% chance currently at InTrade. It is however, only tied with Frost/Nixon in the key two metrics of Top 1000 Voters and Males Aged 45+ at 8.2 and 7.4 for Slumdog and 8.1 and 7.5 for Frost/Nixon, respectively. The Reader has them both beat at 8.2 for Males 45+ and 7.5 for Top 1000 voters.

Since it would be such an upset, I feel justified in making the prediction that either Frost/Nixon or The Reader will win Best Picture, upsetting Slumdog Millionaire. You heard it here first.

Tuesday, January 20, 2009

Benchmarks for Obama's Presidential Success

"Remember Red, hope is a good thing, maybe the best of things, and no good thing ever dies." - The Shawshank Redemption

Lower the National Debt
: There are many ways that Obama could accomplish this: Cutting back on "wasteful" spending, ending the war in Iraq, slashing entitlements, ending the war on drugs, whatever. This is a blanket category for simply lowering the national debt as a percentage of GDP. At the end of the third quarter of 2008 (October 1), the national debt was $10,024,724,896,000 (from here), while the GDP was $14,412,800,000,000 (from here), for a percentage of debt of 69.6%. Assuming that the debt of other countries has stayed constant since 2007, that would put us at the 18th highest level in the world (based on this). Here are some possible futures:

Debt as a percentage of GDP above 75% = F
Debt as a percentage of GDP less than 75% = D
Debt as a percentage of GDP less than 69.6% = C
Debt as a percentage of GDP less than 65% = B
Debt as a percentage of GDP less than 60% = A

Slash Per Capita Healthcare Spending so as to be Comparable to OECD Countries: This should be done while maintaining health outcomes as comparable to these countries. Here is the most recent data (taken from here, OECD countries defined here, explanations on data techniques here):

The US is currently the worst in the world at this category with comparable countries. The median on this list is currently $3187 per capita. In 2005 (the most recent statistics), our value was 200% of this median. If after Obama's presidency that number is,

250% or greater = F
220% or greater = D
180% or greater = C
150% or greater = B
Less than 150% = A

If there is an accompanying dramatic reduction or improvement in health outcomes, which I would not expect, then I will use my discretion to factor that in. I expect that the WHO will collect this data sometime near 2012.

If Obama's GPA is a 2.5 or above, I will consider him a "good" president, if he gets a 3.5 or above I will consider him a "great" president. This post was inspired by the inimitable Robin Hanson's proposal here. Check back in four years from today for the results.