Showing posts with label Rating Systems. Show all posts
Showing posts with label Rating Systems. Show all posts

Tuesday, July 24, 2012

Book Review: Great Flicks By Dean Simonton

Attention conservation notice: Review and notes from a book discussing an academic topic that will likely only interest you insofar as it generalizes to other topics, unless you are both a huge stats and film nerd.

I'm fascinated by movie ratings and what they tell us about: 1) the best ways to use rigorous methods to study the quality of a subjective output, 2) how variable people's assessment of quality are, and 3) how people conceptualize their own opinion in the context of everyone else's. Dean Simonton is a giant in the psychology of creativity, and I loved his book Creativity in Science. So, as soon as I saw this one, I clicked "buy it now" on its Amazon page

My typical gripe against academic investigations of movie ratings is that they discount imdb.com, a huge resource with millions of data points, segregated by age, gender, geographical location, on an incredibly rich array of movies. So, soon after buying the book, I searched in the Kindle app for "imdb" and found very few results. This predisposed me to disliking it. 

A few of my other gripes: 

1) It takes awhile to get used to Simonton's academic writing style. 

2) The book takes few risks stylistically. Each chapter feels like it could be its own separate article. Thus, he does not take full advantage of the long-form medium. 

3) When he discussed a few of the measures (such as the correlation between different award shows), I felt that there was some issues with his account of the causality. Surely there is some, non-negligible probability that people take the ratings of others' into account when they make their own judgments. He mentions this sometimes, but not enough for me, and ideally he'd come up with some creative way to try to get around it.

4) Finally, there are a few typos. I actually like seeing typos, because it makes me think that I am learning from a more niche source that others are less likely to appreciate, but YMMV.

By midway through the book, Simonton had won me back to a large extent. His analyses of his data were very well-done and he supplies tables so you can look at the regression coefficients yourself. And there are many good nuggets, such as: 

- the best predictors of higher ratings are awards for better stories (e.g., best screenplay and best director), as opposed to visual or musical awards
- having individuals on the production team who play multiple roles (such as writer, cinematographer, and editor all at once) makes the film more successful, presumably due to creative freedom 
- some amount of repeated collaboration over multiple films with the same individuals, but not too much, is optimal for winning awards (i.e., there is a trade-off between stimulation and stagnation) 
- higher box office returns are inversely correlated with success at awards shows
- the typical film is unprofitable; "about 80% of Hollywood's entire profit can be credited to just a little over 6% of the movies"  
- the curse of the superstar: "if a star is paid the expected increase in revenue associated with his or her performance in a movie then the movie will almost always lose money" (this is because revenue is so positively skewed) 
- divides movies into two types: those that are extremely successful commercially, and those that are extremely successful artistically (people often use the former to subsidize the latter) 
- negative critic reviews have a more detrimental impact than positive reviews have a boosting effect on box office returns
- on ratings, critics and consumers have similar tastes, although consumers' tastes are more difficult to predict, presumably because their proclivities are more diverse
- for a consumer, the most important factor for whether they will watch a movie is its genre (#2 is word of mouth) 
- dramas do worse in the box office, better at the awards shows; comedies are the reverse
- PG-13 movies make the most money; some romance, but no actual nudity, is best (and lots of action but no gore) 
- on average, sequels do far worse in ratings and awards than the original movies
- the greater the involvement of the author in an adapted movie, the less money it will make (they interfere more and might care more about "artistic integrity" than making money) 
- directors tend to peak in their late 30s; they have more success in their late 20s than their late 50s, on average
- divides directors into two types: conceptual (innovative and imaginative; think Welles) and experimental (technical and exacting; think Hitchcock)
- conceptual directors express ideas through visual imagery and emotions, often leave behind one defining film, and decline quickly
- experimental directors emphasize more realistic themes, slowly improve their methods, and their best films often occur towards (but almost never *at*) the end of their careers
- female actors make less money than their male counterpoints, and the best picture award correlates much better with best male actor than best female actor
- awards for scores are much better predictors of a film's quality than awards for songs

All in all, this book is far from perfect, but it is likely the best full-length treatment of quantitative movie ratings available. If you are interested in the topic, and occasionally find yourself doing things like browsing the rating histograms on imdb, then this is essential reading. 

Saturday, March 31, 2012

The Valiant Never Taste Death But Once

After reading this interesting excerpted article from Dick Teresi's book The Undead, which discusses the difficulties in defining death by a single, consistent set of criteria and the social qualms that stirs, I decided to check out the Amazon reviews. The associated ratings were (and still are) quite shockingly bad! They follow the classic "so bad it's good" distribution, with 5 5-star ratings, 1 3-star rating, and 33 1-star ratings. So, given that I am always up for a good controversy, I decided to read and review it myself. Ultimately I mostly side with the critics, giving it two stars. If you are interested in the subject matter, I'd suggest instead Kenneth Iserson's Death to Dust, which is a bit older but much more level-headed and thorough treatment of similar issues. 

Sunday, March 25, 2012

Comp Exams For Each Course

The solution I propose is comprehensive exams at the end of each course, much like Advanced Placement exams, that thoroughly and objectively distinguish students on merit alone. The emphasis in each classroom would then shift from fighting the teacher for high grades to cooperating with the teacher to learn the material necessary to perform on the exam.
That's from Andrew Knight, in an essay discussing problems that will not be new to anyone who is or has recently been in school; more here. This is exactly what I wanted during most of my science and math courses. The alternative is to place a greater emphasis on big standardized tests like the SAT, but there can be so much variability in results from just one day.

One question is whether such exams could be a part of classes that are less fact-based, such as history and english. There is actually a machine learning competition for automated essay grading going on right now. I don't pretend to know the answer to this question, but even if it is currently infeasible, that shouldn't stop the tests from being used in math, science, and foreign language classes. 

Wednesday, December 7, 2011

Reviewing Newt's Reviews


That is a histogram of his Amazon book ratings, and here is the source (HT TC). He is quite the pushover, and the above is a classic example of why a good rating system must rate the raters or suffer from bias. What about the content? Two trends stick out:

1) He is highly positive. Among his most common adjectives are "masterful," "remarkable," and "brilliant." Even when he gives a book four stars, he rarely says anything negative, and in fact it's usually not clear why books didn't get the top score of five.

2) He is highly technical. By this I mean that the majority of his sentences are devoted to strict summary rather than analysis. This makes sense, as it is probably smarter for a politician to say something obviously factual (and thus unassailable) than to take a risk.

Monday, September 26, 2011

The Value Of Thinking About Rating Systems

A new project I have just started is going to generate personalized movie ratings for users. The way it works is as follows. You rate the movies you have seen. Then the system finds other users with similar tastes to extrapolate how much the you will like some other movies. It is currently written entirely in Python.
That's from Sergey Brin's 1996 resume. Prior, of course, to co-founding Google.

Correlation, causation, or aberration? You tell me. 

Saturday, August 13, 2011

Searching For The Imdb Of Books, Part II

As watching imdb's top 250 most highly rated movies has proven to be such a smashing success, I have long yearned to find (and fleetingly, to develop) a similarly authoritative list for fiction books. The keys for a good list are: 1) a large sample size, 2) shrinkage estimation of ratings to the average, 3) a continuous scale (the more levels, the better, but yes we'll often have to settle for five stars), 4) defenses against gaming, and 5) a wide index of titles. To the best of my knowledge no site fulfills all of these requirements. These are the current contenders:

Amazon ReviewsUpside: They have a huge incentive to index all available books and are proficient at combining ratings across different editions of the same text. They also have a useful "was this review helpful to you?" tool which could eventually be employed to rate the raters and thus weight the overall ratings. Downsides: Their insistence on showing the average rating in half-star increments (typically 5, 4.5, 4, or 3.5) means that it involves manual calculation to distinguish between the two radically different scores of 4.24 and 3.76. I also often don't trust the resistance of their ratings to gaming. But most damningly, there is simply no attempt to create a good list of the most highly rated fiction books. Filtering by "highest average rating" in "literature and fiction", their #7 best fiction book of all time is currently Jim Gorant's The Lost Dogs: Michael Vick's Dogs and Their Tale of Rescue and Redemption, which is probably a fine book, but I think the author would be insulted to hear that it was considered fiction, and I think more than three-quarters of the english profs across the country would be insulted to hear it called literature.  

One-Time Votes: By this I refer to ad-hock competitions of various websites which ask users to vote on their favorite books. There are many of these strewn across the web, for example, check out NPR's top 100 science fiction and fantasy books, or Modern Library's top 100 novelsUpside: These tend to get large sample sizes (NPR had >60,000 votes), which makes them more accurate and harder to game. Downside: The process is not iterative and requires manual input to update, so they won't last or scale. More troublingly, many (such as NPR's) only allow the option to select one's favorite books, without voting others down, which unfairly favors books with high variance as opposed to just high average quality. 

Google Books: The site aggregates ratings from elsewhere on the web, including major vendors and online "bookshelves." Upside: Transparent code, takes ratings from diverse sources, and has a clean layout. Downside: Like Amazon, also displays ratings in half-star increments (et tu, google?). But their biggest problem is that different editions of books are stored in different locations and the ratings are not aggregated across editions. See, for instance, the first four results of a search for "pride and prejudice" (here, here, here, and here). Now, even if they did manage to output one total score per novel, it still doesn't seem very google-like to actually curate such a list themselves. But in that case, it wouldn't be hard for someone else to scrape the ratings and convert them into a ranked list. 

Library Thing: Upside: They have scale, with over 10 million ratings, and they already have some pretty cool statistics (check out the most "connected" people--Napolean is #1). They also do have a top 25 books by ratingDownside: They need to split the rankings for non-fiction and fiction. At this point I've given up on searching for a canonical non-fiction ranked list, as those ratings are so context-dependent and world-view driven. And they need to do a better job of categorizing in general. For example, the movie for LoTR:Two Towers, while an awesome movie and in imdb's top 250, should not be among the highest rated 25 books. More importantly, the editors of the site have not implemented a rating system that punishes books with fewer ratings. Instead, books simply need a minimum total of 20 ratings to make the list. This is bothersome, but easily improved, as the editors could simply study and implement the imdb method

Good Reads: Upside: As far as I can tell, this is the largest "bookshelf" site with the most user ratings. Huge potential. Downside: They've made no attempt to publish a list of the highest rated books across the site! All I can ask is, what is holding you back, GoodReads editors? Qualms about alienating authors whose works won't make the list? Fears of being labelled imperialistic? These are both hogwash. Our time is scarce and in order to be informed consumers we need to know what the best books are. If you are worried about the arbitrariness of the minimum votes cut-off, then publish multiple lists with different scaling parameters. You will thank me later when the list gets out-of-control traffic. Indeed, a group of passionate GoodReads users recently called for such a list. To this valiant effort I can only say, Viva la RĂ©sistance!

Saturday, July 16, 2011

Bill Simmons' Rating Nihilism

Earlier he claimed, on scarce evidence, that "Rotten Tomatoes scares me as a metric," because "people are idiots," and that "their 'top critics' rating is much more useful."[1]

Now he explains that:
I believe Michael Jordan is the greatest basketball player ever, and I can prove it. I believe The Breaks of the Game is the greatest sports book ever, but I can't prove it. Books can't be measured that way — they hit everyone differently, so when we're evaluating them, we can only say, "You can't mention the greatest books (or albums, paintings, TV shows, movies or whatever) without mentioning that one." That's as far as you can go.
I don't get it. He thinks you can sort of rate movies (if you trust only the experts), but you can never rate books except for saying which ones are "among the best"? This is inconsistent.

He's right that there is a key difference between sports and film/writing, although it is not, as he claims, that the latter "hit[s] everyone differently." The difference is that in sports there is a known goal--for the team to win. That means that, at least theoretically, it is possible to tease out which player stats tend to correlate with winning, and then use those stats to evaluate players.

But notice the causality here. We can't evaluate individual players well until we know which stats are generally good indicators that a player will help eir team win. Intuition does not necessarily serve well here, an insight upon which books have been written and careers have been made.

In film/writing there is no such clear objective, and thus the ratings by individuals who have seen/read them must be subjective. So instead of evaluating statistics based on how they correlate with the objective of winning, we must instead evaluate rating systems based on inter-rater reliability. The goal is that if you added more independent ratings by unbiased raters, there should be as small of a deviation as possible between the new and old ratings.

The obvious suggestion is that, if we want better opinions, we need more of them to average out more of our random biases, like how hungry we were when we first saw the movie.

Again, the input must be subjective. But once we've decided upon the best rating system, its output is objectively our best estimate of that film/book's quality. Just as in sports, personal intuition is not the best estimate of quality, and to believe otherwise is simply hubris.

Perhaps it should not surprise us that a key opinion maker is arguing that we should only trust key opinion makers, instead of wide-scale opinion aggregators. But the rest of us don't have to buy it.

####

[1]: Rotten Tomatoes ratings have many problems, like the fact that they threshold scores into "good" and "bad" and count the percentage of each instead of employing a continuous scale. But that is a straw man for the claim that open, aggregated movie ratings are bad.


Monday, April 18, 2011

Milestones In...

I've recently discovered Nature's milestones index, which links to timelines of the major advances in the research of many fields: light microscopy, gene expression, development, etc. These were chosen by panels of many experts. For example, these 40 helped decide the milestones in cancer research. The timelines have links that explain why each milestone was important, like this one on the first methods of DNA sequencing. Awesome.

I wonder if there's some way that we could allow people to vote on these milestones in a similar way that others have set up for people to vote on milestones in computer science? If so, we could tap into what seems to me like the most productive form of crowdsourcing, where experts define the field, and then the masses rank the entries in that field.

Monday, April 11, 2011

Ranking Ideas In Science

Last summer I bought and read The 100 Most Important Science Ideas after noticing it in a bookstore (my first mistake--I should have checked the ratings online first). I learned a fair amount from it, but I have to say it fares miserably in its attempt to actually rank ideas in science. First, it only covers three subjects: genetics, physics, and math. Second, even within those subjects, the topics are listed merely by date of discovery, not importance. Finally, there was little to no space devoted to methodology.

Subsequent attempts to find lists of the most important science ideas, via google searches and cold e-mails to potentially knowledgeable people, have also left me empty-handed. Uncool.

A good system to rank science ideas, both historically and as they are published, would be so money. The historical list would be really useful for educating the next generations and as outreach to the public. And dynamic, post-peer review ratings would help researchers use their precious time reading the best papers, instead of relying solely on the impact factor of the journal.

Given the above, you can imagine my immense pleasure to see Scott Aaronson's announcement today of a site that allows anyone to vote on milestones in computer science.

There are at least a couple of ways this voting could be done. The first way, as they currently have the site set up, is that users can pick and choose to vote any individual idea on the list up or down. The advantage of this is that users can choose to vote only on the ideas that they actually know something about.

The second way is that the site could present two options to users, the users would choose which of those two are better, and then an algorithm would use those preferences to rank all the ideas. The advantage of this is that it's more fun. Indeed, you might recall that a similar system was employed by the young Mark Zuckerberg in facemash. Wait, you haven't seen The Social Network? C'mon now, it's #190 on the top 250. Step your game up.

Anyway, bravo to Jason, Ammar, and Scott. Now we just need to create similar lists for all other scientific disciplines, incentivize people to vote on them, and aggregate the results. We'll also def need some kind of normalization to account for the fact that computational pursuits will have at least 10x the votes, because those people are on their computers like all day.

Saturday, April 9, 2011

When Can We Measure Grit?

Jonah Lehrer's interesting, 15,000 character article about measuring NFL quarterbacks concludes by saying that we have neglected grit in favor of IQ because "grit can't be evaluated in a single afternoon". But this is clearly not true, as earlier in the same article he notes that Angela Duckworth has developed a survey for grit that predicts (well) both Westpoint cadet graduation rates and spelling bee performance. Here (pdf) is Duckworth's rating system for grit, including "self-report and informant-report versions of the Grit Scale, which measures trait-level perseverance and passion for long-term goals." So what's the deal?

I suspect grit is shunned as an aptitude test not because it is un-measurable, but because it is game-able. That is, if NFL scouts started judging players on how they rated themselves 1-10 on perseverance and passion, the players would all give themselves 10's on everything, except maybe one or two 9's to maintain some semblance of honesty. With millions of dollars on the line, wouldn't you?

Still, it does seem to me that you could measure grit in an afternoon, if you wanted to. You'd just have to test it when the player doesn't suspect she is being tested.

Thursday, February 10, 2011

So Bad It's Good, On imdb

My roommate was recently trying to convince me to watch The Room, employing the argument that, although it is poorly rated at a 3.2, it is "so bad it's good." The question is, can we quantify this?

The intuitive way seems to be to look at the distribution of scores, and indeed that is what Tomasz Węgrzanowski has suggested. The idea is that your typical movie tends to have a single peak at its mode, and the percentage of votes will drop off monotonically on both sides of that peak. Movies that are "so bad it's good," on the other hand, will have a bimodal distribution, with lots of high scores (9's and 10's) and lots of low scores (1's and 2's), and few in-between.

For example, The Hunt For Red October is a fairly standard movie, and you can see that it does show a bell-shaped trend, albeit with a ceiling effect:


So what does The Room's rating distribution look like? Frighteningly bimodal:


Eventually we started watching the movie. It is truly disgustingly bad, and in fact it's hard to even call it a movie, as the plot seems to be just a cheap excuse for softcore porn. Infamously, they seem to use the same sex scene twice, although the director vehemently denies this. Whatever. It's hard for me to evaluate whether the movie is "so bad it's good," but the above ratings speak for themselves.

Bottom line: It's a better sign if a movie is rated higher rather than lower. But, holding average rating constant, you should prefer movies with a wider distribution of votes, as they will tend to be more interesting.

Saturday, January 29, 2011

Towards A More Risk Loving imdb

In my view, the "true" rating of a movie is what the average opinion would be if everybody who watches it:
  • has some basic knowledge of art and human affairs in general (for example, a working knowledge of the canon of classics), but no knowledge about the movie in particular (no previews, hype, etc);
  • is watching alone, so nobody in the theater is laughing at unfunny times;
  • experiences neither hunger, thirst, polyuria, tiredness, stress, excessive marijuana-induced paranoia, nor any inclination to rub tongues with the person next to em;
  • watches on a ridiculously large screen with speaker cables made of pure silver.
Since the above conditions will never all hold, we will all be biased in one or another way when we watch a movie. So we must be wary of the opinion of any given rater. That is, if the first person to watch a movie gives it a 10/10, our estimate for the true rating of a movie should not be a 10, but instead should be adjusted down towards the average.

The above is all obvious. What's less obvious is that there's no easy way to decide how much one should scale down the rating. That depends on how much of a risk you're willing to take with low sample sizes.

Consider The Passion of Joan of Ark, whose 8.3 rating should be enough to place fairly high on the top 250. For example, Sin City also has an 8.3 and it's currently #104. However, TPoJoA only ranks #210, because its paltry 11k votes push down its score so much.

Here's what I'm proposing. Let us, the users, choose our own scaling parameter. Let us define how much of a risk we want to take in trusting smaller sample sizes. Let us choose our own destinies. Because the current system smacks of hegemony, and I, for one, will only stand for it because I have better things to do.

Wednesday, December 29, 2010

Hindsight Is 2010

The NYTM's Year in Ideas is consistently good. My favorites for this edition were D.I.Y. Macro, Performance Enhancing Shoes, Relaxation Drinks, and The 2000's Were a Great Decade. Inspired by their ideas, my retrospective for the year will consist of twelve articles / blog posts, one for each month, that seem especially representative of the year in ideas.

January: "Lessons from a pandemic," Nature editorial, 685 words. The H1N1 virus ended up not being that lethal, but it could have been, and this article highlights the lessons. Among them is that six months is too long of a time for vaccine production, given that viruses now easily spread around the world in "a matter of weeks."

February: "The biomechanics of barefoot running," editor's summary, 283 words. Running on one's toes is healthier than running on one's heels, even though most running shoes promote the latter. This finding is largely academic for me personally, as over the years I have come to loathe jogging. Nevertheless, it is emblematic of the larger "back to nature" craze that has taken over in 2010. This includes the paleo diet, probiotics, and restroom posture designed to prevent hemorrhoids.

March: "Snake oil? The scientific evidence for health supplements," by David McCandless and Andy Perkins, infographic. Aside from being fascinating, this is a good example of an effort to harness the academic lit for the benefits of the masses. Also, it is representative of the open data movement, as the authors transparently aggregate their data set in a google doc, which anyone can view.

April: "The data-driven life," by Gary Wolf, 5808 words. Discusses the growing trend of self-experimentation. More generally, he discusses how many more people are using tech and data to inform decisions, trumping their raw intuition.

May: "The moral life of babies," by Paul Bloom, 6026 words. He discusses how our preference towards actors who "do the right thing" emerges very early. That is, it presumably emerges far earlier than the babies would be cogent enough to consciously reason about morality. This is part of a movement in psychology that is emphasizing the arbitrariness of our beliefs and decisions.

June: "Smarter than you think: IBM's supercomputer to challenge 'Jeopardy!' champions," by Clive Thompson, 6609 words. At any given point, the AI iteratively calculates the probability that an answer is correct, and then checks whether that probability passes a certain threshold. This probabilistic thinking seems to be invading fields beyond just machine learning, so it's important to understand.

July: "New developments in AI," by Steve Steinberg, 5496 words. An innocuous and perhaps unfortunate title, but a tour de force of a blog post. He discusses trends in smart cars and massive knowledge-bases, and speculates on how they will affect society. One sentence that's particularly near and dear to my heart is when he writes, "consider that 'what is the best burrito in SF' (an opinion), and 'what do most people consider the best burrito in SF' (a fact) are normally considered equivalent."

August: "A world without mosquitoes" by Janet Fang, 1929 words. She discusses whether we should try to eliminate all of these nasty, virulent insects. The downside is that it would mess with biodiversity in ways difficult to predict, while the upside is that it could save millions of lives. We will face plenty of these type of trade-offs in the coming years, specifically with respect to climate change and geoengineering, and more generally in changing aspects of our natural world that we disapprove of.

September: "Jumping to joy," by Robin Hanson, 212 words. He wonders whether we should experiment more with different lifestyles, and what our failure to do so implies about our precarious sense of self. Questioning which of our selves is the "real" one is trendy these days, boosted in part by things like the implicit association test. Experimentation is also enjoying a resurgence, championed by Dan Ariely.

October: "Lies, damned lies, and medical science", by David Freedman, 6022 words. Explains the problems with current scientific publication and data dissemination systems. Many scientists broadly agree with these critiques of their infrastructure, but lack personal incentives to change them. 

Movember: "Hangover theory and morality plays," by Steve Waldman, 1986 words. He discusses the need to frame causes of the recession in moral terms that anyone can understand, synthesizes relevant economic theories, and holds no punches. It'd be hard to describe the ideas of 2010 without including reactions to the recession.

December: "The hazards of nerd supremacy: The case of Wikileaks," by Jaron Lanier, 4704 words. Wikileaks is one of the defining stories of the year. He explains that we might support the hackers in our hearts, because we perceive them to be the underdogs, but that in our heads we should be much more skeptical.

It's been a fun year of blogging and thanks as always for reading.

Sunday, November 7, 2010

Trust The Ratings Of Others

An '09 paper (link here, pdf here, HT to TC) claims that people make more accurate emotional predictions about a future event when they are simply told how someone else reacted to that event, as opposed to when they are given info about the event. This is somewhat counter-intuitive, so let's look at their evidence.

One of their tests was speed dating. Each guy submitted a photo and some demographic info about himself. Then, each girl predicted how much she would enjoy the date based on either this photo / info or the enjoyment rating of a girl who earlier had gone on a date with the same guy. Next, the guy and girl had their five minute date, (ignore the heteronormativity, my fellow Vassar alums), and finally the girl rated how much she enjoyed herself on a sliding scale of 1 - 100. 

The authors define prediction error as the difference between the girl's predicted and actual enjoyment ratings. Participants made more accurate predictions when they used the first girl's enjoyment rating to predict their own (the avg error was 11.4 +/- 8.7) than when they predicted their enjoyment on the basis of the info (an avg error of 22.4 +/- 10.8).

In classic psyc study fashion, they also asked their participants to say which condition they thought would lead to more accurate predictions. 75% said the info would be more useful than the rating of a girl who had already been on a date with that guy. Oops. Now, indulge me in a few reactions:

1) Why do people underestimate the value of someone else's rating? Probably because people think of themselves and their opinions as more unique than they actually are. This sets up my public choice theory for why popular critics like Anthony Lane tend to be negative and contrarian. Although on the surface this annoys readers, people on a deeper level prefer to read opinions about art that they disagree with, because it allows them to think of their own opinions as more unique.

2) There is only one specified relationship between the study participants: they are all undergrads at the same school. So although the authors toss the word "social network" in towards the end of the paper, their results do not speak to the predictive power of a friend's opinion as opposed to a stranger's opinion. This remains an open question--in predicting their own enjoyment, will people find the opinion of someone in their network more valuable than the average opinion of strangers? Even if you say yes, you must take into account the trade-off of sample size, which is larger when you listen to the masses. The high valuation of sites like facebook relies in large part on the assumption that we'll prefer recommendations from those in our network, but I'm not so sure.

3) A subsequent study looked at how people combine their own mental simulations and third-person reports of other's experiences in making judgments. Corroborating the results of this study, they found that people assign far too much weight to their own simulation of how an event will play out as opposed to the feelings of other people who have actually experienced the event. I myself find this all very relevant to imdb's movie ratings. Remind me again why you trust yourself to judge a movie instead of deferring to the aggregated ratings of others?

Monday, September 6, 2010

The Arguments For And Against Re-Rating Movies

The Prosecution: Changing one's mind about the quality of a given movie, or for that matter any given work of art, is a disgusting practice that ought to be accompanied by ruthless social disapproval. Re-rating allows and even encourages one to incorporate other's opinions into one's own ratings, heavily biasing them. NaĂ¯vely, many assume that this influence will always move ratings upwards and assure themselves that they won't merely follow the opinions of the most popular critics. But the reality of the conformity cycle is much more insidious: you are just as likely to learn that too many others like a movie and thus dislike it. There is no defense against these influences once you have been exposed to them, thus rating must happen early and remain steady despite the greatest of protestations. Ladies and gentlemen of the jury, I believe strongly, and upon contemplation I believe you will come to agree, that re-rating really is the bane of a high-functioning rating system.

The Defense: The vitriol of the prosecution's ad hominem attacks against everyday folks who happen to re-rate now and then, justified only by some childish appeal for purity, is dangerously short-sighted. If you don't understand a movie the first time you see it, that's not necessarily the movie's fault, it could be your own fault too. Thus it's totally understandable that, if you come to understand some angle of the movie better after conscious or unconscious contemplation, your rating might change. Moreover, the quality of a movie cannot be fully judged right after watching, because the quality of a movie is based not only on your experience during the movie, but also the value over replacement of any subsequent thoughts about that movie after watching. Thus a rating must be dynamic; it will change with the ebbs and flows of one's thought processes, the structure and patterns of one's interior life, and yes, maybe even one's interactions with other people. Re-rating is only natural given all of our other human tendencies, and its availability takes much unnecessary pressure and anxiety off of the initial rating. If we want to evolve as people, and more specifically as a society of movie watchers, then we must be willing to accept the consequences of such dynamicity. The defense rests.

The Verdict: Death by reruns of imdb's bottom 100.

Tuesday, June 15, 2010

Niche Finding

Holding quality constant, I tend to enjoy movies more the lower my expectations are. This seems like a fairly universal tendency. For example, the main predictor of a student's enjoyment of a class is the extent of positive deviation in their actual grade from their expectations (here).

My explanation for this tendency is as a mechanism to spur niche finding. In this large world, it is hard to stake out our own identity. Thus, we constantly are on the lookout for things that we enjoy more than others to portray our unique values and thus define us.

As evidence for this, consider how much people love to note that some particular work of art is underrated. The next time you hear someone say something is underrated, probe a bit.

If you disagree with how good that work of art objectively is, they may give some playful rebuttals but won't really mind. However, if you disagree with their assumption that the work of art is rated low by the majority, and thus imply that they are not really unique for liking it, they will get rather annoyed. So, it is not the actual quality of the underrated thing that people mostly care about, but rather their own uniqueness in liking it.

Thursday, April 15, 2010

Any Rating System Trumps None

"Nihilists! Fuck me. I mean, say what you like about the tenets of National Socialism, Dude, at least it's an ethos." - Walter, The Big Lebowski

There are lots of lists that attempt to describe the top movies, with various methodologies. Here are three of the major ones:

1) The American Film Institute determined its top 100 list (here) by having film "experts" create their own top 100 list from 400 nominated movies. Movies were ostensibly judged based on winning awards (read: the Oscars), popularity (box office, syndication, home video sales), historical significance, and cultural impact. It's by far the #1 most cited list that people mention when I bring up top movie lists, probably because AFI's lists have been on TV a decent amount and when it comes down to it Americans watch a shocking amount of TV.

2) Metacritic compiles its top 200 list (here) by averaging the subjective ratings of various movie critics. They ask for user votes but don't actually count them towards the top 200. The big supposed upside of their list is that the ratings should be higher quality because they are based on published reviews. The main disadvantage of their list is the low sample size, which leads to more random noise. For example, Superman II is #2 on their all time list on the basis of a whopping 7 critic votes, while it has the class average of a 6.7 on imdb based on 20,000+ votes. So, Metacritic needs to convert to a Bayesian system that punishes low sample sizes in some way. Unfortunately, they also don't include many older movies, as most of their reviews are from the past 10 years.

3) The internet movie database determines its top 250 (here) via user ratings. Anyone with a valid email address who can pass a CAPTCHA test can rate an individual movie, but you have to have a certain amount of votes and various other qualities for your vote to count towards the top 250. Qualities which imdb doesn't disclose. They use a system that punishes low sample sizes, avoiding the Superman II problem. Compared to the other two lists, theirs is more diverse, either via old movies (as compared to metacritic) or foreign movies (as compared to AFI). Their big problem is recent movies, which start off much higher than they end up as (see: the Dark Knight), but are not punished as such. Admittedly, the fact that many others are against imdb's ratings probably makes me like it more, and I am also biased because I'm currently watching the top 250 have invested a lot of time into it. But I definitely do think that it's the best.

What are some metrics by which we can compare these systems? One way would be to look at other systems that use fan votes as compared to expert votes. For example, the NBA All-Star game relies on fan votes to determine its starters, while it relies on journalists to vote on the MVP. The fan votes tend to be not highly correlated to the quality of the player that year. Allen Iverson has been voted a starter each of the last three years even though his stats have been awful. MVP votes are probably more correlated to player's statistical success, although experts aren't perfect either: Steve Nash probably shouldn't have won it twice.

So, NBA All Star votes might count as evidence against imdb. And perhaps that kind of example is why people don't trust imdb? I would argue that sports are qualitatively different because most people don't actually watch all of the games, whereas most everyone who votes on imdb has actually watched the movies.

Regardless, I think sober, intelligent minds can disagree about the relative merits of each of these systems. Your personal preference will probably depend based on to what extent you believe quality is universal, and how much you trust the opinion of insider elites as compared to normal folks.

Much more troubling is the lack of any system at all, of just wandering through the world like a little boy, lost, looking for his mommy. For example, critic Johnathon Rosenbaum thinks that presenting AFI's list of movies in order is "tantamount to ranking oranges over apples or declaring cherries superior to grapes." His attitude is just pure nihilism, through and through.

Tuesday, April 6, 2010

You Tube Rating Gets Even Worse

In Nov '08 I called You Tube's rating system "a disaster," and in Feb '09 I explained that they don't care. Instead of improving, it has since gotten worse, or, depending on your frame, is no longer really a legitimate rating system at all. From their shoddy explanation:
Ratings have changed from the Star system to a binary "Thumbs-Up ‘Like’" / "Thumbs-Down" system. Anything other than a 1- or 5-star rating is rarely used on YouTube, and so we moved towards a simpler "Like / Don't Like" model.
What's sad is that there is so much potential at You Tube. So many viewers and your typical proportion of willing raters means they could really impact the world. How cool would it be if there were a top 250 for videos, categorized into music videos, activism, stand up comedy, etc?

Sure there might be more 1 / 5 star ratings than you'd like. So why not incentive 2-4 star ratings by weighing them more, throw out some of the extreme ratings like imdb probably does, or better yet, count the rater's deviation from his own average rating? Switching to a 10 star system couldn't hurt.

Instead of a solid rating system, we must rely on recommendations (with small, insular sample sizes) and feedback-propagating lists of "most viewed" videos. With Google's decision, the internet became a little bit less self-aware. I doubt anyone shed a tear. But maybe we should have.

Wednesday, January 13, 2010

Friend or Algorithm?

Mark Sisson poses a question:
Quick. How’d you hear about your favorite book or album of all time? Did you let an online algorithm determine what genre/artist/author/etc you’d prefer? Or did a trusted friend, colleague, or family member make a recommendation? I dunno about you, but I’ll take personal recommendations from people I trust over what some impersonal line of code thinks I should like, given the choice between the two.
This is a pervasive yet ultimately false dichotomy. Rating systems aren't based on what computer algorithms reverse engineer from the raw electromagnetic waves.* They're either based on the average ratings of other average people (like imdb) or the preferences of specific people who share some of your average characteristics (like netflix). That's the reality. Now, can you not trust these because you consider yourself too special to agree with the plebeian majority? You're free to be elitist, but at least admit it.

What's the other main reason to prefer a "trusted" friend over "impersonal line[s] of code"? To signal loyalty to your group or clique. People signal loyalty all the time** so you shouldn't necessarily feel bad about this, but again you might as well admit the truth to yourself and others before you perpetuate the information cascade.

Even though it is a false dichotomy, if I had to choose I'd still take the algorithm all day. Aggregating more opinions leads to less noise in opinion markets! What about you?

####

* Although that would be outrageously baller.
** I don't want to make it seem like I consider myself above this. In fact this very disclaimer is an example of signaling my loyalty to fellow lovers of transparency, as is this one, this one, etc.

Monday, January 11, 2010

The Internet Echo Chamber?

Some of the comments on Charlie Hoehn's recent post focused on whether the internet is merely an echo chamber or if intrinsic quality plays a larger role. As always in these "nature / nurture proxy" debates the winning answer is "somewhere in the middle," and the more useful question is how much each variable can explain.

To the extent that folk's behavior in listening to and downloading music is indicative of folk's propensity to e-mail, re-blog, or re-tweet articles*, then Mathew Salganik and Duncan Watts's two studies of web-based music listens and downloads, here and here, may be helpful in resolving this debate.

The researchers created a music downloading web site and uploaded 48 songs by unknown bands. They then recruited somewhat tech-savvy individuals to listen to, rate, and possibly download the songs. Folks downloaded on average 1 out of 7 songs they listened to, indicating some modicum of selectivity.

In one study, the researchers assigned all incoming visitors to either the "social influence" condition, in which they could see the rating and downloading behavior of others, or the "independent" condition in which they could not. Within the "social influence" condition, visitors were also assigned to one of a few identical "worlds," which should have different rating and downloading trends due to random chance.

When the songs were presented to visitors in a single column sorted by popularity, social influence was at its highest. Participants listened to the most downloaded song about ~45% of the time and the second and third most downloaded songs ~30% of the time, while they only listened to songs downloaded an average number of times ~5% of the time.

Salganik and Watts then used the download trends of individuals in the independent condition to predict download trends of individuals in the social condition. In experiment 2, knowledge of the independent data decreased naive prediction errors for the social influence condition by 16%. In experiment 3, with older and more international demographics, knowledge of the independent data decreased naive prediction errors for the social influence condition by 38%. This averages out to 27% as a rough proxy for the usefulness of independent appeal data for predicting which songs will be succesful in the social influence condition. Not great, but not that bad!

In the next study, the researchers had similar set up but in two of their social influence conditions they used an intervention: inverting the download rankings after 752 visitors (~27% of the overall number) had visited the site. This immediately increases the number of downloads for the previously lower rated songs, but eventually some of the top rated ones begin to climb back:
This study also included a non-inverted social influence condition to compare and an independent condition to measure intrinsic appeal. The r correlation between download ranks in the non-inverted social influence condition and independent condition is a strikingly high 0.82, corresponding to an explained variance of 67%. The inverted social influence conditions have much weaker correlations of 0.40 and 0.45 (corresponding to explained variances of 16% and 20%), but these show that even when social influence is directly manipulated against what folks independently prefer, there is still a positive trend between intrinsic appeal and downloading trends.

Salganik and Watts also mention the rating incompleteness theorem (see here): "On the one hand, by revealing the existing popularity of songs to individuals, the market provides them with real, and often useful, information; but on the other hand, if they actually use this information, the market inevitably aggregates less useful information." So, it's hard to prevent people from becoming biased by other's preferences because looking at them is is often a rational choice designed to save precious time. In other words, it's hard to nudge away from a Nash equilibrium.

* This is not necessarily an apt comparison. Music downloading is much more private and personal, whereas what you choose to blog or tweet about is much more visible and thus will subject you to more public judging. On the other hand, reading and discussing articles on the internet is much nerdier than music listening and thus participants may have less emotional attachment, leading to more quality-driven preferences. I don't know of any more applicable experiments but please get at me if you do.

Bottom Line: To say that "the internet is an echo chamber, full stop" is foolhardy. Based on these music download experiments, it seems that around 25 to 70% of folk's decisions to are based on the intrinsic appeal of the material. There is also reason to expect that this percentage would be higher if the download data were less public and estimates of popularity were more noisy, as they are in real life.