Showing posts with label Math. Show all posts
Showing posts with label Math. Show all posts

Friday, February 1, 2013

The definitive NFL fan base map (LOLJets)

With the Super Bowl coming up this weekend, there's no shortage of football-related stories bouncing around. Most of them are utter nonsense, but thanks to Deadspin and the Harvard College Sports Analysis Collective, we've actually got one pretty fun study to dive into. Yes, this is more sports nerdness, so jump on board.

Using data culled from Facebook, those good folks were able to put together (and then study) a map showing which NFL teams were the most popular (or most "liked") in each county throughout the nation. That enabled us to see, once and for all, what each team's "fan base" really looked like, geographically speaking. Courtesy of Deadspin, here is that map:


While there aren't too many big surprises there (although Alaska is downright bizarre, including a strange patch of Bills fans in the middle of the state), one thing did jump out at me pretty immediately—where the hell are the Jets fans? Oh, there they are... no, not that big green blob that includes southern New Jersey and Delaware—that's Eagles territory. No, it's that little sliver right on the western end of Long Island, comprising basically one county.

Of course, that doesn't mean that the Jets don't have any fans—it just means there isn't any one area in which they're the dominant team, since they're overwhelmed by Giants fans throughout the New York metropolitan area. In fact, the Jets still check in with the 14th-largest fan base according to the study, despite having no real sphere of dominance. Thanks to the HCSAC people, we have the full breakdown for you as well:


Looking closer at this list, it's pretty clear that winning matters, which shouldn't surprise us. The top 3 teams in terms of fan base also happen to be the top 3 teams in terms of historical Super Bowl appearances—the Cowboys and Steelers have 8, the Patriots have 7. And of the top 12 teams on that list, 9 of them have won multiple Super Bowl titles (only the Saints, Bears, and Eagles have not).

Finally, as the HCSAC folks point out, each team that has won a Super Bowl in the last 9 years currently has more than 1.5 million fans, placing them in the top quarter of the league—since both the 49ers and Ravens currently sit on the outside of that top quartile, it'll be interesting to see what kind of fan base jump they may get by winning this weekend.

All in all, the fan base map jives pretty well with our intuitions—the "New England" Patriots moniker is apt, since all of New England minus a small corner of Connecticut leans toward the Pats (they're also big in Canada, and in the U.K.); the Cowboys dominate a huge portion of the country; and Los Angeles, lacking a team, still seems largely to pull for the Raiders, perhaps pining for the olden days. And despite a brief period of dominance at the turn of the century, the Rams can't seem to secure a fan base, nor can the ever-stumbling Jacksonville Jaguars.

Also, the league's fan base continues to skew toward the northern and eastern parts of the country—I ran the numbers to figure out the total numbers of fans by division, and came up with the following:


The East and North divisions make up the top four, combining for more than 65% of the total Facebook fans. Granted, that's aided in large part by the geographical oddity of the Cowboys being in the "East" division, but even if you were to swap the Cowboys with, say, the Rams, you'd still be looking at a 56.2% edge in favor of the North and East versus the South and West. I think it's interesting that the breakdown is in many ways the opposite of what you might expect to see in college football, where the SEC dominates everything—it's possible, if not likely, that the NCAA is pulling share away from the NFL (and the poor Jaguars) in that region.

As one final note, there are some teams who are simply dominant (in terms of fan support) within their divisions—the Steelers boast 64.8% of the total AFC North fans, followed by the Saints with 60.7% of the NFC South, the Colts with 56.8% of the AFC South, and the Patriots with 52.9% of the AFC East. On the opposite end of the spectrum are the Bills (6.9% of AFC East), Jaguars (9.0% of AFC South), Bengals (9.4% of AFC North), and Redskins (10% of NFC East).

But getting back to this weekend, in case you were wondering what the "fan base" breakdown looks like if you consider only the Super Bowl participants, we've got that for you, too. Once again from Deadspin:


Clearly, the nation is leaning heavily toward the 49ers, which is unsurprising given that they've got almost 30% more total Facebook fans than do the Ravens. I apparently should have split allegiances, given that my hometown of Boston is red and my current home state of Virginia is painted purple. Good prediction, in fact—I literally do not care who wins this weekend. Good talk. Enjoy the game.

[Deadspin]
[Harvard College Sports Analysis Collective]

Wednesday, January 23, 2013

More nerd humor

I'm working on a few posts for later today (and later this week), but until I've got those ready to roll, I thought I'd share some more excellent nerd humor that I've come across lately. First, from the always brilliant nerds over at XKCD:


And second, from imgur, courtesy of my man Killagroove:


Yup, it's a big Wednesday around here. More coming your way later today... I think.

Monday, December 31, 2012

The problem with "most likely" outcomes (an NFL playoffs discussion)

On the final Sunday of the NFL regular season (yesterday, for those who weren't paying attention), there's always a number of moving parts as we try to figure out who is going to make the playoffs and who isn't (and also, who is going to be seeded where). Figuring it out is often a challenge, which is why it's nice to have a handy guide at your disposal to help you through the morass.

Luckily, our friends over at Deadspin were nice enough to provide just that, which was an immeasurable help to me as I sat around and rooted for the Patriots and made myself fatter (thanks to Sam Adams and some homemade lasagna). A Patriots win and a Texans loss meant that my Patriots earned themselves the #2 seed and a first-round bye, which was interestingly contrary to what Deadspin had told me to expect. To wit:
The most likely scenario [in the AFC] is that every team which has something at stake wins—they're almost invariably playing teams that don't—and thus the playoff order is exactly what you see above [Texans, Broncos, Patriots].
That line, when I first read it at 2pm or so, stuck with me as I watched the afternoon's games. What does "most likely scenario" really mean? It turns out that the way that we define our terms has an important impact on the way that we understand and respond to the world before us. That's what I'm about to explain.

Yes, it's true, in each individual game, it's "most likely" that the favorite will win—that's what being a favorite means. Therefore, when you start stringing together potential scenarios, the "most likely combination of outcomes" is, indeed, the combination in which all of the favorites win their games. But that doesn't necessarily make it the "most likely scenario"—there's a subtle but very important difference. Bear with me for a second here, because I'm about to get nerdy.

Let's start from the top here, considering a three-game sample. Let's say that each of the three top seeds coming into yesterday (Texans, Broncos, Patriots) had a 60% chance of winning their game (in the grand scheme of things in the NFL, that's a pretty high probability). To determine the likelihood of ALL THREE of them winning their games, which Deadspin said was the "most likely scenario", we just need to multiply the probabilities. In this case, 60% x 60% x 60% = 21.6% , so the likelihood of all the favorites winning was a little less than 1 in 4 odds.

There are 8 possible groupings of winners in this scenario (Texans/Broncos/Patriots would be one, Colts/Broncos/Patriots would be another, Texans/Chiefs/Dolphins a third, etc, etc, etc), and of those 8 possible groupings, the one where the favorites all win is indeed, as we said, the "most likely combination of outcomes". Here's a super-nerdy chart that shows that point, using the 60% probabilities that I used above.


When you look at it this way, you start to see that the "most likely scenario" isn't that all three teams will win, but that one of the other seven scenarios will occur (in fact, the "at least one upset" scenario is more than three times as likely here, with probability 78.4%).

Sure, any one of those individual outcomes is less likely than the individual outcome of "no upsets", but the reality of the matter is quite different. When we start to group the possible outcomes, we see things with a little bit more clarity. I think that a more realistic way of presenting the available data is the following:
Probability of exactly one upset:                       43.2% 
Probability of exactly two upsets:                     28.8%  
Probability of zero upsets:                            21.6% 
Probability of three upsets:                                6.4%
When you group the scenarios this way, you can see that all three teams doing what they're "supposed" to do is, in fact, far from the "most likely scenario" (NOTE: in order for it to become the "most likely scenario" the way I define it, the probability of each favorite winning would have to be more than 75%, as opposed to the 60% that I am using; I find 75% to be way too high for any NFL game). The statistics say that we should probably expect at least one upset, and that "exactly one upset" is the most likely scenario—which, unsurprisingly, is exactly what we ended up with.

Of course, in advance, we can't possibly know which game was likely to produce the upset, but it's almost beside the point. What we can know is that the more times we flip a coin, no matter how lopsided toward "heads" the coin may be, the more likely it becomes that it will eventually come up "tails".

Ultimately, the more independent variables (games) you start linking together, the more likely it is that your "most likely outcome" involves an upset (or a couple of upsets) somewhere along the way. In a sense, this is a similar statistical problem to the birthday problem, which I discussed here once before.


So, why does all of this matter? I'll make this part quick. Let's say you're the Broncos. You're sitting at home this week as the #1 seed, on your bye, trying to decide which team to prepare for (remember, the NFL re-seeds after the first round) while the Wild Card Weekend games are being played. With two games being played—Texans hosting Bengals, Ravens hosting Colts—there are four possible scenarios: (1) Texans and Ravens win, (2) Texans and Colts win, (3) Bengals and Ravens win, (4) Bengals and Colts win.

Assuming once again that the favorite has a 60% chance of winning, we get the following probabilities:
Texans/Ravens (Broncos play Ravens):                 36% 
Texans/Colts (Broncos play Colts):                       24% 
Bengals/Ravens (Broncos play Bengals):               24% 
Bengals/Colts (Broncos play Bengals):                  16%
So, using the same logic we used before, the most likely individual scenario is that both favorites (the Texans and Ravens) win their games, and so the Broncos should be preparing to play the Ravens next week... right?

But no, look again. Even though the "Bengals/Ravens" and "Bengals/Colts" scenarios are individually less likely than the "no upsets" scenario, they combine to be more likely. Because we've given the Bengals a 40% chance of winning their game, and because the Broncos will play the Bengals no matter what if they do indeed win, the Bengals are in fact the Broncos' most likely opponent. Sure, it's only by a small amount, but it's still relevant from a preparation standpoint—Denver should spend at least as much time watching Bengals film as Ravens film, if not more.

Any time we use statistics, we need to be careful with what we're really saying when we communicate our findings or beliefs. In the case of the Deadspin piece, the analysis in question wasn't wrong, it was simply imprecise (and possibly incomplete). When we as readers read that something is the "most likely" scenario, we're almost certainly hoping for something better than a 21.6% probability. Personally, I greatly prefer the much higher 43.2% probability that I ascribed to the "exactly one upset" scenario—I especially prefer it as a Patriots fan, whose team benefited greatly from the way things turned out on the field yesterday.

Good statistics (and good math, and good science, and good writing) requires that we be precise with our methods and our communication of our methods. If we're imprecise, we end up saying things that we don't really mean or that just aren't true.

[Deadspin]

Monday, November 19, 2012

Fun with sports math

Sticking with our sports theme for the day—and adding in some statistical analysis because that's what we do around here—I thought I'd share a couple of recent articles about the baseball postseason that I found interesting. First up, from the Freakonomics blog (emphasis mine):
When the playoff in baseball began, 10 teams – and their fans – were very happy.  But the playoffs being what they are, we knew that only one team – and its fans – would actually be happy when the whole thing was over... 
So what did the Tigers and all the other “losers” (and yes, that includes the Yankees) learn from the playoffs? 
For an answer, let me quote the following from The Drunkard’s Walk: How Randomness Rules Our Lives (a wonderful book by Leonard Mlodinow): 
if one team is good enough to warrant beating another in 55% of its games, the weaker team will nevertheless win a 7-game series about 4 times out of 10.  And if the superior team could beat its opponent, on average, 2 out of 3 times they meet, the inferior team will still win a 7-game series about once every 5 match-ups.  There is really no way for a sports league to change this.  In the lopsided 2/3-probability case, for example, you’d have to play a series consisting of at minimum the best of 23 games to determine the winner with what is called statistical significance, meaning the weaker team would be crowned champion 5 percent or less of the time.  And in the case of one team’s having only a 55-45 edge, the shortest significant “world series” would be the best of 269 games, a tedious endeavor indeed! So sports playoff series can be fun and exciting, but being crowned “world champion” is not a reliable indication that a team is actually the best one. (p. 70-71)
So, no, the "best" team doesn't always win the title, because there's just way too much randomness involved, even in a multiple-game sample as opposed to football's one-game sample. That's why it's entertaining. That's also why the Giants beat the Patriots twice, but I digress.

So if the "best" team doesn't always win the title, maybe the "hottest" team does? Let's ask our friends at the Harvard College Sports Analysis Collective.
Every so often, a playoff series in the NHL, MLB, or NBA will be fought between a team that has just come off of a sweep and a team that has barely survived a competitive 7-game series.  While the latter team is still battling and exerting itself in games, the former will be resting, recovering from the 4-game series, and preparing for the next round... 
Each time a series like this occurs, we are given two contrasting arguments by media figures.  On the one hand, the team that swept has had ample time to recuperate from injuries, rest their bodies and arms, and watch video on both potential teams it could face.  On the other hand, in the large gap of time between games, the team could have “lost momentum,” somehow dissolving the focus and chemistry that had led to the team’s initial success... 
Across the NHL, MLB, and NBA, and looking only at matchups where the previous round was also a best-of-7 series, this scenario has only occurred 29 times throughout history.  The team that has swept has won 20 out of these 29 occasions, and has needed, on average, 5.3 games to defeat its next opponent.  This is not too distant from what one would expect; the teams that swept, in general, have better regular season records, so they tend to be stronger than the opponent who has struggled to emerge from a previous series.  The results, however, are more interesting when broken down by sport. 
Out of the 14 times this matchup has occurred in the NBA playoffs, only twice has the team that went to 7 games in the previous series won the next series...  Much more frequently, the team that has swept in the previous series has gone on to win.  Whether the reason for its winning is that it generally has had better records, or because they were well-prepared and well-rested, is impossible to say for sure. 
The NHL had a similar pattern to the NBA, until 1993; since then, 5 out of 6 teams that went to seven games won the next series against the team that had swept... 
In the MLB, this type of matchup has only occurred four times, mainly because the LCS is the only 7-game series that occurs before another series, and the LCS has not always been a 7-game series.  In all four of these matchups... the team which went to 7 games in the LCS won the World Series...
Although the few data points we have suggest such, concluding that rest is more important in the NBA, whereas momentum is more important in the MLB and NHL is impossible.  In truth, both of these components probably impact the outcome of a playoff series, but probably even more important is how good at winning the team is. Out of these 29 series, 21 of them were won by the team with the better winning percentage (or, points for NHL).  Being well-rested is helpful — but being good is even more helpful.
Alright, then. So, over time, the best team does win more often than not, regardless of how "hot" they are. But in any given playoff series, whether the team is "hot" or "good" seems to take a serious back seat to "luck". Good talk. All of this really bring us right back around to the greatest sports cartoon of all time, from XKCD (in case you were wondering, yes, this entire post was just an excuse to run this cartoon again... I love it):


So, enjoy your sports, by all means. But don't get too carried away with building glowing and complex narratives based on the results of the games. More often than not, it's just a lot of random noise.

[Freakonomics]
[HCSAC]

Friday, October 5, 2012

Password inequality must end

I know it's been a light week for content, but I've still got a few posts for you on this wonderful fall Friday. I feel like keeping things light around here, so I'll avoid politics and the economy for once (which is honestly why I've been quiet this week—as your mother always told you, if you don't have anything nice to say...). Instead, because I like to provide a public service around here, here's some advice for how to safeguard your ATM pin code so that it's hard to crack.
How easy would it be for a thief to guess your four-digit PIN? If he were forced to guess randomly, his odds of getting the correct number would be one in 10,000—or, if he has three tries, one in 3,333. But if you were careless enough to choose your birth date, a year in the 1900s, or an obvious numerical sequence, his chances go up. Way up. 
Researchers at the data analysis firm Data Genetics have found that the three most popular combinations—"1234," "1111," and "0000"—account for close to 20 percent of all four-digit passwords. Meanwhile, every four-digit combination that starts with "19" ranks above the 80th percentile in popularity, with those in the late—er, upper—1900s coming in the highest. Also quite common are MM/DD combinations—those in which the first two digits are between "01" and "12" and the last two are between "01" and "31." So choosing your birthday, your birth year, or a number that might be a lot of other people's birthday or birth year makes your password significantly easier to guess. 
On the other end of the scale, the least popular combination—8068—appears less than 0.001 percent of the time. (Although, as Data Genetics acknowledges, you probably shouldn't go out and choose "8068" now that this is public information.) Rounding out the bottom five are "8093," "9629," "6835," and "7637," which all nearly as rare... 
Some other interesting anedcotes [sic] from the data:  
- Half of all passwords are among the 426 most popular (out of 10,000 total) 
- People prefer even numbers to odd, so "2468" ranks higher than "1357." 
- Far more passwords start with "1" than any other number. In a distant second and third are "0" and "2." 
- Among seven-digit passwords, the fourth-most popular is "8675309," which should ring familiar to fans of '80s music. 
- The 17th-most popular 10-digit password is "3141592654." 
- Two-digit sequences with large numerical gaps, such as "29" and "37," are found often among the least popular passwords.
Wait a minute... half of all passwords are among the 426 most popular? So then, a mere 4% of the passwords are hogging 50% of the password wealth? This is an outrage! This password inequality must end! OCCUPY PIN CODES!!


I am hereby changing my pin code to 2719, because it sounds unpopular. D'oh, forget that I told you that. What I meant to say was, I'm changing it to 1234, because nobody would ever guess that...

[Slate]
 

Monday, September 10, 2012

Statistical significance vs. significance

Alright, it's time to start plowing through some of my unfinished drafts here, in no particular order. I've been sitting on this one for a while, and it follows in the theme of this post and this post, both of which discussed the questionable validity of study results. From the Freakonomics blog... 
A new paper by psychologists E.J. Masicampo and David Lalande finds that an uncanny number of psychology findings just barely qualify as statistically significant.  From the abstract:
We examined a large subset of papers from three highly regarded journals. Distributions of p were found to be similar across the different journals. Moreover, p values were much more common immediately below .05 than would be expected based on the number of p values occurring in other ranges. This prevalence of p values just below the arbitrary criterion for significance was observed in all three journals.
Alright, yeah, I know that's a little stat-wonky/jargony, but the basic point is that a large number of clinical trials that report "significant" results are in fact barely scraping by on the statistical validity scale.

In any statistical study, the "goal" is to show a result that is too extreme to have occurred simply by random chance. A "p-value" of .05 means that there is only a 5% chance that the study result could have occurred simply by chance—low, but not impossible. What we're seeing here is that a large number of "statistically significant" studies are scraping by in this little margin-of-error window just on the "right" side of that 5%. Hence, there's a pretty decent chance that at least some of those studies are reporting something as significant that is actually dumb luck or chance—indeed, probably about one out of every twenty is reporting a significant result when none in fact exists.


Now, I don't really want to go too far down a road talking about bell curves and standard deviations on normal distributions, so I won't. But the point of the matter is, the incentives to report a "statistically significant" result are typically pretty strong, and so we should take a lot of the study results that we read (you know, stuff like "Coffee causes cancer! Also, it prevents cancer and cures cancer, but only when taken in specific doses at pre-determined times over several decades! So drink coffee, and also, don't drink coffee!) with an enormous grain of salt.

A lot of the time, the stuff we're reading is just a reporting of statistical noise and random chance, with a catchy headline attached. So please, people, don't fall prey to the people who want to confuse us with numbers—they're seriously everywhere these days, especially in an election year. Know the statistical background, and you'll be better able to determine for yourself whether a study result is actually significant, or just statistically significant.

[Freakonomics]

Wednesday, September 5, 2012

Cell phone ban update

A couple of years ago, I wrote a post railing against the idiocy of cell phone bans in cars (and Transportation Secretary Ray LaHood's self-described "rampage" against automotive cell phone usage). I wrote then,
Secretary LaHood himself recognized the need for "personal responsibility", which has been in a precipitous decline over the past several decades. You don't create personal responsibility by legislating choices away from the people; quite the opposite, in fact. The more people rely on the government to save them from themselves, the less incentive they will have to make the right choices in the first place (or to deal with any developing problems on their own). Over-reliance on the government has been a key factor behind the disappearance of personal responsibility in America. 
Often, these types of policies and bans are quick-fix responses to avoid taking on more complicated issues. The fact is, distracted drivers existed on the road long before cell phones were prevalent, and they'll exist even if we do ban cell phones in cars entirely. We need to focus on the real problem--irresponsible, unqualified drivers--rather than their symptom, cell phone usage.
I still firmly believe the words that I wrote then—I think the "nanny state" concept is flawed at its core, because it encourages people to outsource responsibility for their actions. That "loss of personal responsibility" dynamic is one that I've more recently discussed in this post and this post, from varying angles. And wouldn't you know it, there's some more recent data that backs up my original hypothesis.
You can take the driver away from the cell phone, but you can't take the risky behavior away from the driver. That's the conclusion of a new study, which finds that people who talk on their phones while driving may already be unsafe drivers who are nearly as prone to crash with or without the device. The findings may explain why laws banning cell phone use in motor vehicles have had little impact on accident rates.
The study involved 108 people, equally divided into three age groups: 20s, 40s, and 60s. For each person, the researchers correlated answers on a questionnaire with data collected from on-board sensors during a 40-minute test drive up Interstate 93 north of Boston...
No cell phones were allowed during these trips. Instead, before they got behind the wheel, the study participants filled in answers about how often they used a cell phone while driving, how they felt about speeding and passing other cars, and how many times in the last year they had been warned or cited for speeding, running traffic lights and stop signs, and other infractions. The team grouped the participants into "frequent users" (those who talked on the phone while driving a few times a week or more) and "rare users" (those who talked while driving a few times a month or less).
Compared with people who rarely talked as they steered, frequent cell phone users drove faster, changed lanes more frequently, spent more time in the left lane, and engaged in more hard braking maneuvers and rapid accelerations, according to the SUV's onboard equipment. Frequent cell phone users, for example, zoomed along about 4.4 kilometers per hour faster on average and changed lanes twice as often, compared with rare users.
"These are not 'oh-my-god' differences," says study leader Bryan Reimer, a human factors engineer at the Massachusetts Institute of Technology (MIT) in Cambridge. "They are subtle clues indicative of more aggressive driving." What's more, he says, other studies have linked these behaviors to an increased rate of crashes. "It's clear [from the scientific literature] that cell phones in and of themselves impair the ability to manage the demands of driving," Reimer says. But "the fundamental problem may be the behavior of the individuals willing to pick up the technology."
Right. The fact of the matter is, if you ban cell phones in cars, there will still be idiots out there who are fiddling with the radio, screwing with their iPad, trying to do their hair (or shave) while driving, or just, you know... speeding and switching lanes all over the place.

Simply put, the people who are most likely to use their phones in their cars ARE ALREADY BAD DRIVERS, so to blame their accidents on their cell phones is basically ludicrous—it's a basic and classic correlation-versus-causation screwup, which we're seeing all over the place these days. We can blame as many accidents as we want on "cell phone use", but the implicit assumption in doing so is that absent a cell phone, the driver in question would have been paying attention and doing everything right. I think that assumption is wrong, and I think that this study helps back up my belief.

As long as we are allowing bad drivers to obtain licenses and continue driving in dangerous fashions, there will be accidents out on the roads. The only thing we should be looking to ban is bad drivers, and I haven't seen any legislation out there that attempts to do that (unless... driverless cars... nahhhhh). If there was, I'd support it in a heartbeat.

[Science]


Tuesday, May 22, 2012

Public opinion surveys, too?

Last week, I ran a post arguing that just about all scientific studies (or at least, their "conclusions") were warmed-over B.S. designed to confuse a public with limited knowledge or understanding of statistics. This week, I came across an article from the Pew Research Center that indicated that public surveys are in danger of becoming similarly unreliable.
For decades survey research has provided trusted data about political attitudes and voting behavior, the economy, health, education, demography and many other topics. But political and media surveys are facing significant challenges as a consequence of societal and technological changes. 
It has become increasingly difficult to contact potential respondents and to persuade them to participate. The percentage of households in a sample that are successfully interviewed – the response rate – has fallen dramatically. At Pew Research, the response rate of a typical telephone survey was 36% in 1997 and is just 9% today. 
The general decline in response rates is evident across nearly all types of surveys, in the United States and abroad. At the same time, greater effort and expense are required to achieve even the diminished response rates of today. These challenges have led many to question whether surveys are still providing accurate and unbiased information. Although response rates have decreased in landline surveys, the inclusion of cell phones – necessitated by the rapid rise of households with cell phones but no landline – has further contributed to the overall decline in response rates for telephone surveys.
The folks at the Pew Center also included this handy chart, which shows just how much deterioration has occurred in just the last 15 years:


There's a lot more to the article than just that blurb, and I'll admit that I'm not fully doing it justice by focusing only on this one dynamic. But I think it's an incredibly important trend, and one that most people are completely unaware of--nobody really knows these days how their statistical sausage is made (let alone their actual sausage, but let's just leave that for another day).

The problem with unreliable public opinion surveys, much like unreliable scientific studies, is that their results often impact future behavior. For example, if I'm considering voting for Ron Paul in an upcoming election, but opinion surveys show that he is polling only about 10%, I'll likely re-consider my vote because he looks "unelectable", regardless of whether or not this is actually the case. On the flip side, a poll that overstates a candidate's popularity could perversely make him more popular among other voters (in fact, I'm pretty sure this is what happened with Rick Santorum). Therefore, public opinion surveys can at times become self-fulfilling prophecies, which is always a bit of a problem.


But response rate isn't even the only problem that opinion polls are currently facing. I've recently participated in multiple phone surveys (one political, one economic), and both times I was literally laughing at the questions as I heard them because they were structured in ways that couldn't possibly reflect my actual thoughts or beliefs--they required overly simplistic answers, or else falsely limited my universe of possible answers. (An example: one caller asked if I aligned myself more closely with the Republican or Democratic party platforms. I responded "neither", and she told me that wasn't an option. I asked her if she was serious. She said she was. I laughed and said fine, flip a coin for me. She refused. So I flipped a coin. It came up tails. I went with the Democrats.)

Our problem is that we've all become increasingly obsessed with data, even as the quality of that data has steadily declined. But if you demand that your data tell a story, they'll tell one for you, even if it's an imaginary story. With response rates this low, that's basically what public opinion surveys have become--imaginary stories to be reported in national news outlets as if they were real news. It's sort of sad, really.

[Pew Research Center]

Friday, May 18, 2012

(Almost) all studies are B.S.

Since I've written here so many times about bad math and bad science and how people use it to confuse the broader public (and often themselves), I felt the need to share this piece from Science Daily that backs up a lot of what I've been saying.
The largest comprehensive analysis of Clinicaltrials.gov finds that clinical trials are falling short of producing high-quality evidence needed to guide medical decision-making. 
The analysis, published May 1 in the Journal of the American Medical Association, found the majority of clinical trials is small, and there are significant differences among methodical approaches, including randomizing, blinding and the use of data monitoring committees. 
"Our analysis raises questions about the best methods for generating evidence, as well as the capacity of the clinical trials enterprise to supply sufficient amounts of high quality evidence to ensure confidence in guideline recommendations," said Robert Califf, MD, first author of the paper, vice chancellor for clinical research at Duke University Medical Center, and director of the Duke Translational Medicine Institute.
People (and companies, and governments) these days are increasingly desperate for "data", but they generally don't have a clue how to interpret it--or, to handle the fact that most studies are, by their nature, "inconclusive". This Freakonomics post similarly pointed out another troubling trend, the increasing rate of research retractions (due to improper methodology, improper interpretation, or both).

Feeding off of that dynamic, journalistic outlets often find ways to massage or outright manipulate the "data" from these studies to suit their storytelling goals. There are, after all, a lot of ways that we can frame statistics in ways to make them misleading--there's the basic crime of omission, where we leave out a lurking variable that actually makes our "conclusions" meaningless, but there are countless other ways in which we can manipulate study results in ways that make them seem more meaningful than they are, or that obscure the actual truth of the matter (I'll save a deeper discussion for another day).

As consumers of media of various types, we all need to be incredibly careful about how we read study results when they are published. We need to educate ourselves about how the studies were performed, determine what the statistics are actually telling us, and decide how we should interpret those statistics, rather than allowing a biased party to mislead us with a headline interpretation. When you read an article that says (for example) that drinking coffee triples the risk of dying from a lightning strike, ask yourself whether that makes any sense at all, or whether there could've been something else going on that explained the study results.


Please, people, don't take "studies" at face value. There's just too many ways to manipulate the results, and too many people with too great an incentive to manipulate them. Remember, scientists who continually return "inconclusive" results from their studies generally won't keep getting grants to perform further studies. It's to their benefit (and can often bring great fame) to report incredibly surprising and unexpected study results--but that doesn't mean their "results" are in any way correct.

Be careful out there. Defrauding the public can be an incredibly profitable business, if you allow it to be. Educate yourselves, and don't buy into the scams. That way you'll be that much more capable of identifying the things you really have to worry about in the world. Like the penguins.

[Science Daily]

Wednesday, May 2, 2012

Catching up on old drafts (a link dump)

I recently promised you all that I had a huge backlog of post-worthy material that I hadn't yet gotten around to writing about. Well today, I took a look at my unfinished draft posts... and it turns out I've got 17 of them, just from the last six weeks. Yeah, I may have the best of intentions, but there's no way I'm going to make it through a 17-post backlog without doing a link dump. So... here goes nothing.

Why Don't Women Patent?
Alex Tabarrok; Marginal Revolution

In this blog post, Tabarrok passes along a recent NBER paper that argued that if women studied science and engineering at the same rate as men, our total number of patents would increase by 24% and GDP would increase by 2.7%. Tabarrok points out that this argument is incredibly specious on multiple levels, and he mocks the finding by suggesting that if we churned out more female construction workers, the construction industry would boom and we would be building tons more houses.

While gender inequality is a serious issue that requires deeper discussion, doe-eyed (and naive) estimates like these only cheapen the argument. I'm reminded of Rob Reid's brilliant TED talk on the "$8 billion iPod", which points out the absurdity of "Copyright Math". This NBER study seemingly suffers from many of the same statistical shortcomings, and Tabarrok and I think it's another clear example of bad math.



Yelp, You Cost Me $2000 by Suppressing Genuine Reviews, Here’s How You Fix It 
Justin Vincent; justinvincent.com

Justin Vincent passes along his personal story about how Yelp's user reviews caused him to make a terrible decision when choosing a moving company, largely because the site's algorithms had blocked a number of negative user reviews that were, in fact, genuine.

I think this is an interesting follow-up to my previous post about "astroturfing" and the difficulties of determining which web product reviews are legitimate, and which aren't. I think we're all still figuring this puzzle out (both as companies and as consumers), and until we do figure it out, our best bet is simply to rely on good old word-of-mouth marketing. Yes, I mean word-of-mouth from real actual people that we know and talk to, not faceless "people" on the internet...

For Regulars and Restaurants, Many Happy Returns
Richard Morgan; Wall Street Journal

The Wall Street Journal's Richard Morgan shares a feel-good story about a loyal New York City man (Bruce Davis) who sits down at the same bar stool at the same Greenwich Village restaurant every night of his life, racking up an average of... wait a minute, $4,000 a MONTH in charges? Are you serious?

If ever there was an article that displayed the obvious gap in our country between the haves and have-nots, this article was it. Recall that median HOUSEHOLD income in the United States is just a hair over $30,000 (in pre-tax dollars, of course), while Mr. Davis spends nearly $50,000 in after-tax money at one restaurant alone. Yes, Mr. Davis' America is not most people's America, but then again, the Wall Street Journal is not most people's newspaper. It's been making that fact abundantly clear in recent months...

Joey Votto's New Contract Is Like a Mortgage-Backed Security
Jack Dickey; Deadspin

This post from Deadspin almost certainly deserves better than to be buried here at the bottom of a link dump, but so be it. In discussing Joey Votto's monster contract extension from the Cincinnati Reds in early April, Dickey made the comparison to a mortgage-backed security. I didn't see the parallels at first, but he did a terrific job of laying them out, and I think that the whole piece is worth a read, especially if you're a cable TV subscriber (yes, this impacts you whether or not you're a sports fan).


This just may be the best piece of sports-related journalism I've read so far this year--it in fact crosses over into financial territory, and it makes a whole lot of sense (it's also sort of terrifying). Given that the recent eye-popping $2 billion purchase of the Los Angeles Dodgers by a Magic Johnson-led group used similar modeling assumptions (and we know those are never wrong, right?), it's clear that the math around cable rights fees is a very pertinent topic. Read this piece and you'll understand how you're footing the bill for Joey Votto's contract, even if you don't know who Joey Votto is.

Friday, March 30, 2012

Mega Millions: should you take the lump sum or the annuity?

If you're thinking of playing Mega Millions for the massive jackpot this weekend (don't bother, by the way, I've already got this thing locked up), there are some important things you should think about first. Yes, "what should I name my boat?" is a very important question, but before you get there you should probably figure out whether to take the lump sum or the annuity. I'm glad you asked. Here are my tips. If you're lazy or hate math, here's the Cliff Notes: take the lump sum.

According to the Mega Millions website, the currently-advertised jackpot of $540 million will result in a lump-sum payment of $389 million. That's cash in hand, today. Or, you can spread that $540 million out over 26 years, pulling in about $20.77 million per year from now until 2037. In essence, by taking the annuity, you're lending your money to the lottery administrators for 25 years.

So what kind of interest rate are you getting for being such a generous lender? Without boring you too much with the math, when I compared the lump sum to the annuity, it came out to an implied interest rate of about 2.84%. With U.S. Treasuries for a similar term (we'll look at the 30-year bond) currently yielding around 3.25%, this seems like a pretty bad deal for the lottery winner (though a very good deal for the lottery administrators).


The other thing to take into consideration is taxes. By any historical standard, tax rates (especially on large amounts of income) are at staggeringly low levels. With government deficits and debt reaching nightmarish proportions, it's pretty much a given that this rate will rise (possibly significantly) at some point over the next 25 years. Therefore, it's to your benefit to generate all of the income (and pay all of the taxes) today, rather than waiting and paying some of those taxes in the future when rates are likely to be higher.

When I ran through the numbers, I found that if tax rates were to remain at 35% for the first 10 years of your annuity payments, and then rise relatively modestly to 45% for the remainder of the 25 years, the implied interest rate on your generous lending would drop to about 2.02%, an incredibly paltry return. If those tax rates increased to 55% in Year 10, your implied interest rate would drop all the way down to 1.08%.

When you consider interest and taxes, it becomes clear that it's not really in anyone's best interest to lend out long-term money right now (hey, maybe that's why nobody can get a mortgage these days...), least of all to the government, who might want to take a bigger bite in future years. So take the lump sum, and try your best not to blow all your many millions in one place. I will try my best to take my own advice when I win this thing later tonight.

Saturday, February 4, 2012

My (completely non-standard) Super Bowl prediction

With all eyes on Indianapolis for Super Bowl XLVI--and my Patriots involved in the big game again--I feel the need to join the parade and weigh in with my prediction for how things will go.

Before I begin, I have to say that I've been overwhelmed by the massive amount of Giants love so far from the "experts". Yes, they've played well lately (otherwise they wouldn't be here), and yes, they present matchup problems for the Patriots in multiple regards--hence the win in the regular season (and no, Super Bowl XLII doesn't mean a thing from a matchup/prediction standpoint).

But they don't build those big buildings in Vegas by accident, and the oddsmakers still give the Pats the edge, even given the murky health of tight end Rob Gronkowski. That might have something to do with the fact that the Patriots were 13-3 in the regular season, while the Giants were 9-7 and actually got outscored by their opponents--but I digress. The point is, these teams are pretty dead even as far as matchups go, and I'm surprised to see any consensus at all among the assorted media.


Now, as you'll know if you've read me often, I have significant philosophical problems with the typical reward system surrounding punditry. Experts are rewarded for making big, bold statements at every turn, knowing that they will get tons of attention for their correct picks, but face little accountability when they are wrong. Predictions (and predictors), therefore, tend to be overly bold, overly confident, and needlessly specific and certain.

All of these dynamics are at play with the standard practice of picking not just the winner of the Super Bowl, but the exact score. I won't be doing that here. Instead, I'm going to be the stat-geekiest stat geek around, giving you a range of possible outcomes, with accompanying probabilities. Hooray! Math fun for everyone! And because I'm not actually making a real pick, I can't possibly be wrong! Alright, fine, I'll make a standard pick at the end, just for kicks, because I know you want it. But I'm not happy about it.

Now... let's get to it. I'm presenting you with three potential game scenarios: The Shootout, The Slugfest, and The Blowout. For each scenario, I'll give you both the probability of that type of game happening, and the conditional probability of each team winning if that scenario develops. Make sense? Alright, cool.

Game Scenario #1: The Shootout (Probability: 55%)  

Recent Super Bowl examples: Super Bowl XLV, Super Bowl XXXVIII, Super Bowl XXXII

Neither of these teams has a particularly great defense. The Giants and Patriots allowed the 27th-most and 31st-most yards of any team in the league this year, respectively, and neither was in the top half of the league in terms of points allowed (the Giants were actually statistically worse than the Patriots, allowing 25 points per game vs. New England's 21.4). Furthermore, both offenses were quite explosive, with Tom Brady and Eli Manning both finishing in the top 5 in the NFL in passing yards.


Therefore, "The Shootout" is far and away the most likely game scenario to develop in Super Bowl XLVI. Picture both teams scoring over 24 points, with the winner posting at least 30. Who has the edge in this scenario? The statistics say the Patriots, despite their shaky secondary. That's mostly because of Brady, who has a higher completion percentage, more touchdown passes, and fewer interceptions this year (and for his career) than Manning.

In games this year where both teams scored at least 24 points, the Patriots were 4-1, while the Giants were just 3-3. Simply put, this year's Patriots are designed to win this type of game. Their edge isn't overwhelming here, but it's significant. In a shootout, the better quarterback usually wins, and Eli is still no Tom Brady.

Edge: Patriots (60% to 40%)

Game Scenario #2: The Slugfest (Probability: 30%)

Recent Super Bowl examples: Super Bowl XLII, Super Bowl XL, Super Bowl XXXVI

This scenario is essentially the opposite of The Shootout, and it would be a repeat of the last time these two teams met in the Super Bowl--a 17-14 Giants victory. Picture both teams scoring fewer than 20 points, and a few big plays (think: helmet catch) making the difference.


Despite the limitations of both defenses, this scenario isn't terribly unlikely. Both defenses have their strengths--generally speaking, it's the Giants' pass rush and the Patriots' ability to force turnovers--and those strengths have shown themselves at various times in the playoffs. Neither defense has allowed more than 20 points in a game so far this postseason (the Giants have allowed an average of 13 per game, the Patriots have allowed 15), and the defenses have combined for 17 sacks and 8 turnovers in the playoffs.

Ultimately, though, this kind of game has to favor the Giants. The Patriots are at their best when they're playing a wide-open game, and they certainly don't want to have to rely on their defense to make stops if they can avoid it. That worked against the Ravens, but barely. It's unlikely to happen again, and the Patriots will need to avoid the turnovers that plagued them in the AFC Championship.

Edge: Giants (70% to 30%)

Game Scenario #3: The Blowout (Probability: 15%) 

Recent Super Bowl examples: Super Bowl XXXVII, Super Bowl XXXV, Super Bowl XXIX

Ah, yes. The Super Bowl blowout. It's been so long since we've had one of these that we've almost forgotten that they were a mainstay of our childhood (assuming that we grew up in the '80s and '90s). Beginning with Super Bowl XV in 1981, 12 of the next 19 Super Bowls were decided by two touchdowns or more, and several of those were severely lopsided affairs (seven were 20+ point blowouts, of which four were 30+ point routs).

Those games have gladly become a thing of the past--seven of the last twelve Super Bowls have been decided by a touchdown or less, and there hasn't been a true "blowout" since the Buccaneers' Super Bowl XXXVII rout nine years ago. But there's still a chance of the blowout returning, even if it's the least likely of the three game scenarios.


But if the blowout does rear its ugly head again, chances favor the Patriots--by a wide margin. Since Tom Brady took over as Patriots quarterback in 2001, no team has won more games or lost fewer games by two touchdowns or more. The Patriots are 61-14 in those types of games (in the regular season; they're an additional 5-2 in the postseason), good for a staggering 81% winning percentage. No other team in the league comes close, and this disparity is a big part of the reason that the Patriots have been so dominant for the last decade--the Patriots rarely get run off the field, so they've pretty much always got a chance.

The Giants, for their part, are an even 33-33 in blowout games over the same span (plus 2-1 in the postseason), and they're only 10-10 over the last three seasons, during which the Patriots have gone a staggering 19-3. The only good news for the Giants is that their two postseason blowout wins have both come in the last month, over the Falcons and the Packers. But history indicates that a third such win is highly unlikely. If the Super Bowl rout returns, chances are it'll be Brady & Belichick who bring it back.

Edge: Patriots (80% to 20%)

So what do you get when you put it all together? Well, an awesome little stat geek matrix, that's what! 


For the record, then, the most likely outcome (of the 6 possible) is a Patriots win in a Shootout, with the Patriots also holding the overall edge. Interestingly, though, the 2nd and 3rd most likely scenarios are both Giants victories--in a Shootout and in a Slugfest, respectively. The take-home lesson is that if the Patriots can avoid a Slugfest--something they were unable to do in Super Bowl XLII--they've got a great chance at winning this game.

Okay, since you made me do it, here's my "standard" pick. My matrix, which is mathematically proven to never be wrong, tells me that a Patriots Shootout win is the most likely, so that's what I'm going with. I have 33% certainty that my pick is correct.

The Crimson Cavalier's needlessly precise, certain-to-be-wrong Super Bowl prediction

Patriots 34, Giants 24

With Gronkowski playing but limited, Brady is forced to use his secondary targets more frequently. He does so very effectively, as both Chad Ochocinco and Kevin Faulk haul in touchdown passes. Wes Welker and BenJarvus Green-Ellis contribute the other touchdowns for the Pats, as New England's offensive line does an admirable job of controlling the Giants' pass rush.

Meanwhile, with the Patriots employing bracket coverage on Victor Cruz for much of the game, Manning must also look to his secondary options for production. He is similarly effective, with Jake Ballard and Ahmad Bradshaw making big contributions in the passing game. A slick receiving touchdown from Bradshaw keeps the Giants close, and they have a chance to tie the game in the late stages. 

Trailing 31-24 with under four minutes to play and the ball in Patriots territory, Manning chooses a bad time to throw his first interception of the day--he tries to force a ball to Cruz in the slot, and Julian Edelman comes up with the crucial pick. A screen pass to Aaron Hernandez brings the ball into Giants territory, where a Stephen Gostkowski field goal gives the Patriots the clinching points.

There. Now that you know what won't happen, you can go ahead and enjoy the game. Do I get to be on ESPN now?

Friday, January 13, 2012

When statistics lie (NFL edition)

One of the primary reasons I started this blog back in 2010 was to shed light on some of the statistical tricks that politicians, corporate executives, and talking heads use (and misuse) so as to confuse, mislead, and otherwise divide the general population. There's any number of tricks that these individuals will use, but almost all of them rely on the same basic dynamic--highlighting one piece of data while ignoring or obfuscating another that would allow for a more complete and realistic picture. It's basic hand-waving misdirection that would make any amateur magician proud.

To see how widespread these statistical games have become, note that I've previously highlighted this dynamic with respect to cancer, budget deficits, oil prices, student loan debt, tennis, Rick Perry, apple juice, education, airlines, and investing. I've also poked fun at the issue here and here. Simply put, given the nature of the public discourse these days, I don't think you can be an informed person (or make good decisions) unless you fully understand statistics and the way that they can be manipulated. The most recent example of this is the GAO's recent takedown of the Obama administration's claim that TARP made money--in a nutshell, the claim is technically true, but only if you ignore a lot of other things that are also true. Pretty standard statistic-manipulating stuff.

With the NFL Playoffs continuing this weekend with a big game between the Patriots and Broncos, a couple of sports journalists have taken a few liberties with a similar statistic, one that allegedly speaks volumes about the Patriots. Here's the statistic, courtesy of SI's Kerry Byrne:
If there's a legitimate statistical and historical reason to doubt the validity of New England's No. 1 seed and 13-3 record, it's the fact that they faced one cream puff after another -- and then lost each time they faced something close to the iron of the NFL. New England did not beat a single team with a winning record in the 2011 season.
We track something over at Cold, Hard Football Facts.com called Quality Standings -- how well you perform against Quality Teams, or teams with winning records. It's an effective way to separate the contending wheat from the pretending chaff each NFL season. Super Bowl champs typically prove along the way that they can consistently beat Quality Opponents. And that historic fact is not good news for the Patriots.
Not only did they face fewer Quality Opponents than any team in football this year (two), but also they lost to both of them (Steelers, Giants). Would the Patriots have gone 13-3 had they faced eight Quality Opponents like the lowly 2-14 Rams? What if they faced the league-high 10 Quality Opponents who made the Peyton Manning-less season in Indianapolis such a daunting challenge?
Cool. Great statistic, right? The Patriots, as it turns out, didn't beat a "quality opponent" all year. The problem is, the meaningfulness of this statistic depends 100% on a completely arbitrary definition of what comprises a "quality opponent".

A deeper look at the Patriots' 2011 schedule reveals that Brady & Co. ended the season with a staggering 7 wins (and no losses) against teams that finished the season 8-8--including their next opponent, the Broncos. Not a single one of these six teams (the Pats played and won two games against the 8-8 Jets) counts as a "quality opponent", so the Patriots get no credit for the victories, despite the fact that every one of them had a winning record in games not played against the Patriots.

Of course, had the Patriots instead lost each of those 7 games against teams that finished 8-8, those teams would have finished at least 9-7, and therefore the Patriots would have finished 0-9 against "quality opponents". If a team is only a "quality opponent" if you lost the game, but not if you won the game, then clearly there's a problem with your definition of "quality opponent".


A closer look at the Gang of Six reveals that three of them (Chargers, Jets, Eagles) were preseason favorites to reach the Super Bowl, and that nearly all of them could have reached the playoffs had they in fact beaten the Patriots--one of them, the Broncos, made it (and won their first-round game--ironically defeating the "quality opponent" Steelers) despite their loss. It seems far-fetched and dishonest to refer to none of these teams as "quality opponents", because the statistic itself is self-referential and relies on circular references.

Ultimately, this is just another example of the way that people misuse statistics to show a slightly skewed version of the world. In this case, the authors are hoping that you won't notice the weird statistical anomaly that defined the Patriots' 2011 schedule. In the case of TARP, the Obama administration is hoping that you won't notice that the "profitability" of TARP depends entirely on the massive expansion of the Federal Reserve's balance sheet (the so-called "money printing" you've been hearing so much about).

It is exceedingly rare that the truth of the world can be easily distilled into one catch-all statistic, but that doesn't keep our favorite talking heads from trying. It's our job to know when they're telling the truth, and when they're using smoke and mirrors. More often than not, it's the latter. Go Pats.

[SI]

Friday, January 6, 2012

On multiple moving variables

In complex systems with multiple changing variables, it's often difficult to determine which variable is producing which outcome--or which interactions are working together in which ways. It's in large part what makes solving our economic problems so difficult, and it's a dynamic that I first explored in this post. As I wrote then,
I'm often writing here about the perils of bad science, and bad statistics, and how people try to draw conclusions from data that is inherently biased (or at least not properly controlled). In complex systems (and almost everything in human interaction is a complex system), it is almost impossible to reliably isolate and determine the impact of just one variable--but that never stops people from trying...
[In] any system with that many moving parts, it's absolutely impossible to know which one of them is having a positive impact, negative impact, or no impact on the overall outcomes. Remember that next time you see a politician trying to take credit for the supposed successes of "his" policy (or, in the case of Ben Bernanke, simultaneously taking credit and blame-shifting)--there's simply no way of knowing what's helping and what's hurting when we can't isolate just one variable.
So what brings me to revisit this dynamic today? Not the economy (for once), but the NFL. As we prepare for the first weekend of the football playoffs (Wild Card weekend, which hopefully won't be as big a letdown as Week 17), we're all still basking in the glow of a record-setting season. Both Drew Brees and Tom Brady broke Dan Marino's long-standing single-season passing record (and Matt Stafford and Eli Manning weren't far behind), leading to countless columns trying to discuss the record's significance and the reasons that led to its sudden shattering.


Many of the explanations correctly focused on a series of league rule changes meant both to protect players and to open up the passing game--the illegal contact penalty, the "Brady rule", and most recently the renewed emphasis on "helmet to helmet" hits on defenseless wide receivers. All of these--in addition to even wider usage of artificial turf and domed stadiums, and even a mild winter so far--helped to push passing numbers higher across the board.

But throughout all of this, one of the simplest variables has largely been ignored, and it's a variable that could easily help explain why this year, of all years, saw multiple quarterbacks take down the record. Let's turn things over to ESPN for a minute.
Before the NFL season, one of the rule changes that received the most discussion was how moving kickoffs up five yards to the 35-yard line would affect the return game...
Did the rule change deprive fans of excitement? From a numbers standpoint, the answer is yes. Nine kickoffs were returned for a touchdown in 2011, compared to 23 in 2010...
The change becomes glaringly obvious when looking at average starting field position. The average drive after a kickoff started just past the 22-yard line, down almost five yards from the previous year.
With each team averaging 189 possessions this season, that adds up to an additional 888 possible yards per team, which may account for part of the offensive boom in 2011.
Now, just because there are 888 yards "available" to be gained certainly doesn't mean that teams are going to gain them. But when you're talking about some of the more powerful offenses in the game, backed up closer to their own endzones, it's a pretty fair bet they're going to come out throwing. Even if we only give these QBs credit for 300 or 400 of that 888 yard figure, that's enough to bring both Brady and Brees back to even with (or slightly below) Marino's record figure. And yet, this ESPN article is the first time I've heard that statistic mentioned, with most people electing to focus more on the "helmet to helmet" rule.


Realistically, neither rule change can be given 100% credit for having spurred this offensive explosion, but similarly neither can be ignored. In a complex system with several moving variables, it's the interaction among all of them that leads to the outcome we observe. Are Brees and Brady's feats any less laudable just because there were a couple of rule changes? Not really. But when we're dealing with small margins between "record" and "not a record", every little bit counts.

This year, the kick returners' loss was the quarterbacks' gain--but that's not the story you'll end up hearing most often. The lesson, as always, is beware of the popular story--it's almost always too simple to actually be correct.

[ESPN]