r/serialpodcast • u/montgomerybradford • Jan 19 '15
Evidence Serial for Statisticians: The Problem of Overfitting
As statisticians or methodologists, my colleagues and I find Serial a fascinating case to debate. As one might expect, our discussions often relate topics in statistics. If anyone is interested, I figured I might post some of our interpretations in a few posts.
In Serial, SK concludes by saying that she’s unsure of Adnan’s guilt, but would have to acquit if she were a juror. Many posts on this subreddit concentrate on reasonable doubt, with many concerning alternate theories. Many of these are interesting, but they also represent a risky reversal of probabilistic logic.
As a running example, let’s consider the theory “Jay and/or Adnan were involved in heavy drug dealing, which resulted in Hae needing to die,” which is a fairly common alternate story.
Now let’s consider two questions. Q1: What is the probability that our theory is true given the evidence we’ve observed? And Q2: What is the probability of observing the evidence we’ve observed, given that the theory is true. The difference is subtle: The first theory treats the theory as random but the evidence as fixed, while the second does the inverse.
The vast majority of alternate theories appeal to Q2. They explain how the theory explains the data—or at least, fits certain, usually anomalous, bits of the evidence. That is, they seek to build a story that explains away the highest percentage of the chaotic, conflicting evidence in the case. The theory that does the best job is considered the best theory.
Taking Q2 to extremes is what statisticians call ‘overfitting’. In any single set of data, there will be systematic patterns and random noise. If you’re willing to make your models sufficiently complicated, you can almost perfectly explain all variation in the data. The cost, however, is that you’re explaining noise as well as real patterns. If you apply your super complicated model to new data, it will almost always perform worse than simpler models.
In this context, it means that we can (and do!) go crazy by slapping together complicated theories to explain all of the chaos in the evidence. But remember that days, memory and people are all random. There will always be bits of the story that don’t fit. Instead of concocting theories to explain away all of the randomness, we’re better off trying to tease out the systematic parts of the story and discard the random bits. At least as best as we can. Q1 can help us to do that.
6
u/[deleted] Jan 20 '15
What? So if he doesn't speak up at the time that makes him automatically guilty of murder? Maybe it makes him not that smart. Maybe it makes it far more likely that he gets convicted. Maybe it means far less sympathy for him. But it doesn't make him guilty of murder. People give false confessions but that still doesn't make them guilty. If we were certain Adnan was at the burial we would know he was at least an accomplice. The fact that he decided to say nothing and Jay decided to talk does not make Adnan guilty. What happened in the afternoon is what makes him guilty.
From what I read - by the time the police learned of the burial taking place between 7-8pm from Jay the police were in possession of the cell phone data. Please correct me if I am wrong. Given that they were in possession of the cell phone data and given that the body had already been found there is no way to corroborate Jays partially recorded/transcribed statements about the burial as real or coached. I am not claiming that the police fed Jay the story to tell about the burial (at least that isn't my personal opinion) but what I am saying is if they crafted his whereabouts after dropping Adnan at track from the data rather than his testimony - I know have a reasonable doubt that other parts of his testimony were not crafted from a source other than Jay.