Video Transcript: Next Steps with Data - Part 2
Welcome. Let's continue this section of the course, where we look at next steps with data by extending what we learned in the last lecture. In the last lecture, we learned about analysis of variance. Now we're going to extend that even further to something we call multiple comparisons. So, remember what we asked
ourselves last lecture, does winter have lower average total users? Winter looks to be lower on average. Well, what did our analysis of variance actually tell us, though? What does it mean if we reject the null hypothesis on that F test that I showed you, that one way analysis of variance. What was the alternative hypothesis to that analysis of variance? It said that we had enough evidence that at least one category is different, but hold on a second, that doesn't help. I know at least one category is different, but which category is different? That's the next step we do after Anova. Anova can only tell you that something is different, but what we want to do now is we want to be able to test all the individual pairs of categories defined where all the differences actually are. This process is called multiple comparisons, or sometimes ad hoc testing. So, again, analysis of variance says a difference exists somewhere. Multiple comparisons dives deep to figure out where the differences actually are. Now, you may be thinking, why don't we just jump straight into a multiple comparisons? Well, multiple comparisons, if you think about it, we have a lot of different things we're comparing. So, we have four seasons, which means I need to test spring to summer, spring to fall, spring to winter, summer to fall, summer to winter, fall to winter. That's six different tests I need to run, and maybe none of them are actually different. Analysis of variance can tell us that analysis with variance would be able to tell us is something even here before you start diving down and start doing lots and lots and lots of hypothesis tests, and so that's the beauty of analysis of variance. Analysis of variance can basically steer you away and say there's nothing different here, you don't need to worry about it, or it could say you should dig further, there is some kind of difference, but like I said, we're going to be doing lots of hypothesis testing, which leads to an idea. There is a problem with multiple comparisons. Let me try and help you see it. You have a coin which lands on heads, let's say 50% of the time when we flip it. Well, what's the probability of flipping a head on your first flip? It's not a trick question, I promise. It's 50% Well, what's the probability of flipping ahead on your second flip? Again, not a trick question. It's also 50%. Ah, here's the harder one. What is the probability of flipping at least one head in the two flips? So, hold on, let's think about how we can flip a coin twice, and what could actually happen if we flip a coin twice. We can get two heads, that would be at least one. So, that's a yes. We could get heads and tails, that's at least one head, so that's a yes. We can get tails and heads, so that's at least one, so that's a yes. And we can get tails and tails. Well, there's no head there, so that's a no. But wait a minute, out of those four different outcomes, we can get at least one head three of the four times, so although the probability of flipping a head on our first flip is 50% and
the probability of flipping a head on our second flip is 50% the probability that at least one of the flips provides a head is 75% Now you may be thinking, what does this have anything to do with multiple hypothesis tests? I'm going to change some of the wording around here. Let's say you have a hypothesis test. Let's say that hypothesis test makes an error 5% of the time when you perform it. Remember what we talked about with hypothesis testing, they're not always right. The whole point of hypothesis testing is you are testing on a hypothesis with sample data, so you will make mistakes. Now, let's imagine the mistake you make happens only 5% of the time. Well, what's the probability. Of making a mistake on your first test, 5% Well, what's the probability of making a mistake on your second test? That's also only 5% but just like the flip of a coin, what's the probability of making at least one mistake in these two tests it's higher than 5% In fact, it's nine and a half or nine and three quarters percent. See, the problem: each individual hypothesis test has a 5% error. We call these comparison wise errors. Basically, what's the chance of you making an error on a single test? However, if you perform many, many, many, many tests, the probability of making at least one error on all of those tests is called the experiment-wise error rate. What's the probability that at least one of these tests didn't tell me the truth? That's going to increase with the more tests that you run. Again, go back to our coin flip. What's the probability of you getting at least one heads when you keep flipping a coin? That probability is going to keep going up and up and up and up, the more times you flip each individual flip only is a 50% chance of heads, but if we were looking at all the flips and said what's the chance of getting at least one head, that's going to keep increasing, so this is where things become a problem again. Comparison wise error rate is the error rate for an individual test or an individual comparison. The experiment wise error rate, though, is the error rate across all of the comparisons. It's basically looking at what's the proportion of experiments in which at least one error occurs now. Tests like we talked about in an earlier section, and confidence intervals like we talked about in an earlier section, typically control for what we call comparison wise error rates. They control for alpha, but ideally, when we run many different hypothesis tests, we want to control for the error rate across the entire experiment. Let me show you just how bad it can get. Let's imagine that your comparison-wise error rate is just 5% So I compare two groups to each other, that's just one comparison. Then I have a 5% chance of making a mistake, and when you only have one comparison, the experiment-wise error rate and the individual error rate is the same. But let's imagine I had three groups I wanted to compare. Well, that's three comparisons: group A to group B, group A to group C, and group B to group C, with three different comparisons, each one of them having a chance of a 5% error. I actually get a 14% chance that at least one of them has made a mistake. What if we only had four groups to compare, like our example, where we have four seasons. Well, with four seasons, we're making
six comparisons. We're running six hypothesis tests. The chances of making at least one mistake in six hypothesis tests, when each hypothesis test has a 5% chance of making a mistake that's jumped all the way up to 26% In fact, let's just look at five groups. If you were comparing five groups, that would be 10 comparisons. The chance in those 10 comparisons of you making at least one error is all the way up to 40% so you can see this gets out of hand really, really quickly, and because it gets out of hand really, really quickly, we need to try and account for it. So, how can we account for this? How can we actually figure this out? Luckily, someone already has both two Tukey and Kramer came together to develop new tests called the Tukey Kramer test that statistically compare averages between two groups, while controlling for all of the pairwise comparisons that will be considered. In other words, it controls for the experiment wise error rate. You basically say, okay, I'm going to change my hypothesis test just a little bit, because I know I'm going to be running many different hypothesis tests, and so Tukey and Kramer came up with a different version of the hypothesis test. To be able to account for this now, again, we're not going to go through all of the math here in this course. This is just again laying the idea and the foundation for things to consider in future studying around statistics. So let's again bring up our ANOVA example. This is the same slide I showed you last time. The average total users is the same across seasons. It's again the null hypothesis being the average of spring equals the average of summer equals the average of fall equals the average of winter. That leads to a test statistic of 128.8 and a p value that's really, really small. So we rejected the null hypothesis and said at least one of the seasons has a different average total number of users, however, if we wanted to do a multiple comparisons, we could get a Tukey Kramer multiple comparisons chart like this. What you see here are a bunch of confidence intervals. So let me highlight these confidence intervals for you. The center dash in each one of those lines is the difference between the two seasons on average, but you notice there's a whisker coming out on each side of those center dashes that provides you the actual confidence interval, specifically a 95% confidence interval. If those 95% confidence intervals contain zero, that's that vertical dash line. Then the two seasons are not statistically different from each other. In fact, there's only one combination where that's the case. Fall is only 264.2 total users a day lower on average than the spring. That being the case, I cannot statistically say that fall and spring are different than each other in terms of their average. However, every other comparison I can say something. In fact, let's look at all of the comparisons with winter. Winter does not seem to be close statistically to any of the other months, in fact, the biggest differences we see are from winter to every other month. Winter to summer is the biggest difference, followed by winter to spring, then winter to fall, and then summer to spring, fall to summer, and then fall to spring. So again we see this idea, winter. Now we can say it, winter is
statistically different from every other season. The analysis of variance told us something was different. The multiple comparisons can tell us what's different. It's not just one thing that's different. Winter is different than spring. Winter's different than summer. Winter's different than small than fall. Summer's different than spring. Fall's different than summer. It's just fall and spring. I can't say they're different than each other. So, in summary, if you reject the null hypothesis on a one-way analysis of variance, that means there's evidence that shows at least one category is different. Well, once that difference is detected, though, you probably are wondering which one of the pairs of categories is actually different. To be able to do that, we go through a process called multiple comparisons, or ad hoc testing. However, that process of testing many, many, many, many individual pairs can basically lead to an uncontrolled experiment wise error rate. We can basically have at least one error occurring across all those tests at a much higher rate than the individual errors in each test. Luckily, we saw that people have come along to be able to correct for this problem, and in doing so, we can then reliably compare all these different combinations again. Hopefully, this gets you excited about future study and statistics. This is sort of the idea about more complicated things that we can do. You talked about hypothesis testing when comparing to just a number. Last lecture, we started comparing multiple groups to each other, and now we can compare all of those pairs of groups to each other, so we can finally see where some differences lie, and we can finally answer that question: Is winter really different than all the other seasons, which statistically it is, so that is the end of this lecture. I look forward to seeing you in the next.