Welcome. Let's continue this section of the course, where we look at next steps  with data by extending what we learned in the last lecture. In the last lecture, we learned about analysis of variance. Now we're going to extend that even further  to something we call multiple comparisons. So, remember what we asked  

ourselves last lecture, does winter have lower average total users? Winter looks  to be lower on average. Well, what did our analysis of variance actually tell us,  though? What does it mean if we reject the null hypothesis on that F test that I  showed you, that one way analysis of variance. What was the alternative  hypothesis to that analysis of variance? It said that we had enough evidence  that at least one category is different, but hold on a second, that doesn't help. I  know at least one category is different, but which category is different? That's  the next step we do after Anova. Anova can only tell you that something is  different, but what we want to do now is we want to be able to test all the  individual pairs of categories defined where all the differences actually are. This  process is called multiple comparisons, or sometimes ad hoc testing. So, again,  analysis of variance says a difference exists somewhere. Multiple comparisons  dives deep to figure out where the differences actually are. Now, you may be  thinking, why don't we just jump straight into a multiple comparisons? Well,  multiple comparisons, if you think about it, we have a lot of different things we're  comparing. So, we have four seasons, which means I need to test spring to  summer, spring to fall, spring to winter, summer to fall, summer to winter, fall to  winter. That's six different tests I need to run, and maybe none of them are  actually different. Analysis of variance can tell us that analysis with variance  would be able to tell us is something even here before you start diving down and start doing lots and lots and lots of hypothesis tests, and so that's the beauty of  analysis of variance. Analysis of variance can basically steer you away and say  there's nothing different here, you don't need to worry about it, or it could say  you should dig further, there is some kind of difference, but like I said, we're  going to be doing lots of hypothesis testing, which leads to an idea. There is a  problem with multiple comparisons. Let me try and help you see it. You have a  coin which lands on heads, let's say 50% of the time when we flip it. Well, what's the probability of flipping a head on your first flip? It's not a trick question, I  promise. It's 50% Well, what's the probability of flipping ahead on your second  flip? Again, not a trick question. It's also 50%. Ah, here's the harder one. What is the probability of flipping at least one head in the two flips? So, hold on, let's  think about how we can flip a coin twice, and what could actually happen if we  flip a coin twice. We can get two heads, that would be at least one. So, that's a  yes. We could get heads and tails, that's at least one head, so that's a yes. We  can get tails and heads, so that's at least one, so that's a yes. And we can get  tails and tails. Well, there's no head there, so that's a no. But wait a minute, out  of those four different outcomes, we can get at least one head three of the four  times, so although the probability of flipping a head on our first flip is 50% and 

the probability of flipping a head on our second flip is 50% the probability that at  least one of the flips provides a head is 75% Now you may be thinking, what  does this have anything to do with multiple hypothesis tests? I'm going to  change some of the wording around here. Let's say you have a hypothesis test.  Let's say that hypothesis test makes an error 5% of the time when you perform  it. Remember what we talked about with hypothesis testing, they're not always  right. The whole point of hypothesis testing is you are testing on a hypothesis  with sample data, so you will make mistakes. Now, let's imagine the mistake you make happens only 5% of the time. Well, what's the probability. Of making a  mistake on your first test, 5% Well, what's the probability of making a mistake on your second test? That's also only 5% but just like the flip of a coin, what's the  probability of making at least one mistake in these two tests it's higher than 5%  In fact, it's nine and a half or nine and three quarters percent. See, the problem:  each individual hypothesis test has a 5% error. We call these comparison wise  errors. Basically, what's the chance of you making an error on a single test?  However, if you perform many, many, many, many tests, the probability of  making at least one error on all of those tests is called the experiment-wise error rate. What's the probability that at least one of these tests didn't tell me the  truth? That's going to increase with the more tests that you run. Again, go back  to our coin flip. What's the probability of you getting at least one heads when you keep flipping a coin? That probability is going to keep going up and up and up  and up, the more times you flip each individual flip only is a 50% chance of  heads, but if we were looking at all the flips and said what's the chance of  getting at least one head, that's going to keep increasing, so this is where things become a problem again. Comparison wise error rate is the error rate for an  individual test or an individual comparison. The experiment wise error rate,  though, is the error rate across all of the comparisons. It's basically looking at  what's the proportion of experiments in which at least one error occurs now.  Tests like we talked about in an earlier section, and confidence intervals like we  talked about in an earlier section, typically control for what we call comparison wise error rates. They control for alpha, but ideally, when we run many different  hypothesis tests, we want to control for the error rate across the entire  experiment. Let me show you just how bad it can get. Let's imagine that your  comparison-wise error rate is just 5% So I compare two groups to each other,  that's just one comparison. Then I have a 5% chance of making a mistake, and  when you only have one comparison, the experiment-wise error rate and the  individual error rate is the same. But let's imagine I had three groups I wanted to compare. Well, that's three comparisons: group A to group B, group A to group  C, and group B to group C, with three different comparisons, each one of them  having a chance of a 5% error. I actually get a 14% chance that at least one of  them has made a mistake. What if we only had four groups to compare, like our  example, where we have four seasons. Well, with four seasons, we're making 

six comparisons. We're running six hypothesis tests. The chances of making at  least one mistake in six hypothesis tests, when each hypothesis test has a 5%  chance of making a mistake that's jumped all the way up to 26% In fact, let's just look at five groups. If you were comparing five groups, that would be 10  comparisons. The chance in those 10 comparisons of you making at least one  error is all the way up to 40% so you can see this gets out of hand really, really  quickly, and because it gets out of hand really, really quickly, we need to try and  account for it. So, how can we account for this? How can we actually figure this  out? Luckily, someone already has both two Tukey and Kramer came together to develop new tests called the Tukey Kramer test that statistically compare  averages between two groups, while controlling for all of the pairwise  comparisons that will be considered. In other words, it controls for the  experiment wise error rate. You basically say, okay, I'm going to change my  hypothesis test just a little bit, because I know I'm going to be running many  different hypothesis tests, and so Tukey and Kramer came up with a different  version of the hypothesis test. To be able to account for this now, again, we're  not going to go through all of the math here in this course. This is just again  laying the idea and the foundation for things to consider in future studying  around statistics. So let's again bring up our ANOVA example. This is the same  slide I showed you last time. The average total users is the same across  seasons. It's again the null hypothesis being the average of spring equals the  average of summer equals the average of fall equals the average of winter. That leads to a test statistic of 128.8 and a p value that's really, really small. So we  rejected the null hypothesis and said at least one of the seasons has a different  average total number of users, however, if we wanted to do a multiple  comparisons, we could get a Tukey Kramer multiple comparisons chart like this.  What you see here are a bunch of confidence intervals. So let me highlight  these confidence intervals for you. The center dash in each one of those lines is  the difference between the two seasons on average, but you notice there's a  whisker coming out on each side of those center dashes that provides you the  actual confidence interval, specifically a 95% confidence interval. If those 95%  confidence intervals contain zero, that's that vertical dash line. Then the two  seasons are not statistically different from each other. In fact, there's only one  combination where that's the case. Fall is only 264.2 total users a day lower on  average than the spring. That being the case, I cannot statistically say that fall  and spring are different than each other in terms of their average. However,  every other comparison I can say something. In fact, let's look at all of the  comparisons with winter. Winter does not seem to be close statistically to any of  the other months, in fact, the biggest differences we see are from winter to every other month. Winter to summer is the biggest difference, followed by winter to  spring, then winter to fall, and then summer to spring, fall to summer, and then  fall to spring. So again we see this idea, winter. Now we can say it, winter is 

statistically different from every other season. The analysis of variance told us  something was different. The multiple comparisons can tell us what's different.  It's not just one thing that's different. Winter is different than spring. Winter's  different than summer. Winter's different than small than fall. Summer's different  than spring. Fall's different than summer. It's just fall and spring. I can't say  they're different than each other. So, in summary, if you reject the null  hypothesis on a one-way analysis of variance, that means there's evidence that  shows at least one category is different. Well, once that difference is detected,  though, you probably are wondering which one of the pairs of categories is  actually different. To be able to do that, we go through a process called multiple  comparisons, or ad hoc testing. However, that process of testing many, many,  many, many individual pairs can basically lead to an uncontrolled experiment wise error rate. We can basically have at least one error occurring across all  those tests at a much higher rate than the individual errors in each test. Luckily,  we saw that people have come along to be able to correct for this problem, and  in doing so, we can then reliably compare all these different combinations again. Hopefully, this gets you excited about future study and statistics. This is sort of  the idea about more complicated things that we can do. You talked about  hypothesis testing when comparing to just a number. Last lecture, we started  comparing multiple groups to each other, and now we can compare all of those  pairs of groups to each other, so we can finally see where some differences lie,  and we can finally answer that question: Is winter really different than all the  other seasons, which statistically it is, so that is the end of this lecture. I look  forward to seeing you in the next. 



Last modified: Wednesday, July 1, 2026, 8:17 AM