Welcome. We are finally here. We are at the last section of the course. This  section of the course serves to lay a foundation, at least through ideas about  where to take statistics from here, about what another statistics course would  cover. Now that you have the foundation in statistics, now that you've looked at  a variety of different things, everything from collecting data to analyzing data,  exploring data, looking at testing with data, building intervals with data. Now,  what we can do is think about what are the next steps we can take with data.  The very first next step we can take is what we call analysis of variance.  Everything we have studied in hypothesis tests and confidence intervals has  focused on one population parameter, essentially an average, for example,  compared to a single number. I think that the average number of total users on a daily basis is equal to 4000. However, sometimes we like to compare multiple  parameters against each other, like comparing two or more averages. For  example, I think that the average number of daily users in winter is lower than  the average number of daily users in summer. See the difference there. I'm not  comparing one thing to a number there, I'm comparing two different things.  When comparing averages between two groups of data, we must think about  how spread out the data actually is. In fact, when comparing many averages,  two or more statistically with each other, we call this an analysis of variance,  also known as ANOVA. The key to calling this an analysis of variance, even  though we're actually comparing averages, is because we need to account for  the spread in the data when we start comparing means. Let me show you how  close are these two values. Let's say you had two groups, group one and group  two, you see the distributions for those two groups here, and those distributions  respective averages. So, you have the average of group one compared to the  average of group two. Well, how close do those values appear to be? The  spread of these two distributions is really not overlapping all too drastically. Most of the values in the left hand distribution are below most of the values in the right hand distribution. So, when looking at these two distributions and looking at  these two numbers, you would say they don't appear to be all too close to each  other. Again, most of the values on the left don't even overlap with most of the  values on the right. However, watch carefully, how close are these values? Now  be mindful, the mu one and the mu two on your screen have not changed in  terms of actual distance. However, when looking at the question, how close are  these values, these two distributions have much wider spread. In fact, the  spread of these two distributions seems to overlap a lot, so much so that these  two values now appear to be rather close to each other. These two distributions  of data don't appear all too separate, so you see the spread of the distributions  that you're comparing really influences your notion of whether or not the  averages that you're comparing are actually equal, so common questions that  ANOVA can help us with, for example, do accountants on average make more  than teachers, do people treated with one of the two new drugs have an 

average higher T cell count than people in the control group? Do people spend  different amounts on average depending on which type of credit card they have? All of these things can be answered with analysis of variance, as well as this  question, does winter have a lower average total users compared to the other  seasons? Remember, we pose this question all the way back at the beginning of this course. Now we have enough under our belts, at least an idea to be able to  think about how we can solve a problem like this, so it looks like winter is lower  on average. However, what did we just see? This is a plot of the averages. We  need to look at how spread out each of these seasons actually are, and we can  see that although winter does appear on average to have a lower value, a lot of  these seasons have a lot of spread in terms of total daily users, so that being  the case, we need to ask ourselves the question, is winter really lower than all  the other seasons or is winter just happened to be a little bit lower here, but not  low enough for us to really think it is. That's what a one way analysis of  variance, also known as a one way ANOVA, can help us figure out. A one way  ANOVA is a design in which independent samples are obtained from two or  more categories. Then we test whether the categories have equal means. So,  for our example, the variable of interest is total users, so we want to look at the  average total users, but across what. Well, our explanatory variable is season,  so I want to look to see if average total users is the same or is different across  seasons. Well, the number of categories I have is four, so I'm looking to actually  see if the average total users change across the four seasons, so how do we do something like this? Well, again, this is a hypothesis test, isn't it? Hence, the  bolded word testing. Here we have a hypothesis test. We want to test two  different means, or here are four different means to see if they're equal to each  other, and with any hypothesis test, what did we learn? The first step to  hypothesis testing is you have to state your null and your alternative hypothesis.  The null hypothesis to a one-way analysis of variance would say that the  average of all of the categories here. Let's imagine we had k categories, but the  average of all of the categories is equal to each other. That's what we're going to assume going in. So, again, the null hypothesis, our initial claim, what we initially think is that all of the groups have an equal mean the alternative to that, though  the opposite of everything being equal is careful here, at least one of the groups  has a different mean. Now you may be thinking the opposite of everything is  equal is everything is unequal, but that's not true. The opposite of everything is  equal is at least something does not match with something else, so at least one  of the means is different from one of the other means. So that's what we're  going to test. We're going to look at these four seasons. We're going to assume  the average daily total users between these four seasons is the same. However, the alternative to that is going to be at least one of the seasons is different than  one of the others in terms of the average total daily users. Now there are some  assumptions we have to pull off with this to be able to do a one-way analysis of 

variance. What you have to assume is that the groups are normally distributed.  You have to assume the groups have equal spread or equal variance, and you  have to assume an independence of observations. But what all these things  mean in our context. Well, groups being normally distributed would mean that  total users across each season follows a normal distribution. Groups having  equal variance or spread means that the total users across each season has  equal variance. So, the spread of winter is going to equal the spread of summer, which is going to equal the spread of spring, which is going to equal the spread  of fall. And the idea of independence of observations is that total users from  each day don't depend on each other. Again, we're just talking about the general idea of analysis of variance. We're not going to get into testing these in this  course, but this should hopefully lay the groundwork for any future courses you  may want to take in statistics. So, when looking at an analysis of variance, we  have two different sources of variation. Those two different sources of variation  can either occur within a category. Or between different categories. Let's talk  about each of them. Between sample variability, basically the variation that  occurs between different categories is variability in that variable of interest that  exists between categories of that explanatory variable, so essentially it's what  can the actual categories explain. Now, the within sample variability, on the other hand, that's the variability in the variable of interest that exists within a category  of an explanatory variable, basically what the categories can't explain, so  between sample variability for us it would be something like the difference  between fall and spring, between sample variability would say the season is  explaining these differences. The within sample variability would basically be  how much variability exists within fall, how much variability exists within spring.  Those are the things that season can't explain, like I can explain why potentially  summer is higher than winter, that is explained by the season, but why summer  one day in summer is different than another day in summer, the season alone  can't explain that. I would need additional information, so again, between  sample variability in our case would be the idea of summer versus winter or  spring versus fall. Within sample variability, it would be summer versus summer.  Why is one day in summer different than another day in summer? Let's take a  look at this visually again. Let's go back to our first two charts, the between  sample variability is how far away the averages are from each other. So, here,  group two compared to group one, the further they are away from each other,  the bigger the between sample variability. The within sample variability is the  fact that everyone in group one doesn't have the same value, and everyone in  group two doesn't have the same value. So, there are some variabilities that  exist within groups that the group itself doesn't explain. Again, maybe there are  some viable reasons why, for example, days in summer are different than each  other, but summer alone doesn't explain it. Maybe it's other information like rain  or weekend or weekday or holiday or something along those lines. However, just

the season alone can't explain that. So, the within sample variability is the  variability that exists within a category. The between sample variability is the  differences in the categories themselves. So let's look at these two things on  these distributions again. Look at the between sample variability here, but look  at how much bigger the within sample variability is, here, that's not a  coincidence. When the between sample variability is bigger than the within  sample variability, we have a lot of evidence to say that these groups are  different. However, when the between sample variability is smaller than the  within sample variability. We don't have a lot of evidence that these groups are  too different. Maybe this is another visual that will help explain it. Let's imagine  you had some idea of total variability, basically the total variability that exists in  your analysis of variance. Part of the reason days are different than each other  is explainable, that is the variability between groups for us it is the variability  between seasons. It is why a day in December is different than a day in July.  However, there also exists variability within a season, so a day in July versus  another day in July. So you take that total variability and break it down into two  groups, and basically you compare the ratio of those two groups. If the within  group, the within sample variability is much bigger than the categories averages  aren't that different, however, if the between group variability is much bigger  than the categories differences or the categories averages are different again if.  Most of the difference between the day in July and the day in December is  explained by season, then it's because of season, but if the difference between  a day in July and another day in July isn't explained very well, then again that's  within a season. Okay. so I'm going to show you all the rest of the steps of the  hypothesis test, even though we don't know how to do them. That's okay. Again,  this is just laying a foundation with a little bit more experience and statistics  under your belt. You could take that null hypothesis, where the average total  users is the same across seasons, or in other words, the average of spring is  equal to the average of summer, which is equal to the average of fall, which is  equal to the average of winter. And you could look at your data, and you could  calculate a test statistic. This test statistic comes from a completely different  distribution we haven't even talked about, called the F distribution. Again, that's  not going to be covered here in this course, but the same idea holds true, the  bigger this number, the smaller the p value, and in fact that's exactly what we  have. The p value for this test statistic is less than 0.0001 in other words, very  rare, so the chances of us seeing the sample averages that we did, and the  sample variances in each of the groups that we did underneath the null  hypothesis that all these groups are equal to each other is really, really small, so we're going to reject the null hypothesis. At least one of the seasons has a  different average total daily users then one of the other seasons. Wonderful.  Again, let's summarize. I know we didn't talk a lot about math, but the main goal  of this lecture was just to talk about idea again. You have all the foundation to be

able to build into something like this in another kind of statistics course, but for  right now we're just going to talk about the general idea. A one-way analysis of  variance is a design in which independent samples are obtained from two or  more categories, then we're going to test whether these categories have equal  means, but remember the way to understand if two categories have equal  means. We also have to look at the variation, and variation variability can come  from two places. It can come from the fact that these categories are different  between different categories, or it can come from the fact that there's a lot of  variability within a category, so between sample variability is the variability in  whatever you're interested in that exists between categories. Within sample  variability is the variability of whatever you're interested in within a category, and from there you can see all the things that we can now do. So, you've looked at  hypothesis testing earlier in the course when comparing to a single number.  Now you can compare any group's averages that you'd like. See how much  more powerful statistics can now get. That is the end of this lecture, and I look  forward to seeing you in the next one.



Last modified: Wednesday, July 1, 2026, 8:16 AM