Video Transcript: Next Steps with Data - Part 1
Welcome. We are finally here. We are at the last section of the course. This section of the course serves to lay a foundation, at least through ideas about where to take statistics from here, about what another statistics course would cover. Now that you have the foundation in statistics, now that you've looked at a variety of different things, everything from collecting data to analyzing data, exploring data, looking at testing with data, building intervals with data. Now, what we can do is think about what are the next steps we can take with data. The very first next step we can take is what we call analysis of variance. Everything we have studied in hypothesis tests and confidence intervals has focused on one population parameter, essentially an average, for example, compared to a single number. I think that the average number of total users on a daily basis is equal to 4000. However, sometimes we like to compare multiple parameters against each other, like comparing two or more averages. For example, I think that the average number of daily users in winter is lower than the average number of daily users in summer. See the difference there. I'm not comparing one thing to a number there, I'm comparing two different things. When comparing averages between two groups of data, we must think about how spread out the data actually is. In fact, when comparing many averages, two or more statistically with each other, we call this an analysis of variance, also known as ANOVA. The key to calling this an analysis of variance, even though we're actually comparing averages, is because we need to account for the spread in the data when we start comparing means. Let me show you how close are these two values. Let's say you had two groups, group one and group two, you see the distributions for those two groups here, and those distributions respective averages. So, you have the average of group one compared to the average of group two. Well, how close do those values appear to be? The spread of these two distributions is really not overlapping all too drastically. Most of the values in the left hand distribution are below most of the values in the right hand distribution. So, when looking at these two distributions and looking at these two numbers, you would say they don't appear to be all too close to each other. Again, most of the values on the left don't even overlap with most of the values on the right. However, watch carefully, how close are these values? Now be mindful, the mu one and the mu two on your screen have not changed in terms of actual distance. However, when looking at the question, how close are these values, these two distributions have much wider spread. In fact, the spread of these two distributions seems to overlap a lot, so much so that these two values now appear to be rather close to each other. These two distributions of data don't appear all too separate, so you see the spread of the distributions that you're comparing really influences your notion of whether or not the averages that you're comparing are actually equal, so common questions that ANOVA can help us with, for example, do accountants on average make more than teachers, do people treated with one of the two new drugs have an
average higher T cell count than people in the control group? Do people spend different amounts on average depending on which type of credit card they have? All of these things can be answered with analysis of variance, as well as this question, does winter have a lower average total users compared to the other seasons? Remember, we pose this question all the way back at the beginning of this course. Now we have enough under our belts, at least an idea to be able to think about how we can solve a problem like this, so it looks like winter is lower on average. However, what did we just see? This is a plot of the averages. We need to look at how spread out each of these seasons actually are, and we can see that although winter does appear on average to have a lower value, a lot of these seasons have a lot of spread in terms of total daily users, so that being the case, we need to ask ourselves the question, is winter really lower than all the other seasons or is winter just happened to be a little bit lower here, but not low enough for us to really think it is. That's what a one way analysis of variance, also known as a one way ANOVA, can help us figure out. A one way ANOVA is a design in which independent samples are obtained from two or more categories. Then we test whether the categories have equal means. So, for our example, the variable of interest is total users, so we want to look at the average total users, but across what. Well, our explanatory variable is season, so I want to look to see if average total users is the same or is different across seasons. Well, the number of categories I have is four, so I'm looking to actually see if the average total users change across the four seasons, so how do we do something like this? Well, again, this is a hypothesis test, isn't it? Hence, the bolded word testing. Here we have a hypothesis test. We want to test two different means, or here are four different means to see if they're equal to each other, and with any hypothesis test, what did we learn? The first step to hypothesis testing is you have to state your null and your alternative hypothesis. The null hypothesis to a one-way analysis of variance would say that the average of all of the categories here. Let's imagine we had k categories, but the average of all of the categories is equal to each other. That's what we're going to assume going in. So, again, the null hypothesis, our initial claim, what we initially think is that all of the groups have an equal mean the alternative to that, though the opposite of everything being equal is careful here, at least one of the groups has a different mean. Now you may be thinking the opposite of everything is equal is everything is unequal, but that's not true. The opposite of everything is equal is at least something does not match with something else, so at least one of the means is different from one of the other means. So that's what we're going to test. We're going to look at these four seasons. We're going to assume the average daily total users between these four seasons is the same. However, the alternative to that is going to be at least one of the seasons is different than one of the others in terms of the average total daily users. Now there are some assumptions we have to pull off with this to be able to do a one-way analysis of
variance. What you have to assume is that the groups are normally distributed. You have to assume the groups have equal spread or equal variance, and you have to assume an independence of observations. But what all these things mean in our context. Well, groups being normally distributed would mean that total users across each season follows a normal distribution. Groups having equal variance or spread means that the total users across each season has equal variance. So, the spread of winter is going to equal the spread of summer, which is going to equal the spread of spring, which is going to equal the spread of fall. And the idea of independence of observations is that total users from each day don't depend on each other. Again, we're just talking about the general idea of analysis of variance. We're not going to get into testing these in this course, but this should hopefully lay the groundwork for any future courses you may want to take in statistics. So, when looking at an analysis of variance, we have two different sources of variation. Those two different sources of variation can either occur within a category. Or between different categories. Let's talk about each of them. Between sample variability, basically the variation that occurs between different categories is variability in that variable of interest that exists between categories of that explanatory variable, so essentially it's what can the actual categories explain. Now, the within sample variability, on the other hand, that's the variability in the variable of interest that exists within a category of an explanatory variable, basically what the categories can't explain, so between sample variability for us it would be something like the difference between fall and spring, between sample variability would say the season is explaining these differences. The within sample variability would basically be how much variability exists within fall, how much variability exists within spring. Those are the things that season can't explain, like I can explain why potentially summer is higher than winter, that is explained by the season, but why summer one day in summer is different than another day in summer, the season alone can't explain that. I would need additional information, so again, between sample variability in our case would be the idea of summer versus winter or spring versus fall. Within sample variability, it would be summer versus summer. Why is one day in summer different than another day in summer? Let's take a look at this visually again. Let's go back to our first two charts, the between sample variability is how far away the averages are from each other. So, here, group two compared to group one, the further they are away from each other, the bigger the between sample variability. The within sample variability is the fact that everyone in group one doesn't have the same value, and everyone in group two doesn't have the same value. So, there are some variabilities that exist within groups that the group itself doesn't explain. Again, maybe there are some viable reasons why, for example, days in summer are different than each other, but summer alone doesn't explain it. Maybe it's other information like rain or weekend or weekday or holiday or something along those lines. However, just
the season alone can't explain that. So, the within sample variability is the variability that exists within a category. The between sample variability is the differences in the categories themselves. So let's look at these two things on these distributions again. Look at the between sample variability here, but look at how much bigger the within sample variability is, here, that's not a coincidence. When the between sample variability is bigger than the within sample variability, we have a lot of evidence to say that these groups are different. However, when the between sample variability is smaller than the within sample variability. We don't have a lot of evidence that these groups are too different. Maybe this is another visual that will help explain it. Let's imagine you had some idea of total variability, basically the total variability that exists in your analysis of variance. Part of the reason days are different than each other is explainable, that is the variability between groups for us it is the variability between seasons. It is why a day in December is different than a day in July. However, there also exists variability within a season, so a day in July versus another day in July. So you take that total variability and break it down into two groups, and basically you compare the ratio of those two groups. If the within group, the within sample variability is much bigger than the categories averages aren't that different, however, if the between group variability is much bigger than the categories differences or the categories averages are different again if. Most of the difference between the day in July and the day in December is explained by season, then it's because of season, but if the difference between a day in July and another day in July isn't explained very well, then again that's within a season. Okay. so I'm going to show you all the rest of the steps of the hypothesis test, even though we don't know how to do them. That's okay. Again, this is just laying a foundation with a little bit more experience and statistics under your belt. You could take that null hypothesis, where the average total users is the same across seasons, or in other words, the average of spring is equal to the average of summer, which is equal to the average of fall, which is equal to the average of winter. And you could look at your data, and you could calculate a test statistic. This test statistic comes from a completely different distribution we haven't even talked about, called the F distribution. Again, that's not going to be covered here in this course, but the same idea holds true, the bigger this number, the smaller the p value, and in fact that's exactly what we have. The p value for this test statistic is less than 0.0001 in other words, very rare, so the chances of us seeing the sample averages that we did, and the sample variances in each of the groups that we did underneath the null hypothesis that all these groups are equal to each other is really, really small, so we're going to reject the null hypothesis. At least one of the seasons has a different average total daily users then one of the other seasons. Wonderful. Again, let's summarize. I know we didn't talk a lot about math, but the main goal of this lecture was just to talk about idea again. You have all the foundation to be
able to build into something like this in another kind of statistics course, but for right now we're just going to talk about the general idea. A one-way analysis of variance is a design in which independent samples are obtained from two or more categories, then we're going to test whether these categories have equal means, but remember the way to understand if two categories have equal means. We also have to look at the variation, and variation variability can come from two places. It can come from the fact that these categories are different between different categories, or it can come from the fact that there's a lot of variability within a category, so between sample variability is the variability in whatever you're interested in that exists between categories. Within sample variability is the variability of whatever you're interested in within a category, and from there you can see all the things that we can now do. So, you've looked at hypothesis testing earlier in the course when comparing to a single number. Now you can compare any group's averages that you'd like. See how much more powerful statistics can now get. That is the end of this lecture, and I look forward to seeing you in the next one.