Video Transcript: Testing Hypotheses with Data - Part 2
Welcome. Let's continue this section of the course where we're talking about testing hypotheses with data by now moving into the next step in hypothesis testing, gathering your data and calculating the test statistic. The test statistic
basically summarizes the amount of information that your sample provides you can sort of think of the test statistic like evidence in a court case, and in all honesty, it'd be so much easier if court cases actually just did this with a single number, because that's what we're going to calculate. We're going to calculate a single number, and that single number is going to represent all of our evidence, and again, like I said, it'd be so much easier if court cases were like that, you know. The prosecution stands up, looks confidently at the jury, and goes 27 and the jury goes, "Oh, wow, okay, it's done, it's over, you know what, guilty. Whereas the prosecution may stand up again, trying to look all confident, and they say something like 1.2 and the jury's like, I'm not impressed. Nope, nope, innocent until proven guilty. So we're going to say not guilty again. See how much easier it would be if that were the case. But either way, so test statistics again are sort of a numerical summarization of the amount of information, the amount of evidence that a sample provides, and the nice part is that test statistics have a common form. A test statistic is basically your sample statistic minus the null value divided by the standard error. Luckily, you've seen all three of those numbers before. Again, the statistic is just the samples information for sample means, it would be exactly that, it would be just the sample mean for sample proportions, it would be just the sample proportion. So, whatever number you're testing from your sample, the statistic, that's what this first number is. The null value comes from the null hypothesis, it's the number that you think the null hypothesis is, so you have some initial thoughts, some initial number in your null hypothesis, so that is what this null value basically represents. So, again, with this numerator in this ratio, we're looking at what does our sample say, and what did we think was going on beforehand? That is going to be what we're looking at the difference between. Remember, in our last lecture, we were looking at the difference, right? We said, well, here's our sample average daily users. Let's subtract off the total number of daily users, or the average daily users we thought it was. Of course, last but not least, we have the standard error. This comes from that idea of variability from the sampling distribution of the statistic. Again, statistic, we can compare where a number is on a distribution, but you have to have some idea of scale, you have to have some idea of variability, right? So, if I told you that I was only estimating, well, I was only off in estimating something by $30 maybe that's impressive, maybe it's not. If I was talking about the US GDP, then wow, if I was only off by $30 I was really close. If I was talking about how much money I spent in McDonald's yesterday. Well, that no, that was off by a lot. So, again, it all has some notion of scale. That's what the standard error is. The standard error has some measure of scale. So, when it comes to sample means, remember when we were dealing
with confidence intervals, when it came to confidence interval, sample means had to look at a t distribution because we didn't know the value of the population standard deviation. It's actually the same thing here. So, when looking at the t distribution, we're just going to call the test statistic for sample means t, just sort of after the distribution that it comes from, so t is going to equal our statistic, which here x bar minus our null value, our original thought, it's the number from the null hypothesis, we call this mu zero, mu naught, and then we divide it by the standard error of x bar, which is just the sample standard deviation s divided by the square root of n, and that's it. That is our test statistic that will essentially tell us how far away our sample data is from our original thought. If our sample data is really far away from our original thought, we start to question our original thought. If our sample data is really close to our original thought, then you know what, our original thought may be right. That's the general idea. The nice part is we could do the same thing for proportions as well. So again, sample proportions use the normal distribution, so we can take a look at the same idea. We can take a look at a statistic p hat. We can subtract off a null value, maybe some original thought about proportions, like for example, I think that only half of the days in DC are clear or cloudy, so that could be some original thought, and then we divide again by some notion of a standard error, the square root of p zero over, I'm sorry, times one minus p zero over n. So let's summarize the test statistic, really is just some numerical summary of the amount of information that your sample provides, that's really all it is. And to calculate this test statistic, you typically only need three things. First, information that you obtain from the sample itself, basically the statistic you're trying to compare. Next, you need information about your original claim, your null hypothesis. We call this the null value, and last but not least, you need some measure of variability to be able to scale this comparison between the statistic and the null value that we use the standard error, but okay, so you've collected enough data, you have all of this data that's here. How do you know how much evidence is enough evidence? When is something rare enough? Well, that's what deals with the next step of hypothesis testing, the p value, as well as the significance level. When we talk about p values, the p stands for probability, and the p value is the probability you got your sample or something even crazier, more extreme than your sample underneath the null hypothesis. So again, remember, why do we think the null hypothesis is true? Well, that's your original thought, that's your original claim. It's just like a court case, remember, innocence until proven guilty. So we're going to assume innocence, we're going to assume the null hypothesis is true, and the p value is going to represent the probability we see all of the evidence that we do, the probability we get the sample that we do underneath that null hypothesis. Again, that's how a case in a court of law should work. A jury should go in, for example, and not assume guilt. They should assume innocence, and if the prosecution can gather enough evidence, then the jury starts to question
that original thought of innocence and says, "You've captured enough evidence to say for us to say that this person is guilty, right? That's the idea, but again, we have to figure out what's the probability we got our sample. Go back to that coin
flip example I gave you. I flip a coin once, it's heads, you still think the coin is fair. I flip it two times, both of them are heads, you probably still think the coin is fair. I flip it five times in a row, and it's all heads. Yes, could that happen? Sure, but the probability of that happening really small. You're no longer starting to believe that null hypothesis. Now, if the p value is low, so again, if that probability is really small, this implies that the sample we obtain from the population is extremely rare. Again, if we were to assume that null hypothesis is true, again think about it in that fair coin example, could a fair coin get five flips in a row that were heads. Yes, it could. It's probable. I'm sorry, it's possible. It's just not probable. Could it happen? Yes. Is it likely to happen? No. And so this leads us to question the validity of that null hypothesis and If we no longer believe the null hypothesis, we reject that null hypothesis if the p value is low enough. If the chances are small enough, we're no longer going to believe that null hypothesis is true. The question, though, is how low is low enough. How rare is rare enough? That's what we call the significance level alpha. If the p value is less than or equal to the significance level, the sample is so rare. We no longer believe the null hypothesis is true, and we call this rejecting the null hypothesis. So, you can think about it this way: going in, you think to yourself, you know what, I'm going to have him flip this coin, and if what he flips happens less than 5% of the time I'm no longer going to believe that that coin is fair, and so I flip a coin once, it's heads, 50/50 chance, I flip a coin twice, it's still heads, now there's only a 25% chance, I flip it three times, it's heads, there's a 12.5% chance, I flip it four times its heads, it's a 6.25% chance. I flip it five times its heads, and now it's a three point something percent chance. And you're like, that's it. I said, if this got less than 5% I would no longer believe it. You are below 5% that is my significance level. For example, I am no longer going to believe that null hypothesis that this is a fair coin. That's the idea. The significance level is sort of that definition of how rare is rare enough. The p value is the probability you actually saw what you did, so maybe it's rare, maybe it's not. We don't know that ahead of time. That's the whole idea. That's why we're testing these things. If it is rare enough, if the p value is smaller than the significance level, then we no longer believe the null hypothesis. We reject the null hypothesis. Now, there are two possible outcomes from a hypothesis test. Again, if the p value is smaller than or equal to alpha, we no longer believe the null hypothesis is true. We call this rejecting the null. However, if the p value is bigger than alpha, so again I flip the coin four times in its heads, there's only a 6% chance of that happening, but you said that if it had to be less than 5% then you'd no longer believe me. So, at 6% you're still good, you still believe the coin is fair. So, you do not have enough evidence to say the coin isn't fair, or you do not have
enough evidence to say the null hypothesis isn't actually true, so here we say we do not reject the null hypothesis. This goes back to that idea of guilty versus not guilty, like I mentioned in the last lecture. We don't actually say innocent at the end of a court case. We're not collecting data for innocence; we assumed innocence. We're collecting data for guilt. If you have enough data, if you have enough evidence for guilt, then you reject innocence. You claim guilty. If you do not have enough evidence for guilt, then you just say not guilty. You don't have enough evidence to reject innocence, so innocence still holds true. That's the idea. Let's try and view this, you know, on a graph. It may help a little bit. So, let's imagine we were testing the alternative hypothesis that mu was less than some number, so we're going to say that mu is equal to some number, that's our initial null hypothesis. So mu is going to equal some number. I'm going to collect a sample that is my sample, I'm going to represent it with t notice how my sample is far away, quote unquote, from the original null hypothesis, but how far away is far enough away. Well, I shaded in the probability that I got a sample that far away, or even further. Again, I'm testing what's the probability in my alternative hypothesis of being below this number, so the chances of me being smaller than or equal to t is that green shaded p value you see. Notice the original shaded area, alpha was the original shaded area. If my evidence was in that original shaded area, if it was smaller than that, then I'm no longer going to believe my null hypothesis. And that's exactly what happened. I'm rejecting the null hypothesis. My p value, that shaded area that's green is smaller than the blue shaded area, so I no longer believe the null hypothesis is true. Again, my sample is so far away from my initial thought, at least according to p value, I no longer believe it. What if it was this? On the other hand. In this scenario, my data was a lot closer to the null hypothesis. In fact, if we look at the green shaded p value here, it actually is rather large compared to the alpha. In fact, my p value is bigger than my alpha, so I don't have enough evidence to reject my null hypothesis. My p value is bigger than alpha. I said I had to be at least alpha far away, and I'm not, so my values are close together according to the p value, and because of that, I can't not believe the null. The null said innocent. I'm going to stay with innocence. I don't have enough data to prove otherwise. The same thing could be done on the other side. We could be testing if a number is bigger than the null hypothesis. Same idea holds true if it's really far away. If your data is really far away from your initial thought, the probability of seeing it is really small, and so it's so small it's smaller than that alpha level. Then you know what, we no longer believe that null hypothesis, because these values are far apart. So things don't have to be done just on one side of the distribution, lower or upper, they can be done on both. So, again, what if I had an alternative where I'm like, I don't know if it's going to be too high, and I don't know if it's going to be too low, but I just think it's different. Again, we can do the same idea. We just sort of have to split our p value into two pieces. We can say, well, maybe we're
high, maybe we're low, and so if we were to add the two pieces of our p value together, is that less than or equal to adding our two pieces of our alpha together? That's the idea. Again, imagine you had, I'm willing to be wrong 5% of the time. Imagine that if it's less than 5% I no longer believe it. Well, okay, but if you can say high or low, then that 5% could be split into two pieces, 2.5% on the low end, 2.5% on the upper end, and we'll do the same thing with our p value, we'll split that into two pieces as well, but again, the notion is these values are still far apart or close together. Are you close enough to the null to where you can't reject innocence, or are you far enough away from your original thought where you can reject that innocence, so again let's take a look at our hypothesis testing coin example. So we have our five flips of a coin. The probability of me getting five heads in a row is 3.125% again just over 3% you no longer believe the coin is fair, but could it be fair? Yes, it could be fair. It could. There's a 3% chance that it is fair, but you just don't believe it, because that number is so small, that's the idea, that's the idea of significance level in p values, the p value is the probability, the significance level is how small is small enough, where you no longer believe the null, so again, this significance level is something, in all honesty, you should do this before you even run the hypothesis test. It basically defines where the unlikely values of the sample are going to be under the assumption that the null hypothesis is true. Another name for this is called the rejection region, so if we were to look back, for example, at this chart here, the blue shaded area highlighted by alpha, that's the rejection region. Any data sample that's in that region, you're no longer going to believe the null, your data is too far away from the null, whereas here our data was not in that rejection region, our data was closer to the null than where that rejection region started. That's another way of thinking about this idea of significance level. There are a lot of typical things that people use: 1% 5% 10% Those are some typical values, so let's summarize the p value is the probability that you got your sample or a sample that's even crazier than the one you've seen, all under the assumption that the null hypothesis is true. Now, if that p value is low, that would imply that the sample that we obtained from the population is extremely rare if we are to assume the null hypothesis is true, but again, How rare is rare enough? That's where the significance level comes in. The significance level defines the unlikely values of the sample statistic, if that null hypothesis is actually true, and so if our p value is below the significance level, we reject our null. If the p values are greater than the significance level, then we don't have enough evidence to reject the null hypothesis. Awesome, so now we've worked through all the steps of the null hypothesis, and in our next lecture, we'll work through an example, but that is the end of this lecture, and I look forward to seeing you in the next one.