“I don't understand why a T-statistic needs the data to follow a normal distribution.”
asked on Stack Exchange · 2,975 views
Statistics is unusual in that you can pass a course by learning which test to run rather than why that test is the right one — look at the variable types, pick from the flowchart, read the p-value off the table.
That works until the question arrives as a situation rather than a recipe. The flowchart doesn't cover 'my sample is large, does that mean I can stop worrying about normality?'
The questions people actually ask are almost never about a formula. They're about when the formula applies — and that is a different skill from being able to compute with it.
Predict before you check
A colleague says: "Our sample is 500, so we don't need to check whether the data is roughly normal — the central limit theorem covers us." What is actually wrong with that?
Knowing n > 30 is a recognition cue. Knowing what the CLT is a statement about is what lets you decide whether the rule applies at all — and those are separate skills. A correct answer here doesn't prove the second one, which is why Vectra keeps testing the distinction long after the first answers turn green.
The gap underneath the gap
Almost every question in this cluster — why a t-statistic assumes normality, why Type I error doesn't shift with sample size, why power relates to the normal distribution at all — is really a question about one object: the sampling distribution of a statistic. Learners who never built a picture of that object end up memorising which test goes with which situation, which is precisely the knowledge that evaporates when the situation is described in words instead of handed over as a table. Backing up to the sampling distribution is where a real diagnosis starts when this is what's underneath — because practising more tests assumes the gap is the testing, and here it isn't.
it doesn't move on until you get it.
Not ready to sign up? Say what you're stuck on and Vectra returns the prerequisites most likely to be responsible — each one graded by how much is actually known about it. In this cluster the usual culprit is the sampling distribution, not the test you were about to reach for.
Or find what's actually missing, free →