I do not think that word means what you think it means.
Let’s start with a question from a reader, one that I have heard many times before:
"I want to know how to prove that something is statistically significant without knowing statistics. How can I evaluate survey results better, starting from zero knowledge of stats?"
With evaluation, there's an inclination to toss around statistical jargon with very few people in the room understanding what it means. And – more importantly – what it doesn't mean.
Here's what's going on:
The statistical language makes things feel weighty. It’s the non-profit equivalent of using “gravitas” to describe a leader. But this sets in motion an anxiety-fueled quest that pursues “Statistical Significance” as little more than a phrase that has been handed down as some kind of Holy Grail.
Meanwhile, everyone can only vaguely guess at what it means. They just know they've been told it's IMPORTANT.
Now, I’m not saying all statistics or statistical language is bullshit.
Statistical tests are just a tool. They are not imbued with magical powers. They do not confer authority. What matters is if they are relevant to the job at hand.
For a non-math metaphor: Imagine using a hammer on a screw. It's a tool. It might do something. But the process is harder than necessary and not as helpful as it should be.
Let’s give you a guide for navigating this jargon as a non-stats person living in a stats-crazy world:
- The real meaning of significance
- When you actually need to worry about it
- When you super-duper do not
- What you (or your boss/funder) might actually mean
- How to talk about it (to said boss/funder)
What does significant even mean?
Simply: statistical significance means that the result you found is more likely caused by something than due to the whims of random chance.
That’s it.
Significant does not mean big.
It just means not random.
When we see a difference between two numbers, we get excited. But the world is a messy, random place. Should we be excited?
Statistics look at the relationships in the numbers and discern – mathematically – how likely it is that the difference is caused by something other than random variation.
When do you actually need to worry about “statistically significant”?
For you, Reader, I’m going to go out on a limb and say: Very, very rarely.
95% of the time* evaluation in our field is best served by descriptive statistics:
- The proportion of people who did or said X
- The average score or rating of people on Y
- The highest and/or lowest ratings achieved
- The distribution of ratings across a spectrum of categories
*As always, I 100% made up that stat. And it’s super-significant.
Descriptive data points tell you what you wanted to find out. They point you in a direction. They are, most of the time, your answer.
For “statistically significant” to have any use, you need to test your result against something else.
Not to be all science-nerd about it, but you need a hypothesis.
Do you have a hypothesis about a relationship? If you can’t name what you are trying to test, compare, or “prove”, you can pass Go. You can collect $200. You already have the stats you are looking for.
Practical examples of when you might have a hypothesis:
- Program experience will cause students to improve their pre to post scores.
- Families will spend more time in the exhibit than adult-only groups.
- Teachers who attend a week-long PD will feel more confident than teachers who attend one-day PD.
You might have a testable situation if you are comparing time periods or comparing groups.
I’ve been doing this for 20+ years, and in most evaluation scenarios where I whipped out the ol' stats machine? It fell into one (or both) of these comparison buckets.
In theory, you could compare your descriptive statistics against some mythical “norm” or benchmark. But seriously, when was the last time you collected data that had some big ol’ benchmark dataset behind it?
Here in non-profit education land, that’s laughable. We have teeny tiny datasets about very specific impacts from very specific programs. This isn’t epidemiology. (I assume they have big datasets. I don’t actually know.)
Which brings me back to my main point: So very often, it’s not even possible to “run stats” on the data you have. Focus on exploring and presenting your descriptive stats well to maximize what they can tell you.
That’s most important.
In a lot of program evaluations, statistics don't apply.
What I see the most is a program that invited everyone involved to give data.
Fun fact about all those stats? They all assume your data is a random sample from a much larger population. The math? It's built to gauge the risk that your results are because you randomly picked up some wacko data points.
If you tried to get data from everyone in a program? You have a census, my friend. You don’t need stats. You need response rate.
The data you see? That’s it. That’s everyone.
The stuff they liked? That’s what they actually liked!
The stuff they learned? That’s what they actually learned!
And, if your Spidey Sense says that all of your cranky participants boycotted your survey? Stats can't solve that anyway.
What you might mean is margin of error.
A question I have heard more than a few times is this. (In other words, if you’ve said or thought this, you are in extremely good and smart company.)
“How do we get a statistically significant sample size?”
That’s not, technically speaking, a thing.
Having a bigger sample size makes it easier to detect “statistical significance” about a hypothesis. To the math, bigger sample = “less chance this pattern is due to random whims of chance.”
But when I hear anxiety about “significance, ” a lot of times people actually mean: Can I trust that these descriptive statistics reflect reality?
There's a way to math this question. But the mathematical margin of error is pretty conservative. Most calculators are going to tell you to get an obscene sample size for a small margin of error and high "confidence".
(Sidebar: "Confidence" is another word that doesn't mean what you think it means, when it comes to stats.)
It's just not realistic for most of you.
Instead, I suggest you apply my Common Sense Margin of Error when reviewing descriptive statistics.
I use the same principles as the math-y version: The fewer people you talk to, the bigger the chance capturing a random oddball will sway your picture. But because most programs get likeminded (not polarized) responses, strong patterns are unlikely to be from random oddballs.
My Common Sense Margin of Error Rules:
- Don’t hold onto any single number as Gospel. Consider it to be pretty close.
- If you see a number that is strikingly higher, lower, or different than other data, pay attention! That means something!
- View small differences in numbers with a large grain of salt. Consider close numbers as basically equivalent.
- The smaller your sample, the more grains of salt you need.
Scripts for when someone asks YOU for stats and you’re not a stats person.
It’s all well and good for me to tell you these things, but WTF do you actually say, if it’s your boss, funder, or colleague asking these questions?
“This dataset includes everyone in the program – we did a census of all our participants. So, this is a complete picture of what resulted.”
“I’m not sure I’m clear on the hypothesis we’d be looking to test statistically. Can you help me identify the relationship we'd be trying to test for here?”
"Point X is substantially higher/lower/different than the rest of the data, which is a real pattern and not the result of noise or error."
“To do X would require statistical expertise, and possibly software, that is beyond our in-house skills. If it’s critical to add that analysis to these descriptive statistics, we’d need to contract with an outside expert to do that piece.”
And if it’s your own Imposter Syndrome insisting you must need Greek letters:
Tell your Imposter Syndrome to read this issue. If that doesn’t work, tell it to call me.
Real World Example:
Timing and Tracking is a type of study we’ve done again and again in exhibitions in museums, zoos, and aquariums.
Because it’s a method that focuses on human behavior, there is a lot more variation than in something like a survey. (Humans. Are. Weird.) It’s one reason we try to get more observations than other methods. There's more "noise" in the data, so we minimize the effect of any one oddball visitor by getting more data.
At the end of a study, we focus largely on the descriptive results. That picture of what people actually did is so valuable. Applying a Common Sense Margin of Error is plenty.
But often, we also formulate some hunches along the way. Maybe: Did the adult-only groups spend longer in the exhibit than family groups?
When we run the averages, we see numbers that look promising. Maybe an average stay-time is 30 seconds or a minute higher for all-adult groups.
But because that average is from a set of observation data that was so very noisy, a statistical comparison can offer a mathematical view on the question: Is the difference in the aggregate a real pattern? Or is it due to a few extreme cases?
It helps us interpret the pattern and the noise. But even more important for our clients? Visualizing that comparison in a way that supports the result. Not just relying on some mysterious p-value like it's a Holy Grail.
Did this raise questions for our next Stats Therapy group session? Hit reply and tell me.