When is statistical significance insignificant?
It's about time we cared more about being insightful, and less about sounding right.
You want a number. A round, defensible, sounds-about-right number. Most importantly, a statistically significant number. You know, a proper number you can put in front of people to prove the point.
This is a sweeping statement, but I would probably guess that 95%* of the time I’ve been asked whether a number is ‘statistically significant’, it’s been by someone who doesn’t know what statistical significance means.
And most of the time, it’s also been the wrong question. Often to a point where it doesn’t even make sense.
This isn’t a dig at people who didn’t do stats or can’t remember a definition they learnt in a classroom 20 years ago. It’s a dig at the mindset we’ve got ourselves into, where it is more important to sound right, than be right.
Some patchy evidence to back that up - take a look at the chart below. The blue line in the chart below is google search trends for “how to sound smart”. The red line is “how to be smarter”.
So rather than using ‘statistical significance’ in a performative manner, when should it be used when it comes to customer research? And if we can’t or shouldn’t use ‘stat sig’, what should we use instead?
First, something that I think is obvious but for some reason, in todays increasingly ‘data-driven’ world, is increasingly forgotten: Sometimes it is simply impossible to get answers from data.
Perhaps because the right data simply hasn’t been collected. Or it’s impossible to collect the data in a non-biased way (ie every single non-compulsory survey).
But most often, because the data is being used to predict something that hasn’t happened yet. And context, it turns out, is actually quite important.
As Yogi Berra said, 'It's tough to make predictions, especially about the future.' And shedloads of historic data only makes it very slightly easier, as most studies into forecasting show.
A deeper piece on the shortcomings of data is for another post. So back to the matter at hand. This is not intended to be a stats class. A would strongly recommend a chat with Claude (or your preferred LLM) to get into the weeds of what statistical significance means. But I want to talk about what decisions you do have control over when it comes to customer research, and how you should approach it.
Significant does not mean significant.
The word borrows all of this authority from the everyday meaning of “significant,” which is “important”. It doesn’t, in statistics.
A p-value below 0.05 (ie ‘significant to a 95% level’) says something actually very narrow: if there were no real effect, you would see data at least this extreme less than five percent of the time.
It does not tell you the probability that the effect is real. It does not tell you how big the effect is, or whether it matters, or whether you should do anything about it, or indeed any sense of causality.
With a large enough sample you can make a difference so tiny it is effectively meaningless be “significant,” because significance is partly just a function of how many people you asked.
Significance is not a finish line.
When it comes to customer research, the question most often asked is ‘how many people do we need to survey to get statistically significant results’?
The thing is, nobody can tell you the answer to the question above, because it depends on more than just sample size. It also depends on properties we will only get from doing the survey itself (variance, size of the effect). And, of course, on our own threshold for ‘statistical significance’. So the question, clearly, is the wrong one.
What you can do, before you collect anything, is work out roughly how many people you would need to reliably detect an effect of a given size.
And that calculation has four inputs, none of which is the sample size itself:
The smallest difference you would actually care about. What difference between two treatments or groups would actually get you to change what you do? Sometimes it’s literally ‘any’. Sometimes it’s a lot, if the change is expensive (eg a fundamental change to a product).
How much the thing you are measuring varies. A measure that is all over the place needs far more people to pin down than one that barely moves.
Your tolerance for crying wolf, the false positive rate, (conventionally set at five percent and almost never questioned!).
Your tolerance for missing a real effect that is sitting right there, the thing statisticians call power, conventionally eighty percent and even less often questioned.
So when someone asks “how many people for significance,” the honest answer is “it depends,” and what it depends on is mostly a set of choices they have not made yet and often have not realised they need to make.
My point is simply this - statistical significance is, more often than not, the wrong stick to beat a researcher with. It isn’t a measure of insight, or ‘size of a relationship’, but is a function of lots of variables not at all related to how helpful the research actual is.
Don’t get me started on synthetic data
The deeper problem with synthetic data is that it collapses variance. Ask a model what someone thinks of a proposition on a five-point scale and it gives you 3.5, because 3.5 is the safest summary of everything it has ever read on the subject. Ask two thousand real people and you get a mess: a pile at 1, a pile at 5, a thin scatter in between, and a mean of 3.5 that describes almost nobody. The synthetic sample reproduces the central tendency and discards the shape of the distribution, which is where the information actually lives. Ie it assumes everyone of a certain description behave identically, so it clusters results in an unusual way.
Variance is actually very important, because with low variance, small differences show up as being statistically significant.
TLDR; Synthetic Research tends to exaggerate statistical significance, making the idea of statistical significance even more redundant when using synthetic data.
Above all else, statistical significance doesn’t tell you whether that will hold true tomorrow.
This, for me, is the crux of the issue.
Times, they are a-changing. Constantly. And in unpredictable ways.
Statistical significance answers one narrow question: if there were genuinely no difference, in this population, at this moment, measured with this exact instrument, how surprised should I be to have seen data like this? That is a useful question. It is a question about the past. It contains no information whatsoever about whether the effect will survive a different week, a different sample, a different country, or a barely different wording of the question. Times change, constantly, and in ways nobody has a distribution for.
The replication crisis is usually told as a story about fraud and p-hacking. Some of it was. But a great deal of it was simply this: effects that were real enough in one room in one year did not transport to another room in another year, because the context was doing more work than the variable was. Context is not noise around the effect. Very often, context is the effect. I studied behavioural economics at university, and seen paper after paper discredited largely as a result of the above.
So the question worth asking is not “is this difference significant?” but “is this genuinely insightful?”
We need to push harder to move more towards asking ‘is this genuinely insightful?’ - here are some thoughts on how to do exactly that.
5 questions to ask yourself to understand whether you have a genuine insight, without needing statistical significance.
1. Can you state the mechanism in one sentence?
A finding without a why is a coincidence with a decimal point. I used to call this the ‘mirage of precision’ - we have a tendency to assume that the more decimal places, the greater the accuracy, which is simply not true.
So if you cannot say why the 45-year-olds behave that way, you have no idea what would make them stop. Mechanisms are the only part of a finding that travels. Put another way - correlations are local and extremely context dependant; explanations are portable across time and place.
2. Did anyone in the room bet against it?
Before you reveal the numbers, make everyone write down what they think the answer is. Insight is the distance between the prior and the posterior, and if you never record the prior you can never measure the distance. Half of what gets presented as a finding is something the client already believed. Worth knowing. Not worth paying £80,000 for.
3. Does it survive translation?
Ask the same thing three different ways and see what is still standing. If the answer flips when you rephrase the question, you have measured a property of your questionnaire and not a property of a person. This is also the fastest way to catch the synthetic stuff, incidentally: real people are stubbornly consistent about the things they actually care about and wildly inconsistent about the things they do not.
4. Can you find the person it is not true for, and explain them?
Intentionally go looking for the disconfirming case. If your segment is real, the exception will have a reason, and the reason will make the segment sharper.
5. What happens on Monday?
If the result had come out the other way, would a single person have done a single thing differently? If not, the question was never worth asking, and no amount of significance will retrospectively make it so. This is the only test that most research would fail on its own terms, and it is the cheapest one to run.
The thread running through all six
None of these are statistical tests. Every one of them is a test of whether you understand the thing well enough to predict how it will behave when the context changes, because the context, by definition, constantly changes.
Significance tells you that you probably did not get fooled by this sample. For me, a genuine insight tells you what to expect from the next one, and why.
Which is, in the end, the whole job of genuine research.
This is why, at Dose of Reality, we are constantly building a bigger library of deep interviews. Not because we want to prove statistical significance, but because we want to increase the likelihood that users will find something genuinely insightful. Deep research at scale multiplies your chances of insight.
That is ultimately why we exist as a business. To help businesses make genuinely more informed, perhaps more inspired, decisions.
*this is a guess and not a statistically significant number, or indeed a significant, number



How wise...