Vlog

‘Flawed’ use of non-significant data ‘killing further research’

Papers still too quick to interpret results as showing no meaningful effect, ‘polluting’ scientific literature

Published on
August 11, 2026
Last updated
August 11, 2026
Source: Getty/malerapaso

Academic papers are routinely incorrectly treating “non-significant” results as proof that experiments had no effect, according to new research. 

A paper authored by researchers from the universities of Manchester, Oxford and Arkansas warns that there is a “widespread misinterpretation of non-significant results”, which may be killing off areas of interest in future research.

The study, published in PNAS, outlines that “non-significant results” in academic research are typically interpreted as meaning “no difference” or “no effect”, “despite long-standing recognition that this is a fundamental misinterpretation”, it says.

However, the interpretation remains “widespread”, featuring in about 50 per cent of research papers and conference presentations, the paper says.

Vlog

ADVERTISEMENT

This is one of the “most widespread and problematic misinterpretations in the scientific literature”, it explains, adding that a statistically non-significant result “shows only that the data do not provide strong evidence for a difference”. 

It says such a distinction matters because “findings can arise for two very different reasons: either there is no meaningful difference, or a meaningful difference is present but cannot be detected reliably because of limited sample size or high variability”. 

Vlog

ADVERTISEMENT

“Recognising the difference between ‘no evidence of a difference’ and ‘evidence of no meaningful difference’ can substantially improve how scientific results are interpreted, reported, and acted upon.”

It recommends that “equivalence testing” can be used instead to help make the distinction whether “the effect is too small to matter” or proves the evidence “remains inconclusive”, and therefore requires further research to determine its significance.

Co-author Jakub Tomek from the University of Oxford said that “absence of evidence is not evidence of absence” and using data in such ways is “really problematic”.

Tomek said: “This is a very basic issue, yet in life sciences, it routinely passes review and editorial judgement, and then it pollutes the literature because people say ‘there is no difference, this drug doesn’t have an effect’, which kills the development of further research.”

Vlog

ADVERTISEMENT

He said that “we did not detect a significant difference” often becomes shortened to “there was no difference”, especially when summarised in further research papers, “and that’s a huge shift”. 

David Eisner, professor of cardiac physiology from the University of Manchester, added: “Inevitably, some published results where people have concluded that there’s no effect will be wrong.”

He said he hoped their findings can help lead to a shift in data interpretation, and for academics to make it clear where there is evidence that something is not “biologically or medically important”, or whether “the data is so scattered that further work is required” to make definitive conclusions. 

Eisner said that this was only “one of several statistical issues which people need to pay more attention to” and forms part of a greater need to “be more careful with the use of statistics”.

Vlog

ADVERTISEMENT

juliette.rowsell@timeshighereducation.com

Register to continue

Why register?

  • Registration is free and only takes a moment
  • Once registered, you can read 3 articles a month
  • Sign up for our newsletter
Please
or
to read this article.

Related articles

Reader's comments (2)

I do not see a point in being so puritan about this - the analyses should be clearly reported, and scientists are trained to interpret the statistical results correctly regardless of how they are interpreted by the author(s). Being nitpicking about this often results in people disagreeing about phrasing - that 'we found no difference' is, at times, argued by puritans as if it were meant to say 'there is no difference' etc. Then it dwindles to reusing 'approved' stock phrases over and over again whenever a non-statistically significant result is mentioned. This then raises accusations of using GenAI to write the manuscript (i.e., the detection of common stock phrases). The priority is the precision and accuracy of statistical reporting, numerically speaking, and open science (e.g., data availability for verification and additional analysis) and not whether the author(s) phrased it this way or that. If scientists are that gullible as to just take the interpretation of the author(s) without critical reflection on the empirical justification of the interpretation, then it is the poor training of scientists that should be more concerning.
new
Thanks for engaging - this is Jakub, one of the quoted people responding here. There are multiple layers here - first, your shoulds are obviously right, but the "should" often remains not converted into "is", unfortunately. Yes, poor training is a huge issue, and that's why we wrote a book with David recently ("Basic Statistics for Life Scientists: A Concise Handbook of Essential Techniques"), trying to support non-technical readership in learning how to become good users of statistics in a way that they don't just give up and go for whatever gives them the p-value on the side of 0.05 they want at a given moment. A second point - I'd gently push back on the idea that worrying about "there was no difference" as being problematic is puritan. People often read shallowly (even if they >should< not), often don't have an intuition for study power etc. - if a paper claims "there was no difference", that's what many will take home from it. More importantly though, those claims of no difference remain reasonably easily checkable in the paper where they are used, but then the claims get propagated through a cascade of other publications that refer to them. At that point, it stops being a clumsy shortcut, and becomes a real problem. Finally - the paper that the article is linked to (now online here https://www.pnas.org/doi/10.1073/pnas.2611548123) focuses on equivalence testing, which helps discern between "no significant difference" because the real effect is very likely small, and between general uncertainty resulting from high variability and/or low sample size. We tried to be as constructive about this as possible - which is what I think is needed. I understand and agree that many things about current statistical practice are rather infuriating, but just saying people do things badly and should get better without a constructive way forward has not brought much improvement unfortunately.

Sponsored

Featured jobs

See all jobs
ADVERTISEMENT