About
A p-value of 0.0738 is not significant at the 0.05 level. It is also not nearlysignificant, marginally significant, or trending towards significance. It is a number on one side of a line the authors chose in advance. This site collects what gets written when it lands on the wrong side.
Where the phrases come from
The list is not ours. It comes from three published sources, and the 318 phrases drawn from them are marked as confirmed specimens throughout the site.
- Hankins M (2013). Still Not Significant. Probable Error. mchankins.wordpress.com
- Otte WM, Vinkers CH, Habets PC, van IJzendoorn DGP, Tijdink JK (2022). Analysis of 567,758 randomized controlled trials published over 30 years reveals trends in phrases used to discuss results that do not reach statistical significance. PLOS Biology 20(2): e3001562. doi.org/10.1371/journal.pbio.3001562
- Hankins M. Continuing specimens posted publicly on LinkedIn.
What this site adds is the evidence: those lists give the phrases, and this one shows each phrase in the sentences researchers actually published, with the p-value each one was describing.
How a sentence gets here
- Each phrase is searched against Europe PMC, restricted to open-access articles with retrievable full text.
- Every hit's full text is fetched and the phrase located in it. This step is not optional: Europe PMC's index ignores stopwords, so it returns identical results for not quite significant and quite significant. Roughly half of all search hits do not survive it.
- The sentence containing the phrase is extracted, along with any p-value reported inside that same sentence.
The corpus currently holds 514,629 sentences from 460,328 papers across 8,873 journals.
What counts as a specimen
Hankins' original post puts it plainly: results are either significant or not and can't be qualified. A significance test answers one yes-or-no question against a threshold chosen in advance, so an adjective attached to “significant” is not describing a degree, because there is no degree to describe.
Every sentence here falls into one of three kinds.
- A hedge approaches significance without arriving: nearly significant, a trend toward significance. A way of not saying “not significant”.
- A qualification attaches a degree to significance: highly significant, moderately significant. The qualification is the error whatever the p-value, so highly significant (p = 1×10⁻¹⁶) is the same category mistake as highly significant (p = 0.09), only less funny.
- A plain statement reports the result as it is: did not reach statistical significance (P = .977). That is correct usage and no kind of specimen, so it is filtered out of the browser. It is still in the corpus and the API will hand it over (
verdict=fair), because it becomes remarkable when the p-value says the result actually was significant.
One thing the site deliberately does not flag: p = 1. It looks impossible but is not, because discrete tests such as Fisher's exact routinely return it. A p-value of exactly zero really is impossible, since p is the probability of a result at least as extreme as the one observed, and that always includes the observation itself.
What this is not
It is not an accusation. Hedged language is often the honest description of a genuinely ambiguous result, and a hard threshold at 0.05 is itself a convention worth arguing with. What makes the collection funny is the gap between the arithmetic and the adjectives, not the people who wrote them. No author is named anywhere on this site, and every sentence links to its source so you can read what surrounds it.
Maintainer
Citing this collection
@misc{barely_significant_2026,
author = {Sofi-Mahmudi, Ahmad},
title = {Barely Significant},
year = {2026},
url = {https://barelysignificant.xera.ac},
note = {Database, accessed YYYY-MM-DD}
}Licensing
Sentences are quoted from open-access articles under the licences their publishers applied, with attribution and a link back to the source. The code and the phrase list are on GitHub. The data is available through the API.