ExoLabz logo
support@exolabz.ca
What a p-Value Does and Does Not Tell You

What a p-Value Does and Does Not Tell You

A result is reported as significant at p < 0.05 and the finding is treated as established. The p-value did not say that. It answered a narrow question about one dataset under one assumption, and almost every way the number gets used in practice asks it to answer a different question instead.

What the number actually is

A p-value is the probability of observing data at least as extreme as what was observed, if the null hypothesis were true. The conditional clause carries the whole meaning.

It is therefore a statement about data given a hypothesis, not about a hypothesis given data. Those are different quantities and they are not interchangeable. A p-value of 0.03 does not mean there is a three percent chance the result is a fluke, and it does not mean there is a ninety-seven percent chance the effect is real.

The four misreadings that do the damage

  • Treating it as the probability the hypothesis is false. It cannot be, because it was calculated assuming the null is true.
  • Treating 0.05 as a boundary between real and unreal. The threshold is a convention, chosen arbitrarily and defended mostly by habit. There is no mechanism that makes 0.049 meaningful and 0.051 meaningless.
  • Treating a non-significant result as evidence of no effect. Absence of evidence arrives the same way whether the effect is zero or the study was too small to see it. Distinguishing the two requires a confidence interval, not a p-value.
  • Treating a smaller p-value as a bigger effect. It is not a measure of size. A trivial effect measured precisely produces a very small p-value; a large effect measured crudely may not.

Why the effect size is the number that matters

The question worth asking is how large the difference is and how precisely it was determined. A p-value answers neither.

An effect size with a confidence interval answers both at once, and it degrades gracefully: a wide interval that includes zero says the study was uninformative, which is more useful than a bare “not significant”. This is the same logic as reporting an analytical result with its uncertainty rather than as a bare figure, set out in measurement uncertainty.

Where the multiplicity problem comes in

Run twenty independent comparisons with no real effect anywhere and, at a 0.05 threshold, one will come out significant on average. This is arithmetic, not misconduct.

It becomes a problem when the comparisons are run and only the significant one is reported, or when the hypothesis is formed after seeing which comparison worked. A study testing several concentrations, several time points and several readouts is running many comparisons whether or not it describes them that way, and the number of tests performed is part of what a reader needs.

The corrections available for this are conservative and blunt. The more useful practice is to state in advance which comparison is the primary one, and to treat the rest as exploratory — which is also what makes a result worth trying to reproduce, in the sense of why two laboratories get different results.

The peptide literature is unusually exposed to this

Studies in this area tend to be small, to use several readouts, and to be published individually rather than as part of a programme. Each of those raises the proportion of published positives that will not hold up.

Small studies have a specific pathology worth naming: when a study is underpowered, the only results that clear a significance threshold are the ones where the observed effect happened to be large. A literature of small studies therefore over-states effect sizes systematically, even when every individual study is honest. That is part of what makes the reading discipline in what the published literature actually shows necessary.

What to look for in a paper

Whether the effect size is reported at all, or only the p-value. Whether an interval accompanies it. How many comparisons were made in total, not just how many were reported. Whether the sample size was decided in advance. And whether the analysis described is the one that was planned, or one arrived at afterwards.

A paper reporting “p < 0.05” with no effect size and no denominator has given a verdict without the data behind it — the same structural problem as a certificate reporting “complies” instead of a value, described in how a specification limit gets set.

The honest summary

A p-value is a weak, narrow, conditional statistic that has been asked to carry the entire weight of scientific conclusions for decades. It is not useless — it does flag that a pattern is unlikely under a specific null model — but it is one input among several, and the smallest of them.

The practical position is to read effect sizes and intervals first, treat the p-value as a footnote, and treat a single significant result in a small study as a reason to look further rather than as a conclusion.

Leave a Comment

Your email address will not be published. Required fields are marked *

*
*

Legal Disclaimer

The products offered by ExoLabz are intended solely for research purposes. These products are not for human consumption, are not intended for medical use, and have not been approved by the FDA or Health Canada for any therapeutic or diagnostic purpose. ExoLabz makes no claims regarding the safety, efficacy, or intended use of these products outside of a controlled research environment. By purchasing our products, you agree to use them strictly for scientific research and in compliance with all local laws and regulations.

GLP-1 15mg research peptide vial - ExoLabz Canada
0
    0
    Your Cart
    Your cart is empty