ExoLabz logo
support@exolabz.ca
Outliers: When Exclusion Is Legitimate and When It Is Not

Outliers: When Exclusion Is Legitimate and When It Is Not

One value in a set sits well away from the others. Removing it improves the result, and there is usually a reason available if one looks for it. The question of whether that removal is legitimate has a clear answer, and the answer turns almost entirely on when the reason was identified.

The only defensible basis

A value may be excluded when an assignable cause is identified that is independent of the value itself: a documented pipetting error, a well that was visibly contaminated, an instrument fault recorded at the time, a sample that was mishandled.

The independence requirement is the whole test. “This value is far from the others” is not a cause — it is the observation prompting the search. A cause found only because a value looked wrong, and accepted only because it explains that value, is not independent evidence.

This is the same logic that governs an out-of-specification investigation, described in out-of-specification results and retesting, where an assignable cause must be found rather than assumed.

Why statistical tests do not settle it

Several tests exist to flag outliers. They identify values improbable under an assumed distribution, and that is all they do.

Two limitations follow. The assumed distribution may be wrong — real biological and analytical data are frequently skewed or heavy-tailed, and a value that is improbable under a normal model may be entirely ordinary under the true one. And improbable is not the same as erroneous: an unusual value may be the most informative observation in the set, indicating a real subpopulation, a threshold effect, or a genuine instability in the material.

A test can tell you a value is unusual. Nothing in the arithmetic can tell you it is wrong.

What exclusion does to the statistics

Removing the most extreme value reduces the apparent variability, which narrows the error bars and lowers the p-value. Do this routinely and the reported precision of every experiment is overstated.

The effect compounds with small samples. In a set of three, removing one leaves two, from which a meaningful spread cannot be estimated at all. In practice this is where the temptation is strongest and the damage greatest, for the reasons in technical and biological replicates.

Deciding before rather than after

The practice that resolves this is procedural rather than statistical. Write down, before the data are seen, what will disqualify a run: a system suitability failure of the kind in system suitability, a control that falls outside a stated range, a plate with an edge effect, a culture at the wrong confluence.

A rule fixed in advance can be applied without the result influencing it. A rule invented afterwards cannot be, however reasonable it sounds — and reasonableness is not the constraint, since a reason can nearly always be constructed for any particular value.

Better options than deletion

  • Report both analyses. With and without the value, stated plainly. If the conclusion survives both, the question is moot; if it does not, that is the finding.
  • Use a method less sensitive to extremes. A median and an interquartile range describe skewed data without anyone deciding what to delete.
  • Increase the sample. One extreme value in twenty carries little weight; one in three dominates. This is the only remedy that adds information rather than removing it.
  • Investigate it. An unexplained extreme value is a lead. In analytical work it may be a contamination event of the kind in cross-contamination on the bench, which is worth knowing about regardless of what happens to the datum.

Excluding a run is different from excluding a value

Two decisions get called the same thing. Discarding an entire run because a control failed is a judgement about the whole experiment, made on evidence that does not come from the values of interest, and it is the defensible form.

Discarding one observation from within an otherwise acceptable run is a judgement about that observation, and the evidence for it is usually the observation itself. The first kind can be governed by a rule written in advance; the second rarely can, which is why run-level criteria are the ones worth defining and value-level exclusions the ones worth avoiding.

What a reader can see

Very little, which is the problem. A published figure shows the values that survived, and exclusions are frequently not mentioned. Where they are, the useful details are how many, on what basis, and whether the criterion was set in advance.

A paper stating “one outlier was removed” without a reason has described an action rather than justified it. A paper reporting the full dataset and noting an unexplained value has been more useful, even though the figure looks worse.

The position worth holding

Excluding data is sometimes correct and always costly, because it is the one analytical decision that cannot be checked by anyone downstream. The discipline that makes it defensible is entirely about sequence: criteria first, data second.

Where that sequence has not been followed, the honest options are to keep the value or to say plainly that it was removed after the fact and why — not to present the reduced dataset as though it were what the experiment produced.

Leave a Comment

Your email address will not be published. Required fields are marked *

*
*

Legal Disclaimer

The products offered by ExoLabz are intended solely for research purposes. These products are not for human consumption, are not intended for medical use, and have not been approved by the FDA or Health Canada for any therapeutic or diagnostic purpose. ExoLabz makes no claims regarding the safety, efficacy, or intended use of these products outside of a controlled research environment. By purchasing our products, you agree to use them strictly for scientific research and in compliance with all local laws and regulations.

GLP-1 15mg research peptide vial - ExoLabz Canada
0
    0
    Your Cart
    Your cart is empty