Out-of-Specification Results and What a Retest Can Prove
A result falls outside its limit. The obvious next move is to run the test again, and if the second result passes, to report that one. That sequence is the single most criticised practice in analytical quality work, and understanding why explains what a retest can and cannot establish.
The result exists once it is produced
The governing principle is that a measurement is data from the moment the instrument produces it. It cannot be un-produced by a later measurement that is more convenient, and a laboratory that discards results until one passes is not measuring, it is selecting.
This is not a matter of paperwork discipline. If a batch is genuinely borderline, repeated testing will eventually return a passing value by chance alone, and reporting only that value converts a coin flip into a certificate.
What an investigation is for
The formal response to a failing result is an investigation, and its purpose is to determine whether the result reflects the material or the measurement. It proceeds in a fixed order for a reason:
- Laboratory phase first. Was there an identifiable error — a mis-weighed standard, a mobile phase prepared wrongly, a failed system suitability sequence, a transcription mistake? The checks involved are the routine ones described in system suitability.
- An assignable cause must be found, not assumed. “The instrument must have drifted” is not a cause; a calibration record showing the drift is.
- Only then, the material phase. If no laboratory error is identified, the result stands as a property of the sample, and the question becomes whether the sample represented the lot — which is the subject of sampling plans and batch representation.
The three things a retest is not
A retest has a narrow legitimate role, and three illegitimate ones that look similar:
- It is not an appeal. A second result does not overturn a first. Where a laboratory error was identified and corrected, the retest replaces the invalidated result; where no error was found, the two results are both data and the material is out of specification on the balance of them.
- It is not a search. Testing until a passing value appears — sometimes described as testing into compliance — inverts the logic of the measurement entirely.
- It is not a substitute for a wider sample. Re-injecting the same vial tests the injection, not the batch. Re-preparing from the same aliquot tests the preparation. Only a new sample drawn from the lot tests the lot.
How many retests are permitted, drawn from where, and who authorises them, should be written into the procedure before any result is seen. A retest count decided after the first failure is not a rule, it is a negotiation.
Averaging, and when it hides the answer
Averaging a failing result with passing ones is legitimate only where the averaging was part of the method to begin with — a method specifying the mean of three preparations reports that mean, and individual values are not results in their own right.
Where the method specifies a single determination, averaging a failure with a later pass conceals the variability that is itself the finding. Two determinations of 97.2% and 99.0% averaged to 98.1% report a compliant lot and suppress the fact that the method, the sample or the material is varying by nearly two percent — a spread that dwarfs the uncertainty discussed in measurement uncertainty.
Out of specification, out of trend, and out of expectation
Three categories are worth separating, because only one of them breaches a limit:
- Out of specification — the result falls outside an acceptance criterion.
- Out of trend — the result is inside the limit but markedly different from the lot history. A purity of 98.2% against a history clustered at 99.4% passes and still signals that something changed, in the way described in lot-to-lot variation and trending.
- Out of expectation — an anomaly with no limit attached at all: an unexpected peak, a shifted retention time, a mass that matches nothing predicted.
The second and third never appear on a certificate, because a certificate reports one lot against its limits. They are visible only to whoever holds the history, which is one of the structural limits of reading a single report.
What the buyer sees, and does not
Almost none of this is visible downstream. A certificate reports the final result; it does not report how many determinations were made, whether any were invalidated, or what an investigation concluded. A perfectly clean certificate is compatible with a clean first measurement and with a long and well-documented investigation.
That is not a scandal — it is what a certificate is for — but it does mean that the confidence a certificate carries comes largely from whether the issuing laboratory has a procedure for this at all. Accreditation is the ordinary external check on that, within the scope limits described in what ISO/IEC 17025 accreditation covers, and the difference in incentives between testing arrangements is set out in third-party versus in-house testing.
The honest version of a failure
Where a lot genuinely fails, the available responses are limited and none of them are reporting a better number: the lot can be rejected, it can be reprocessed and retested as a new lot with a new identifier per batch and lot numbering, or it can be released against a different, lower specification with that specification stated plainly on the report.
The last option is entirely legitimate and much less common than it should be. A material honestly described as 94% against a 94% limit is more useful to a buyer than the same material described as 98% because the third determination said so.
