Technical and Biological Replicates Are Not Interchangeable
The methods section says n = 3. Whether that means three independent experiments or three wells from one preparation determines what the statistics mean, and the two are reported identically far more often than they are distinguished.
The distinction
A technical replicate is a repeated measurement of the same biological sample. Three wells filled from one dilution, three injections of one vial, three readings of one plate. It measures the precision of the measurement process.
A biological replicate is an independent instance of the thing being studied. Separate cultures, separately treated, ideally on different days from different passages. It measures the variability of the system.
Both are useful and they answer different questions. The error arises when technical replicates are counted as the sample size for a claim about biology.
Why the substitution inflates confidence
Technical replicates are, by construction, more similar to each other than biological replicates are. They share a preparation, a dilution, an operator and a moment.
Treating them as independent observations therefore does two things at once: it understates the variability, because the shared sources of variation are invisible within the group, and it overstates the sample size. Both push the standard error down, so the error bars shrink and any test becomes more likely to clear a threshold — the mechanism described in error bars.
The result is a figure that looks well determined and a p-value that looks convincing, neither of which reflects how reproducible the finding actually is.
What counts as independent
The test is whether the replicates share a source of variation that the claim is supposed to average over. Some cases are clear and some are not:
- Three wells, one dilution, one plate. Technical, unambiguously.
- Three plates, one dilution, one day. Still technical with respect to the preparation — the dilution error, and any adsorption loss of the kind in adsorptive loss to surfaces, is common to all three.
- Three separate cultures treated from three separate dilutions on three days. Biological.
- Three separate cultures, all treated from one stock dilution prepared once. Partly shared — the biology is independent, the compound preparation is not. If the preparation is wrong, all three are wrong together.
That last case is the common one and the awkward one. It is stronger than a technical triplicate and weaker than a true biological triplicate, and the honest description is to say what was shared.
How each should be handled
Technical replicates are averaged. Their mean becomes one observation, and their spread is a quality check on the measurement rather than data about the system. If technical replicates disagree substantially, that is a problem with the assay — pipetting, per pipetting accuracy, or plate position effects — and it should be fixed rather than averaged away.
Biological replicates are the observations. The sample size is the number of them, and the statistics are computed across them.
Stated compactly: technical replicates improve the estimate of each point; biological replicates are what you count.
Why three is the number, and what it costs
Three is conventional rather than derived. It is the smallest number that permits a variance to be estimated at all, and it estimates it badly.
The consequence is that a three-replicate experiment has low power, which means that the effects it detects are the ones that happened to look large — the selection problem set out in what a p-value does and does not tell you. Increasing biological replicates addresses this; increasing technical replicates does not, because they do not add information about the variability that limits the conclusion.
Pseudoreplication, which is the formal name
Counting non-independent observations as independent has a name in the statistical literature: pseudoreplication. It was described decades ago in ecology, where the equivalent error was treating repeated measurements from one site as replicate sites, and it is the same mistake wherever it appears.
Naming it matters because it makes the remedy visible. The fix is not a different test or a correction factor — it is to analyse at the level the claim is made about. If the claim concerns cultures, the analysis is across cultures, with the within-culture measurements averaged first. Where both levels genuinely carry information, a model that represents the nesting explicitly is the correct tool, and averaging is the simple approximation to it.
Reading a methods section for it
The question to ask is what one observation was. A methods section stating “experiments were performed in triplicate” has not answered it — triplicate wells and triplicate experiments are both described that way.
What resolves it: whether replicates were performed on separate days, whether separate cultures or passages were used, whether the compound was diluted separately for each, and whether the reported n refers to wells or experiments. A paper that specifies these can be evaluated; one that does not has left the central question of how reproducible the result is unanswered, which is the broader problem in why two laboratories get different results.
The same structure in analytical work
This is not a biology-specific idea. Re-injecting one vial three times measures the instrument. Preparing three solutions from one weighing measures the preparation. Sampling three times from a batch measures the batch — the distinction underlying sampling plans and batch representation.
In both settings the discipline is identical: decide what the claim is about, and count only the replicates that vary in the way the claim needs them to.
