The Reagent Problem in Reproducibility: When the Material, Not the Method, Is Wrong
Misidentified cell lines and unvalidated antibodies turned part of the reproducibility debate from a statistics argument into a materials argument. Here is the evidence, where synthetic peptides sit in it, and what a buyer actually controls.
Part of the reproducibility problem has nothing to do with statistics or procedure: the experiment was run correctly on the wrong material, or on material that differed from what the next laboratory used. Freedman and colleagues estimated the cumulative prevalence of irreproducible preclinical research at more than 50%, costing roughly US$28 billion a year in the United States alone, and named biological reagents and reference materials as one of four categories of error, alongside study design, laboratory protocols, and data analysis and reporting [2]. The best-documented cases — misidentified cell lines and poorly characterised antibodies — are material failures no amount of careful statistics could have caught.
This article covers that materials side: the evidence, where synthetic peptides sit in it, and what a buyer can control. How to report a compound's source in a paper is the subject of this cluster's pillar; the argument here is why that report matters.

How big the problem is, and how much of it is materials
The widest view comes from a survey. In 2016 Nature asked 1,576 researchers about reproducibility: more than 70% said they had tried and failed to reproduce another scientist's experiments, more than half had failed to reproduce their own, and 52% agreed there was a significant crisis [1]. A survey measures perception rather than frequency, but the perception was broad and consistent across fields.
The economic estimate is more specific about causes. Freedman, Cockburn and Simcoe put United States preclinical research spending at around US$56 billion a year and estimated that more than half of it — a midpoint of 53.3% — goes on work that cannot be reproduced; they then weighted the error sources into four categories, one of which is biological reagents and reference materials [2]. The figures are built from published prevalence ranges rather than direct measurement, and the authors present them as estimates. The part to keep is not the dollar total but the category: materials are a named source of failure, separate from how an experiment was designed or analysed.
Misidentified cell lines: the best-documented case
Cell lines are the clearest example because the failure is binary and detectable. A line distributed as one tissue of origin turns out, on DNA profiling, to be another — often a fast-growing line that overgrew the original. The International Cell Line Authentication Committee maintains a register of such lines: version 14, released in February 2026, lists 608, of which 560 are misidentified with no known authentic stock [4].
The effect on the literature is large. Horbach and Halffman traced 32,755 articles reporting research on misidentified cells, cited in turn by an estimated half a million further papers, and found that the contamination of the literature was not declining over time [3]. Many of those papers may have been run and analysed flawlessly. The material simply was not what its label said.
Precise identification measurably helps. Babic and colleagues found that papers naming their cell lines with Research Resource Identifiers reported problematic lines less often than papers that did not [7]. Looking a resource up in a curated registry puts any warning attached to it in front of the author before the experiment, not after publication.
Antibodies: what validation revealed
Antibodies repeated the pattern. In 2015 Bradbury and Plückthun, with more than a hundred co-signatories, argued in Nature that poorly characterised binding reagents were wasting research money and damaging reproducibility, and that such reagents should be defined by their sequence and produced recombinantly [5]. The following year an international working group proposed five conceptual pillars for validating an antibody in a given application: genetic knockout or knockdown, an orthogonal antibody-independent measurement, independent antibodies against the same target, expression of a tagged protein, and immunocapture followed by mass spectrometry [6].
Two lessons carry over to any reagent. Validation is application-specific: a reagent that performs in one assay may fail in another [6]. And identity has to be pinned to something that does not drift — a sequence rather than a catalogue name that can later be attached to a different product. Vasilevsky and colleagues found that 54% of the research resources in their sample, antibodies prominent among them, could not be uniquely identified from the papers that used them [8].
Where synthetic peptides sit in this picture
Synthetic peptides start from a better position than cell lines or antibodies. They are chemically defined: the intended structure can be written down exactly, and identity can be checked against it by mass spectrometry. There is no passage number, no biological drift, no clone to confuse with another.
That advantage is easy to overstate. A peptide's identity can be confirmed while the material is still not what an experiment assumes. Solid-phase synthesis produces deletion and truncation sequences that can sit very close to the target, sometimes differing by a single residue; purification leaves a counter-ion behind; the lyophilised solid holds water. And the analytical limits that make peptides hard to test mean a single certificate rarely settles all of those questions at once.
Batch-to-batch variation: the under-reported source of variance
For a defined compound, the realistic material risk is not misidentification but variation between lots. Two batches of the same sequence can differ in purity, in which impurities make up the remainder, in the trifluoroacetate or acetate counter-ion left from purification, and in water content. Each changes how much of the intended peptide a weighed milligram contains, and some impurities or counter-ions can carry activity of their own in a sensitive assay.
This variance is under-reported because it is invisible in most papers: few methods sections record a lot number at all. If more than half of resources cannot be identified even at the product level [8], almost none can be traced to the batch. When a result later fails to replicate, the lot is the one variable nobody can go back and check.
| Failure mode | Biological reagents | Synthetic peptides |
|---|---|---|
| Identity error | Misidentified or cross-contaminated lines; wrong clone | Wrong sequence — uncommon, and detectable by mass spectrometry |
| Drift over time | Passage-dependent change; genetic drift | None in the dry solid; degradation once in solution |
| Lot-to-lot variation | Antibody batches differing in specificity | Purity, impurity profile, counter-ion, water |
| Registry to check against | ICLAC register, Cellosaurus, Antibody Registry | None — the written sequence is the reference |
| Main defence | Authentication and application-specific validation | Lot records, lot-specific certificates, retained material |
What a buyer can actually do about it
Most of the reproducibility literature addresses journals, funders and suppliers. A buyer's leverage is smaller but real, and it sits almost entirely in documentation.
- Buy by lot, and record it. Transcribe the lot string at receipt and carry it into every notebook entry that uses the material.
- Keep the certificate for that lot, not a generic one, and note what it does and does not measure.
- Retain a sealed portion of each lot until the work is published, so a later question can be answered by measurement rather than memory.
- Where a result depends on a new lot, run a bridging comparison against the previous lot before relying on it.
- Report the lot, the purity method and the concentration basis in the methods, so the next laboratory knows which variable to control [2][8].
None of this makes a peptide more reproducible in itself. It makes the variance visible, which is the precondition for anyone — the original authors included — controlling it. The surveyed researchers who could not reproduce their own experiments [1] are a reminder that the laboratory most likely to need the lot record later is the one that failed to keep it.
References
- 1,500 scientists lift the lid on reproducibilityNature, 2016
- The Economics of Reproducibility in Preclinical ResearchPLOS Biology, 2015
- The ghosts of HeLa: How cell line misidentification contaminates the scientific literaturePLOS ONE, 2017
- Register of Misidentified Cell Lines (version 14)International Cell Line Authentication Committee (ICLAC), 2026
- Reproducibility: Standardize antibodies used in researchNature, 2015
- A proposal for validation of antibodiesNature Methods, 2016
- Incidences of problematic cell lines are lower in papers that use RRIDs to identify cell lineseLife, 2019
- On the reproducibility of science: unique identification of research resources in the biomedical literaturePeerJ, 2013
