Research · Evidence and Explanation

Replication, Robustness and Scientific Confidence

Scientific confidence should not rest on one study alone. It grows when findings recur, conclusions survive reasonable alternatives and different evidential approaches support compatible claims.

By Yona Ole Lobulu ·

Research note10 min readFoundationalD2.13

Topic
Evidence and Explanation
Read first
3 pieces should be read before this one
Reading time
About 10 minutes of reading
Difficulty
Advanced reading: this page assumes a fair amount of earlier reading

The question

How should replication, robustness and converging evidence change confidence in a scientific claim?

Definition

Scientific confidence is the degree of epistemic support the available evidence warrants for a specified claim. It grows when findings recur, when conclusions survive reasonable alternatives and when meaningfully distinct approaches support compatible conclusions.

A study reports that a psychological intervention improves a behavioural outcome. The study is well designed, the result is statistically significant, and the paper survives peer review.

How confident should we now be that the intervention works?

The answer cannot be read from any one of those facts alone.

A study can provide important evidence while still being affected by sampling uncertainty, measurement limitations, analytical choices or features of the population and design. Peer review can provide scrutiny before publication, but publication itself does not transform an individual result into established knowledge.

Scientific confidence therefore depends on more than whether a study produced a convincing result.

Scientific confidence should usually rest on the relevant body of evidence rather than an isolated finding.

Confidence attaches to claims, not papers

The same study can support different claims to different degrees.

Suppose researchers report that X and Y are associated. They might also argue that X causes Y, or that a particular mechanism explains how X produces Y.

Those are not equivalent conclusions. The distinction between correlation, prediction, causation and mechanism governs what any single study can support.

Evidence can strongly support the association while providing little evidence about causation. Evidence can support a causal effect while leaving its mechanism uncertain.

So asking whether a paper is simply "reliable" or "unreliable" misses something important.

The relevant question is:

How strongly does the available evidence support this particular claim?

Here, scientific confidence means the degree of epistemic support the available evidence warrants for a specified claim. It is not a universal numerical probability attached to a paper or finding. It presupposes the prior question of what counts as evidence for the claim being made.

Confidence attaches to a specified claim, not simply to a paper.

This distinction matters because every subsequent test—replication, robustness or convergence—can strengthen only what it actually tests.

Replication asks whether evidence recurs

Replication exposes a finding or claim to another empirical test.

In this Research Note, replication refers broadly to testing a scientific finding or claim again using new observations, while recognising that replications vary in how closely they reproduce the original procedure.

This is different from reproducibility.

Here, reproducibility asks whether the reported analysis can be regenerated from the same underlying data and analytical procedures. Scientific disciplines do not use these terms completely uniformly, but separating the questions is useful:

Reproducibility: Can the reported analytical result be regenerated?

Replication: Does compatible evidence recur with new observations?

Some replications remain relatively close to the original procedure. Others change meaningful aspects of the method or operational definition while testing the broader claim. These are often described as relatively direct and conceptual replications, respectively. Neither form is universally superior; they expose claims to somewhat different tests.

Nor is there a universal criterion for declaring a replication successful. Compatibility has to be judged relative to the claim, design, effect estimates and uncertainty involved.

That is why replication should not be treated as a binary truth test.

A successful replication does not mean: the claim is now proven.

A failed replication does not mean: the claim has been disproven.

A successful replication can increase confidence that the relevant pattern recurs under the conditions tested. A well-designed and precise failure to reproduce an expected effect should normally reduce or complicate confidence.

Replication can also revise what we think the magnitude of an effect is. An effect may recur while proving substantially smaller than its original estimate.

Confidence that an effect exists is therefore different from confidence about how large, important or practically consequential that effect is.

Replication outcomes are evidence, not binary verdicts on truth.

Robustness asks whether the conclusion survives reasonable alternatives

Replication introduces new observations.

Robustness tests another vulnerability:

Does the conclusion survive reasonable alternative choices?

Research requires decisions. Investigators may need to decide how to process measurements, which observations to exclude, which variables to account for, how to specify an analysis or how to operationalise an outcome.

Many such decisions are legitimate.

The problem is not that researchers make choices. The question is whether the scientific conclusion depends narrowly on one choice when other theoretically defensible and statistically valid alternatives address the same question.

If a result persists across reasonable alternatives, it is more robust than one that appears only under a narrow specification.

But "reasonable" cannot simply mean: the analyses that preserve the preferred conclusion.

Which alternatives deserve consideration must themselves be justified by the scientific question, theory, measurement and statistical assumptions.

And robustness has limits.

A conclusion can survive every alternative tested while remaining vulnerable to an assumption or source of error that none of those alternatives addressed.

Robustness is evidence against dependence on the alternatives tested, not proof that every important source of error has been eliminated.

This is why replication and robustness are complementary.

A result may recur in new datasets but depend on a fragile analytical decision. Another may survive many analyses within one dataset but fail to recur when tested on new observations.

Replication and robustness test different vulnerabilities.

Converging evidence tests a claim through different routes

Confidence can also grow when a claim survives tests that do not all fail in the same way.

This is the logic behind converging evidence and triangulation.

Imagine that compatible conclusions emerge from behavioural observation, self-report, physiological measurement, longitudinal evidence and experimental manipulation.

The strength does not come simply from having more methods.

It comes partly from whether those methods expose the claim to different major sources of bias and error.

If several studies use the same flawed measurement or depend on the same problematic assumption, their agreement may reproduce the weakness along with the result.

But when meaningfully distinct approaches with different vulnerabilities support compatible conclusions, one shared artefact becomes a less sufficient explanation for the overall pattern.

Convergence is most informative when supporting methods do not share the same major vulnerabilities.

This does not require identical results.

Different methods can measure different aspects of a phenomenon, estimate different quantities or apply to different populations. What matters is whether their findings are compatible with the broader claim once those differences are understood.

Nor does convergence guarantee correctness. Partially independent approaches can still share assumptions, blind spots or previously unrecognised sources of error.

It strengthens a claim by exposing it to different tests—not by making error impossible.

More studies do not automatically mean more evidence

Ten papers do not necessarily represent ten meaningfully distinct tests.

Studies can share datasets or overlapping participants; measurement instruments; populations; analytical assumptions; research programmes; underlying sources of bias.

This matters because scientific confidence cannot be calculated from citation count alone.

Publication bias can distort the visible evidence base further. If some kinds of results are more likely to become visible than others, the published literature can appear more consistent than the underlying research process actually was.

Transparency helps by making methods, data, materials and analytical decisions easier to inspect, reproduce and challenge.

But transparency does not guarantee correctness.

Nor does statistical significance.

A statistically significant result can contribute information within an analysis, but significance alone does not establish measurement validity, replication, robustness, causation or overall scientific confidence. Interpretation also depends on the estimated magnitude and uncertainty of the effect, not merely whether a statistical threshold was crossed.

Evidence accumulation depends on quality and meaningful independence, not citation count alone.

Repetition cannot repair the wrong inference

Replication becomes especially easy to overinterpret when the underlying problem sits somewhere else.

Suppose several studies use the same measure and repeatedly obtain the same result.

That recurrence may substantially increase confidence that the measured pattern is real.

But if the measure does not validly represent the broader construct being claimed, merely repeating the same measurement cannot solve that problem. This is a question of measurement validity, not of recurrence.

Repeating the same measurement cannot by itself repair a validity problem shared across the replications.

The same principle applies to causal inference.

Suppose ten observational studies repeatedly find that X and Y are associated.

The accumulated evidence may make us highly confident that the association recurs.

It does not follow that X causes Y.

Other designs or complementary identification strategies may strengthen a causal interpretation, but repetition of the association alone does not make that transition.

The same boundary applies to mechanism.

Repeated evidence from designs capable of supporting a causal effect can strengthen confidence that the effect occurs while leaving the mechanism producing it unresolved.

Replication therefore does not automatically upgrade: association → causation

or: causal effect → mechanism

Replication strengthens only the kind of inference the underlying evidence can support.

Scientific confidence is graded

Replication, robustness and convergence all matter because each reveals something different about the evidential position of a claim.

But none provides a universal switch from uncertain to established.

New evidence should revise confidence in proportion to what it establishes.

A strong replication can increase confidence.

A precise failure can weaken it.

A robustness analysis can expose fragility.

A different measurement or design can reveal that a result generalises—or that it does not.

Apparently conflicting findings can reveal previously hidden differences between populations, contexts or conditions.

And new evidence can sometimes force a claim to be narrowed or abandoned altogether.

This means scientific confidence is not merely about accumulating studies.

It is responsive to what accumulating evidence actually reveals.

Relevant questions include: Was the phenomenon measured appropriately? Does the design support the inference being claimed? Does the finding recur? What is the estimated magnitude, and how uncertain is it? Does the conclusion survive reasonable alternatives? Do meaningfully distinct approaches support compatible conclusions? What contradictory evidence exists? Under which populations and conditions does the claim hold?

There is no universal number of replications that turns a claim into scientific truth.

There is no significance threshold that settles every scientific question.

And there is no formula in which more papers automatically produce stronger knowledge.

Scientific confidence is graded, evidence-responsive and claim-specific.

That does not make strong scientific conclusions impossible.

Some bodies of evidence warrant extremely high confidence.

The principle is simply that the strength and scope of the conclusion should remain proportional to what the evidence actually supports.

Four questions for evaluating a scientific claim

The logic can be compressed into four questions.

Does the finding recur? Replication.

Does the conclusion survive reasonable alternatives? Robustness.

Do different evidential routes support compatible conclusions? Convergence, or triangulation.

What confidence does the total evidence warrant? Scientific confidence.

These questions are complementary. None can be interpreted independently of measurement quality or the kind of claim being made.

A replicated association remains an association unless additional evidence supports causation.

A robust result can still rest on a poor representation of the underlying construct.

Converging approaches can still share an important weakness.

And one failed study does not automatically erase an otherwise strong evidence base.

The value lies in what these tests reveal together.

From findings to scientific confidence

One study can matter enormously.

But scientific claims become more trustworthy when they remain open to further measurement, independent testing, alternative analyses and evidential challenge.

Sometimes those tests strengthen the original conclusion.

Sometimes they weaken it.

Sometimes they reveal that an effect is smaller than first believed, applies only under certain conditions or requires a different explanation.

And sometimes they show that the original claim should no longer be retained.

That is not a failure of scientific confidence. It is part of how confidence becomes better calibrated.

Strong scientific claims are not strong because one study looked convincing or because enough papers accumulated around them. They are strong when the relevant body of evidence continues to warrant strong confidence after serious opportunities to test, challenge, qualify and revise the claim.

Scientific confidence should usually rest on the relevant body of evidence rather than an isolated finding.

Sources and research record5 sources, with findings, strengths and limitations as entered

References

5 sources this piece rests on, as entered in the Library.

  1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

    Consensus report · The National Academies Press, Washington, DC

    The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

    doi:10.17226/25303

  2. Open Science Collaboration (2015) Estimating the reproducibility of psychological science

    Metascientific study · Science, 349(6251) · aac4716

    The large replication project that raised sustained attention to replication, transparency and cumulative evidence. Cited here as historical and metascientific context, not as a universal replication rate.

    doi:10.1126/science.aac4716

  3. Simonsohn, U., Simmons, J. P., Nelson, L. D. (2020) Specification curve analysis

    Methodological article · Nature Human Behaviour, 4(11) · 1208–1214

    Formalises the problem of analytical flexibility: a conclusion can depend on one specification when many other defensible specifications address the same question, and robustness can be examined across that set.

    doi:10.1038/s41562-020-0912-z

  4. Lawlor, D. A., Tilling, K., Davey Smith, G. (2016) Triangulation in aetiological epidemiology

    Methodological article · International Journal of Epidemiology, 45(6) · 1866–1886

    Sets out why convergence is informative when approaches carry different key sources of bias, and why agreement between methods sharing a vulnerability can reproduce the weakness along with the result.

    doi:10.1093/ije/dyw314

  5. Wasserstein, R. L., Lazar, N. A. (2016) The ASA statement on p-values: Context, process, and purpose

    Professional statement · The American Statistician, 70(2) · 129–133

    The statistical profession's statement that a p-value does not measure the importance of a result or the probability that a claim is true, and that scientific conclusions should not rest on whether a threshold was crossed.

    doi:10.1080/00031305.2016.1154108

Further reading

Behind this page

The claims this research note makes, the evidence behind them, and the limits it accepts.

Evidence status

High confidence

Strongly supported, though resting on synthesis or principle rather than a single decisive body of evidence.

Claims

  1. A single study can provide valuable evidence without establishing a scientific claim by itself

    Established

    What this does not assert: This concerns the evidential status of one result, not whether the study was well conducted.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

  2. Convergence is most informative when supporting methods do not share the same major vulnerabilities

    Established

    What this does not assert: Agreement between approaches sharing a bias can reproduce the weakness along with the result.

    1. Lawlor, D. A., Tilling, K., Davey Smith, G. (2016) Triangulation in aetiological epidemiology

      Methodological article · International Journal of Epidemiology, 45(6) · 1866–1886

      Sets out why convergence is informative when approaches carry different key sources of bias, and why agreement between methods sharing a vulnerability can reproduce the weakness along with the result.

      doi:10.1093/ije/dyw314

  3. More papers do not automatically mean more meaningfully independent evidence

    High confidence

    What this does not assert: Studies can share datasets, instruments, populations, assumptions or research programmes.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

    2. Lawlor, D. A., Tilling, K., Davey Smith, G. (2016) Triangulation in aetiological epidemiology

      Methodological article · International Journal of Epidemiology, 45(6) · 1866–1886

      Sets out why convergence is informative when approaches carry different key sources of bias, and why agreement between methods sharing a vulnerability can reproduce the weakness along with the result.

      doi:10.1093/ije/dyw314

  4. Repeating the same measurement cannot by itself repair a validity problem shared across the replications

    Canonical inference

    What this does not assert: Recurrence of a measured pattern is not evidence that the measure represents the intended construct.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

  5. Repeated evidence of an association does not by itself establish causation

    Established

    What this does not assert: Other designs or identification strategies are required for that transition.

    1. Lawlor, D. A., Tilling, K., Davey Smith, G. (2016) Triangulation in aetiological epidemiology

      Methodological article · International Journal of Epidemiology, 45(6) · 1866–1886

      Sets out why convergence is informative when approaches carry different key sources of bias, and why agreement between methods sharing a vulnerability can reproduce the weakness along with the result.

      doi:10.1093/ije/dyw314

  6. Replicated causal effects do not automatically establish mechanism

    Canonical inference

    What this does not assert: Evidence that an effect occurs can leave the process producing it unresolved.

  7. Statistical significance is not equivalent to scientific confidence

    Established

    What this does not assert: Interpretation also depends on measurement, estimated magnitude, uncertainty, design and the broader evidence base.

    1. Wasserstein, R. L., Lazar, N. A. (2016) The ASA statement on p-values: Context, process, and purpose

      Professional statement · The American Statistician, 70(2) · 129–133

      The statistical profession's statement that a p-value does not measure the importance of a result or the probability that a claim is true, and that scientific conclusions should not rest on whether a threshold was crossed.

      doi:10.1080/00031305.2016.1154108

  8. Scientific confidence is graded, evidence-responsive and claim-specific

    Canonical synthesis

    What this does not assert: The Library's working position: no replication count, significance threshold or citation total settles a scientific question.

  9. Confidence must be attached to a specified claim rather than globally to a paper

    Canonical inference

    What this does not assert: The same study can support an association strongly and a mechanism weakly.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

  10. Replication tests whether compatible evidence recurs using new observations

    Established

    What this does not assert: Replications vary in how closely they reproduce the original procedure.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

  11. Reproducibility and replication answer different questions

    Established

    What this does not assert: Disciplines do not use the terms uniformly; the distinction is a convention adopted for clarity.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

  12. Successful and failed replications both update confidence rather than issuing binary verdicts

    High confidence

    What this does not assert: How much either updates confidence depends on the quality, precision and relevance of the replication.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

    2. Open Science Collaboration (2015) Estimating the reproducibility of psychological science

      Metascientific study · Science, 349(6251) · aac4716

      The large replication project that raised sustained attention to replication, transparency and cumulative evidence. Cited here as historical and metascientific context, not as a universal replication rate.

      doi:10.1126/science.aac4716

  13. Replication can change confidence in an effect's magnitude, not merely whether it exists

    Established

    What this does not assert: An effect can recur while proving substantially smaller than its original estimate.

    1. Open Science Collaboration (2015) Estimating the reproducibility of psychological science

      Metascientific study · Science, 349(6251) · aac4716

      The large replication project that raised sustained attention to replication, transparency and cumulative evidence. Cited here as historical and metascientific context, not as a universal replication rate.

      doi:10.1126/science.aac4716

  14. Robustness tests whether a conclusion survives reasonable alternative choices

    Established

    What this does not assert: Which alternatives count as reasonable must be justified by the question, theory, measurement and statistical assumptions.

    1. Simonsohn, U., Simmons, J. P., Nelson, L. D. (2020) Specification curve analysis

      Methodological article · Nature Human Behaviour, 4(11) · 1208–1214

      Formalises the problem of analytical flexibility: a conclusion can depend on one specification when many other defensible specifications address the same question, and robustness can be examined across that set.

      doi:10.1038/s41562-020-0912-z

  15. Robustness does not guarantee that every important source of error has been eliminated

    Canonical inference

    What this does not assert: A conclusion can survive all alternatives tested and remain vulnerable to an assumption none of them addressed.

    1. Simonsohn, U., Simmons, J. P., Nelson, L. D. (2020) Specification curve analysis

      Methodological article · Nature Human Behaviour, 4(11) · 1208–1214

      Formalises the problem of analytical flexibility: a conclusion can depend on one specification when many other defensible specifications address the same question, and robustness can be examined across that set.

      doi:10.1038/s41562-020-0912-z

  16. Replication and robustness test different vulnerabilities

    High confidence

    What this does not assert: A result can recur in new data while depending on a fragile analytical decision, and the reverse.

    1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

      Consensus report · The National Academies Press, Washington, DC

      The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

      doi:10.17226/25303

    2. Simonsohn, U., Simmons, J. P., Nelson, L. D. (2020) Specification curve analysis

      Methodological article · Nature Human Behaviour, 4(11) · 1208–1214

      Formalises the problem of analytical flexibility: a conclusion can depend on one specification when many other defensible specifications address the same question, and robustness can be examined across that set.

      doi:10.1038/s41562-020-0912-z

Sources

  1. National Academies of Sciences, Engineering, and Medicine (2019) Reproducibility and Replicability in Science

    Consensus report · The National Academies Press, Washington, DC

    The consensus statement separating reproducibility from replication, documenting disciplinary variation in terminology, and setting out why replication outcomes are evidence to be interpreted rather than binary verdicts on truth.

    doi:10.17226/25303

  2. Open Science Collaboration (2015) Estimating the reproducibility of psychological science

    Metascientific study · Science, 349(6251) · aac4716

    The large replication project that raised sustained attention to replication, transparency and cumulative evidence. Cited here as historical and metascientific context, not as a universal replication rate.

    doi:10.1126/science.aac4716

  3. Simonsohn, U., Simmons, J. P., Nelson, L. D. (2020) Specification curve analysis

    Methodological article · Nature Human Behaviour, 4(11) · 1208–1214

    Formalises the problem of analytical flexibility: a conclusion can depend on one specification when many other defensible specifications address the same question, and robustness can be examined across that set.

    doi:10.1038/s41562-020-0912-z

  4. Lawlor, D. A., Tilling, K., Davey Smith, G. (2016) Triangulation in aetiological epidemiology

    Methodological article · International Journal of Epidemiology, 45(6) · 1866–1886

    Sets out why convergence is informative when approaches carry different key sources of bias, and why agreement between methods sharing a vulnerability can reproduce the weakness along with the result.

    doi:10.1093/ije/dyw314

  5. Wasserstein, R. L., Lazar, N. A. (2016) The ASA statement on p-values: Context, process, and purpose

    Professional statement · The American Statistician, 70(2) · 129–133

    The statistical profession's statement that a p-value does not measure the importance of a result or the probability that a claim is true, and that scientific conclusions should not rest on whether a threshold was crossed.

    doi:10.1080/00031305.2016.1154108

Read this first

This piece assumes them.

What this opens up

What becomes readable once you have this.

Where to go from here

Next published piece

Mind, Brain and Body Form One System

Mind, brain and body can be distinguished scientifically without being treated as separate systems. Human behaviour emerges through interacting psychological, neural, physiological and behavioural processes within an embodied organism.

Continue through the Library →See where this sits in the graph →

Back to the Library →