Skip to content
‹ Papers
PLoS Medicine

Why Most Published Research Findings Are False

John P. A. Ioannidis· 30 Aug 2005Peer-reviewed

A modelling essay arguing that the probability a published research finding is true depends on prior odds, statistical power, and bias, and that under common conditions most published findings are more likely false than true. Structured here as propositions with a link-out to the open-access original (CC BY).

When John Ioannidis published this essay in 2005, he was not reporting an experiment. He was doing arithmetic about how science itself behaves. His question was simple: when a study announces a positive finding, what is the chance it is actually true?

The answer, he showed, is not fixed. It depends on three things: how likely the hypothesis was before the study ran, how much statistical power the study had to detect a real effect, and how much bias crept in through flexible design and analysis. Feed realistic values into that relationship and, for many common kinds of research, the probability that a positive result is true falls below one half. From the same logic follow his well-known corollaries: smaller studies are less trustworthy, more analytical flexibility makes findings less trustworthy, and even a crowded, competitive field lowers the odds that any single positive result holds up.

What the paper does not do is measure how many published findings are in fact false. It is a model, not a survey. Its power is in reframing reliability as a predictable consequence of study design rather than a matter of individual error, and that reframing reshaped how a generation thinks about reproducibility. Read it as a way of reasoning, not as a headcount.

A landmark meta-argument whose logic is sound and trivially reproducible, but whose famous headline is an illustrative model output under chosen assumptions, not an empirical measurement. Read it as a framework for reasoning about reliability, not as a measured head-count of false findings.

Evidence tierweak
theoretical model / simulation

As evidence about the real world this is the lowest empirical tier: no data, no held-out sample. Every conclusion is an analytic consequence of assumed pre-study odds, power and bias, illustrated with plausible values.

Effect vs significancen/a

There is no measured effect or significance test. The 'PPV < 0.5' figure is the output of a deterministic equation under selected priors, not an estimate with a confidence interval.

Reproducibilitystrong

Fully reproducible: the model is a handful of closed-form equations anyone can re-derive; there is no data or code dependency. Peer-reviewed and open access (CC BY).

Generalizabilitymoderate

The framework is deliberately field-agnostic, but the specific numbers depend on field-specific priors, power and bias that are asserted for illustration rather than measured.

Conflictsstrong

No funding and no competing interests declared; a methodological argument with no obvious beneficiary.

Noveltystrong

Reframed research reliability as a positive-predictive-value problem and became one of the most-cited papers in science — a genuinely new synthesis, not a delta on prior art.

Claim 1

For most study designs and settings, it is more likely for a research claim to be false than true.

MethodAnalytic model of the positive predictive value (PPV) of a research finding as a function of pre-study odds R, statistical power (1-beta), and Type I error alpha.Samplen/a (theoretical model, not an empirical sample)EffectPPV < 0.5 across the plausible parameter ranges tabulated.UncertaintyNo confidence interval — a deterministic model; sensitivity shown across parameter values, not sampled.
Modeling the Framework · Table 4
Claim 2

The smaller the studies conducted in a scientific field, the less likely the research findings are to be true.

MethodCorollary of the PPV model: small n lowers power, which lowers PPV for a fixed pre-study odds.Samplen/a (corollary)EffectDirectional: PPV decreases as power decreases.UncertaintyModel-derived, not empirically estimated.
Corollaries · Corollary 1
Claim 3

The greater the flexibility in designs, definitions, outcomes, and analytical modes, the less likely the findings are to be true.

MethodCorollary introducing a bias term u into the PPV model; analytic flexibility increases u.Samplen/a (corollary)EffectDirectional: PPV decreases as bias u increases.UncertaintyModel-derived.
Corollaries · Corollary 4
Claim 4

The hotter a scientific field (more teams involved), the less likely the research findings are to be true.

MethodCorollary: with many independent teams chasing significance, the probability at least one reports a false positive rises.Samplen/a (corollary)EffectDirectional: PPV decreases as the number of competing teams rises.UncertaintyModel-derived.
Corollaries · Corollary 6
confidence capped at low

This is a methodology paper, not a market catalyst. The appraisal's weak evidence tier caps every chain below at low confidence — read these as diffuse, long-horizon second-order effects, not trade ideas.

early-stage drug licensing / translational biotechlonger diligence cycles; a premium on preregistered, replicated assets over single-lab results

As reproducibility pressure mounts, pharma and biotech raise the internal evidentiary bar before licensing external preclinical research.

low confidence (capped from moderate)5-10 years, diffuse
research-tooling / reproducibility-infrastructure vendorsstructural tailwind for preregistration, data-availability and replication tooling

Crowded, hype-driven research areas carry a higher base rate of non-replication, which slowly re-rates the credibility premium of incumbents with reproducible track records.

low confidencelong, uncertain
Read the open-access original (doi:10.1371/journal.pmed.0020124)