Skip to content

Explainer · Frontier

What Does 5 Sigma Actually Mean?

Five sigma is the rule that turns a measurement into a discovery, and it is quoted far more often than it is understood. It does not mean a result is 99.9999 percent certain to be real, which is the way it is almost always reported. This piece explains what a p-value actually is, why particle physics sets the bar at five sigma rather than three, and why a frequentist significance and a Bayesian odds can look at the same data and reach opposite-sounding verdicts. It uses four results the site already covers, from the muon to dark energy, and one famous case where six sigma was undone by a loose cable.

Abstract technical illustration on a dark slate background showing acoustic soundwave interference patterns on the left connected to glowing electromagnetic field lines and particle physics orbit trajectories on the right.

On 4 July 2012, CERN announced the Higgs boson at five sigma, and a room full of physicists applauded a number most people in the wider audience could not define. Five sigma is the rule that decides when a measurement becomes a discovery. It is also badly misunderstood, routinely reported as “99.9999 percent certain the particle is real,” which is not what it says. This is what a sigma actually measures, why the bar sits at five, and why two honest statisticians can look at the same data and disagree about whether anything has been found.

What a p-value is, and the thing it is not

Start with the quantity underneath the sigma, which is the p-value. Suppose there is no new particle, no signal, only the familiar background. The p-value is the probability that random fluctuations in that background would, by chance alone, produce a bump at least as large as the one you saw.

Small p-value, rare fluke, so the background-only story looks strained. That is the whole logic. A sigma count is just this probability rewritten as a distance, in standard deviations, out along the tail of a bell curve. One sigma is a p-value around 0.16, hardly rare. Three sigma is about one in 740. Five sigma is roughly one in 3.5 million.

Now the part almost every headline gets wrong. The p-value assumes the null hypothesis is true and asks how surprising the data would be. It does not tell you the probability that the null hypothesis is true. Five sigma is not a 99.9999 percent chance the Higgs is real. It is a statement that if the Higgs were not there, data this striking would appear about once in 3.5 million tries. The direction of the conditional matters, and reversing it is the single commonest error in science reporting.

Significance is not size

A high sigma count says an effect is probably there. It says nothing about how big the effect is. A tiny difference measured precisely enough can reach five sigma, and a large one measured sloppily may not clear three. Sigma answers “is something there?”, never “how much?”. Keeping those two questions apart is the beginning of reading these claims well.

Why the bar is five, and not three

Five sigma is a convention, not a law of nature. Particle physics keeps two bars: three sigma to call something “evidence,” five to call it a “discovery.” Most journals will not let the word discovery near a title below five. The statistician Louis Lyons, who has written the standard defences of the threshold, gives three reasons it sits so high.

History. Too many three and four sigma signals have simply evaporated once more data arrived. The graveyard is large enough that the field learned to distrust anything short of five.

The look-elsewhere effect. Scan a wide enough range and a fluctuation somewhere becomes almost guaranteed, even when nothing real is present. Lyons’s colleague Glen Cowan makes the point with a plot of pure random noise in a hundred bins: the eye picks out several convincing “peaks,” and every one is empty. The p-value for one specific spot overstates things. The honest number corrects for all the places you could have looked, and that correction costs you significance.

Systematics. A sigma is the signal divided by the uncertainty in the background. If you have underestimated that uncertainty, because some instrumental effect is mis-modelled, a three sigma bump can be pure artefact. Demanding five buys a margin against the systematic errors you have not thought of.

Lyons himself is not dogmatic about it. He has argued that a fixed bar is crude. An extraordinary claim, such as neutrinos outpacing light, should demand far more than five, while a well-predicted one like the Higgs might reasonably need less. The rule endures anyway, partly for its simplicity and partly for fairness. A shared, stated-in-advance threshold stops credit going to whichever team is willing to publish on the flimsiest evidence.

The sigma scale at a glance

Here is the whole scale in one picture, with four results from this article marked on it. Watch the dot furthest to the right. OPERA cleared the discovery bar comfortably, and it was still wrong.

The sigma scale, from pure chance to discovery

The bell curve and the significance scale, with famous results The bell curve of pure chance 0 1σ 2σ 3σ 4σ 5σ 3σ: evidence 1 in 740 5σ: discovery 1 in 3.5 million too thin to see Chance of a fluke this big or bigger 1 in 100 1 in 10,000 1 in a million 1 in 100 million 1 in 10 billion 3σ evidence 5σ discovery 1σ 2σ 3σ 4σ 5σ 6σ 7σ Dark energy 3.2σ Muon g-2 4.2σ Higgs 5σ OPERA 6.2σ The 2023 Bell test lies far below this chart.
Top: how background fluctuations are spread when nothing real is there. The tail beyond 3σ is a sliver; beyond 5σ it is too thin to draw. Bottom: the same tail on a logarithmic scale, each gridline a hundred times rarer than the last, with four results from this article placed by their significance. Odds are one-sided, the particle-physics convention.

Six sigma, and wrong: the OPERA cable

In September 2011 the OPERA experiment reported neutrinos arriving from CERN at Gran Sasso about 60.7 nanoseconds early, which is to say faster than light. After a repeat run the significance reached 6.2 sigma, past the discovery bar, a result that would have overturned special relativity.

It was wrong. A fibre-optic connector carrying a GPS timing signal to the master clock was not fully screwed in, adding a delay of around 73 nanoseconds. A clock oscillator running fast pulled about 15 nanoseconds the other way. The two corrections nearly cancelled and left roughly 58 nanoseconds, almost exactly the anomaly. The neutrinos had obeyed the speed limit all along.

The lesson is the one the five sigma rule cannot fully enforce. OPERA’s statistics were immaculate. The 6.2 sigma described the precision of the measurement, not its accuracy, and a loose cable lives in the gap between the two. No p-value protects against an apparatus that is quietly lying to you.

The anomaly that vanished without new data

For twenty years the muon’s magnetism was the most watched crack in the Standard Model. Measured, it came out slightly larger than theory predicted, and the gap refused to close. Against the 2020 theory prediction, Fermilab’s 2021 measurement put the discrepancy at 4.2 sigma, tantalisingly near discovery.

Then the gap dissolved, and almost none of the movement came from the experiment. The hard part of the prediction is the contribution of the strong force. It can be computed two ways: from older collider data, or from lattice simulations of the theory on a grid. The two disagreed. When the 2025 theory update adopted the lattice value, the predicted number rose to meet the measurement, and the discrepancy fell to statistical noise. Fermilab’s own final result, more precise than ever, agrees with the revised theory.

What the muon teaches

A sigma measures the gap between two numbers, and both numbers carry uncertainty. Here the measurement barely moved while the theory did, and a near-five-sigma anomaly evaporated. Significance was never a verdict on reality. It was a running score on the distance between an experiment and a calculation, and the calculation caught up. The fuller account sits in our piece on the muon g-2 anomaly.

The other question: how much do you now believe?

A p-value is deliberately one-sided. It scores how badly the no-signal story fits, and stays silent on how much you should believe the alternative. Many physicists, especially in cosmology, want that second question answered directly, and for it they turn to Bayesian model comparison.

The tool is the Bayes factor: given the data, how much have the odds between two models shifted? Its signature feature is a built-in penalty for complexity. A model with a spare free parameter can accommodate many possible datasets, and that flexibility counts against it. To win, the richer model must fit not just better but enough better to justify the freedom it claimed. Add a parameter that the data barely constrain and the Bayes factor drifts back toward even odds, its verdict being that the data did not really move you.

Cosmologists read the result on the Jeffreys scale, stated by Roberto Trotta. Odds around 3 to 1 count as weak, 12 to 1 moderate, 150 to 1 strong. Because the scale runs on the logarithm of the odds, climbing one rung takes roughly ten times more supporting data. Evidence, on this view, accrues slowly and grudgingly.

When the two frameworks disagree, on purpose

Because the two approaches ask different questions, they can return answers that look contradictory while both are right. The clearest live example is dark energy.

After a 2026 recalibration, the case for dark energy evolving over cosmic time stood at 3.2 sigma by the frequentist count. The same dataset, weighed as a Bayes factor, gave odds of only about 5 to 1, which the Jeffreys scale files under weak. One number sounds like a near-discovery; the other sounds like a shrug. Neither is wrong. The frequentist figure asks how rare the data would be with no evolution, and the answer is fairly rare. The Bayesian figure asks whether adding evolution, with its extra parameters, was worth it, and the answer is not especially. We walk through the physics in our piece on what dark energy is.

Each framework has a matching weakness, and honesty means owning both. The frequentist p-value is hostage to the look-elsewhere effect and to how well you modelled your systematics. The Bayesian evidence depends on the prior you place on the extra parameter, and a defensible change of prior can shift the verdict by a full rung on the scale. There is no framework that removes judgement. There are only two ways of making the judgement explicit.

A p-value done right

None of this makes the machinery worthless. It makes it a tool with a domain, and there are places it works almost perfectly.

A Bell test is the cleanest case. The null hypothesis is sharp: the world runs on local hidden variables. There is one number to compute, no free parameters to tune, no wide spectrum to scan for a bump. When a 2023 superconducting experiment violated the Bell bound, it reported a p-value below ten to the power of minus 108. That is a frequentist statement with no look-elsewhere ambiguity and no systematics quietly inflating it, which is exactly why it carries such weight. The framework is not the problem. Using it outside its domain, and misreading what it says, is.

How to read a significance claim

A few habits separate a durable claim from a fragile one.

Ask what the null was. A p-value is only as meaningful as the hypothesis it tries to reject, and a sharp null beats a vague one. Ask whether the number is local or global. A raw local p-value that ignores the look-elsewhere effect is optimistic by design. Ask where the uncertainty is. If a significance rests on a theory prediction as much as a measurement, the prediction can move, as the muon showed. Ask which framework is talking. A frequentist sigma and a Bayesian odds are answering different questions, and quoting one while imagining the other is how a weak result gets dressed as a strong one.

Established, contested, unproven

Established. The definitions. A p-value is a tail probability under the null, sigma is that probability as a distance, and five sigma is about one in 3.5 million. The two-tier convention, evidence at three and discovery at five, is standard across particle physics. The look-elsewhere effect and systematic uncertainty are real and quantifiable.

Contested. Whether five sigma is the right bar. Lyons and others argue a claim-dependent threshold would serve better, stricter for the implausible and looser for the well-motivated. More broadly, whether frequentist or Bayesian methods are the sounder default is a decades-old argument statisticians have not settled. To a real extent it is a question about what you want a number to mean.

Unproven, by construction. No amount of sigma proves a discovery is real. The figure is always conditional on a model of the background and the apparatus, and OPERA is the standing reminder that a flawless statistical significance can sit on top of a loose cable. The number is a guide to how seriously to take a result. It is never the last word.

Note on sourcing

The threshold conventions and the three justifications for five sigma come from Louis Lyons’s papers on discovery significance and from Glen Cowan’s lectures on statistics for physics searches, all peer-reviewed or standard references. The Bayesian scale is Roberto Trotta’s, from his cosmology review. The muon g-2 figures trace to the Fermilab final report and the 2025 Theory Initiative white paper; the dark energy comparison rests on a 2026 preprint, flagged as such in the linked article. The OPERA account and its resolution are drawn from the collaboration’s own revised measurement and the contemporary reporting of the cable fault.

References

  1. L. Lyons, Discovering the Significance of 5 sigma, preprint (2013)
  2. L. Lyons, Statistical Issues in Searches for New Physics, preprint (2014)
  3. G. Cowan, K. Cranmer, E. Gross and O. Vitells, Asymptotic formulae for likelihood-based tests of new physics, European Physical Journal C 71, 1554 (2011) doi:10.1140/epjc/s10052-011-1554-0
  4. G. Cowan, Statistics for Searches at the LHC, preprint (2013)
  5. G. Cowan, Practical Statistics for Particle Physics, CERN Yellow Reports (2019)
  6. R. Trotta, Bayes in the sky: Bayesian inference and model selection in cosmology, Contemporary Physics 49, 71 (2008) doi:10.1080/00107510802066753
  7. R. Trotta, Applications of Bayesian model selection to cosmological parameters, Monthly Notices of the Royal Astronomical Society 378, 72 (2007) doi:10.1111/j.1365-2966.2007.11738.x
  8. E. Linder and R. Miquel, Cosmological Model Selection: Statistics and Physics, International Journal of Modern Physics D 17, 2315 (2008)
  9. Muon g-2 Collaboration, Final Report on the Measurement of the Positive Muon Anomalous Magnetic Moment at Fermilab to 127 ppb, preprint (2026)
  10. OPERA Collaboration, Measurement of the neutrino velocity with the OPERA detector in the CNGS beam, preprint, revised version (2011)

What does five sigma mean?

It means that if there were no real signal, random background fluctuations would produce an effect at least as large as the one observed only about once in 3.5 million times. It is the conventional threshold for claiming a discovery in particle physics.

Does five sigma mean 99.9999 percent certain?

No. That is the most common misreading. Five sigma is a statement about how unlikely the data would be if there were no signal. It is not the probability that the discovery itself is real, which is a different quantity the p-value does not give you.

What is a p-value?

The probability of observing data at least as extreme as what was seen, assuming the null hypothesis, usually "background only, no new effect," is true. A small p-value means the background-only explanation fits the data poorly.

What is the difference between three sigma and five sigma?

Three sigma, a roughly one in 740 chance, is the conventional threshold for calling a result "evidence." Five sigma, about one in 3.5 million, is the threshold for calling it a "discovery." The gap reflects how often smaller signals have turned out to be flukes.

Why is the discovery threshold so high?

Three reasons. Many three and four sigma signals have vanished with more data. The look-elsewhere effect makes a fluctuation somewhere almost certain in a wide search. And underestimated systematic errors can fake a moderate signal. Five sigma buys a margin against all three.

What is the look-elsewhere effect?

The fact that if you search a wide range of possibilities, finding an unlikely-looking fluctuation somewhere becomes probable even when nothing real is present. A p-value for one specific location overstates the significance unless it is corrected for all the places you could have looked.

Does a higher sigma mean a bigger effect?

No. Sigma measures statistical significance, meaning how confident you are that an effect is present, not its size. A very small effect measured precisely enough can reach high significance, while a large effect measured poorly may not.

What happened with the OPERA neutrinos?

In 2011 the OPERA experiment measured neutrinos apparently travelling faster than light, at 6.2 sigma. The result was later traced to a fibre-optic cable that was not fully connected, plus a fast clock oscillator. Together they accounted for the anomaly, and the neutrinos were within the speed limit.

How can an anomaly disappear without new data?

The muon g-2 case shows how. Its significance measured the gap between a measurement and a theoretical prediction. When the prediction was updated using lattice calculations, the predicted value shifted to match the measurement, and the gap closed, even though the measurement barely changed.

What is a Bayes factor?

A measure of how much the odds between two competing models shift in light of the data. Unlike a p-value, it directly compares two hypotheses and includes a penalty for model complexity, so a model with extra parameters must fit substantially better to be favoured.

Why do frequentist and Bayesian analyses sometimes disagree?

Because they answer different questions. A frequentist p-value asks how unlikely the data would be with no signal. A Bayesian analysis asks how much the data shift the odds between two models, charging the more complex one for its extra freedom. The same data can look significant to one and weak to the other.

What is the Jeffreys scale?

A convention for interpreting Bayes factors, set out by Roberto Trotta for cosmology. Odds of about 3 to 1 count as weak evidence, 12 to 1 as moderate, and 150 to 1 as strong. Because it runs on the logarithm of the odds, moving up one level takes roughly ten times more data.

Which framework is correct, frequentist or Bayesian?

Neither is simply correct; they answer different questions and have different weaknesses. Frequentist significance is vulnerable to the look-elsewhere effect and to mismodelled systematics. Bayesian evidence depends on the prior chosen for the extra parameters. The honest approach states which is being used.

Can a five sigma result still be wrong?

Yes. Significance is always conditional on a model of the background and the apparatus. If that model is flawed, as with a mismodelled systematic or a faulty instrument, a very high significance can still be wrong. OPERA at 6.2 sigma is the standard example.

How should I read a significance claim in the news?

Ask what the null hypothesis was, whether the p-value is corrected for the look-elsewhere effect, whether the significance depends on a theory prediction that could move, and whether the number is frequentist or Bayesian. Those four questions separate a robust claim from a fragile one.

Leave a response

Your email address is not published.

Newsletter

Get new Quantum Nature articles by email.

New research write-ups and explainers, sent only when there is something worth reading — no fixed schedule, no filler.

Browse recent articles

Free. Unsubscribe in one click.