Skip to content

July 2026: Government declines to widen the Thirlwall terms of reference (16 July) · new 100-page insulin report to the CCRC challenging the trial evidence (9 July) · Thirlwall report still expected no earlier than September · inquests relisted to 2027 · Shoo Lee Panel: no medical evidence of deliberate harm.

Lucy Letby Facts

Statistical explainer

Multiple comparisons and false positives

When many comparisons are run, the probability that at least one comparison appears significant by chance grows quickly. Hospital cluster investigations that search for any pattern linking a suspect to events are especially exposed.

Statistical analysis2Expert/professional
Last updated

Why it matters in the Letby case

Statistical critique of the Letby investigation argues multiple-comparisons effects were not adequately controlled in the public presentation of patterns.

Prosecution position

The investigation and its public presentation drew on a range of patterns connecting Letby to the events, of which the shift chart was the most prominent.

Expert challenge / post-conviction reading

Statistical critique argues multiple-comparisons effects were not adequately controlled. When many possible patterns are searched, the chance that at least one looks striking by coincidence rises quickly, and a hospital investigation searching for any link between a suspect and a set of events is particularly exposed.

What remains uncertain

How many comparisons were actually run is not on the public record, so the size of the correction that should apply cannot be calculated.

Common questions

What is the multiple-comparisons problem?

If you test enough hypotheses, some will look significant purely by chance. The more you look, the more careful you have to be about what you find.

Why are cluster investigations exposed to it?

Because they search broadly for any factor linking a suspect to a set of events, which is exactly the situation in which coincidental patterns emerge.

Does this mean the pattern was coincidence?

No. It means the strength of a pattern has to be discounted by how much searching produced it, and that discount was not applied.

Can the correction be worked out now?

Not reliably — how many comparisons were run is not on the public record.

Found an error or missing source? Send a correction with the page URL and supporting source.