Method notes · 04
Someone has to get lucky
An unlikely streak for one named person can be an ordinary event among thousands of people, cards, signs, dates, and outcomes.
In brief
- Ask how many opportunities existed for an impressive result, not only how rare the winner looks.
- Choosing the most unusual group after seeing the data changes the probability question.
- Multiplicity corrections protect a family of claims; they do not erase effect sizes or uncertainty.
The named person and the crowd
Ten correct Zener guesses in a row has probability one in 5 to the tenth power—1 in 9,765,625—for a particular person beginning at a particular trial. That sounds astonishing. Give millions of people repeated chances and the probability that somebody produces such a run can become substantial.
The same change of denominator appears when twelve signs, twenty-two tarot cards, several outcomes, and many dates are inspected. A result selected because it is extreme cannot be judged as though it were the only result anyone intended to examine.
A family of chances
At a five-percent threshold, one valid null test has a five-percent chance of crossing the line. Twenty independent null tests have about a 64-percent chance that at least one crosses it. The tests in real data are often correlated, so that number is not universal, but the direction is: more opportunities create more accidental highlights.
The family must be defined by the question. If the article claims that any sign behaves differently, all twelve sign comparisons belong together. If four tarot contrasts were singled out before collection, those four form the confirmatory family; the other cards can still be described as exploration.
What Holm correction does
Holm’s procedure orders the p-values from smallest to largest and asks the smallest to clear the strictest threshold. If it succeeds, the next receives a slightly easier threshold, and so on. It controls the chance of one or more false rejections across the family without assuming that every test is independent.
A correction is not a punishment for looking. It is an accounting rule for the number of claims the same evidence was allowed to win. We still publish raw estimates and intervals because a corrected decision does not tell a reader whether an effect is large, precise, or useful.
Sources
- A simple sequentially rejective multiple test procedure
Holm, Scandinavian Journal of Statistics 6(2), 1979 · Peer reviewed · SRC-1F3E358F - The ASA statement on p-values
Wasserstein & Lazar, The American Statistician 70(2), 2016 · Peer reviewed · DOI 10.1080/00031305.2016.1154108 · SRC-A95BDF57