Occult Research Observatory
CONTINUOUS OBSERVATION · EST. MMXXVI · ALL DATA PUBLIC
Language: EnglishPolski

Experiment I · Open

The Blind Horoscope Test

Twelve horoscopes rise each morning with their signs effaced. If the stars write anything personal into them, their readers should find their own at better than one part in twelve.

EXP-001 · STATUS: Collecting · v1.3

Opening the register…

Collecting — no figures published yet

The register is open and entries are arriving. Nothing appears here until a day closes: every entry stays sealed until 00:00 UTC, so the first figures publish after the first full day. That delay is what stops anyone, ourselves included, from watching a result take shape and choosing when to stop.

The wager

Readers shown all twelve of a day's horoscopes — unlabelled and shuffled — will select the one written for their own sun sign more often than chance. Chance is one in twelve: 8.33%. The astrological claim predicts more; the null predicts exactly that, forever.

The analysis is fixed in advance: a two-sided exact binomial test against p₀ = 1/12 at α = .05, with conclusions drawn only at sample sizes of 1,000, 5,000, 10,000 and 25,000. Because this is filed before any data exists, there is no room to move the goalposts once results arrive.

The filed protocol

The full preregistration is a document in the site's public repository. Its SHA-256 is printed below and computed from the file itself, so you can download it, hash it, and confirm that the plan described on this page is the plan that was filed. Every revision is a commit you can read.

Version v1.3EXP-001

SHA-256 84d4a7053997a1994154ce1f6280bfd3021c09b832bcebda3f3329217e6c1b87

Read the protocol · Revision history

Fig. I — What chance looks likep₀ = 1/12 · 95% · POWER 80%
0%5%10%15%20%Hit rate1001,00010,000Sample size (log scale)chance — 8.33%10.8%9.4%9.1%8.8%rates consistent with pure guessingsmallest effect we could resolve
Both curves are computed from the design, not observed. As the sample grows, the band of luck narrows and the smallest resolvable effect falls toward the chance line — which is exactly why the milestone schedule exists. A genuine effect holds its level while the band sinks away beneath it.
Sample size
Chart of the null hypothesis: the 95% band of hit rates consistent with chance narrows as sample size grows, and the smallest detectable true rate falls from about 15% at n=1,000 toward 9% at n=25,000.
Sample sizeBand of chanceSmallest detectable
1,0006.6%10.0%10.8%
5,0007.6%9.1%9.4%
10,0007.8%8.9%9.1%
25,0008.0%8.7%8.8%

The instrument

Source
Twelve readings a day, commissioned from a practising astrologer who is credited by name and paid for the work. The same author throughout a run — switching mid-experiment would quietly change what is being tested.
Disclosure
The author knows the readings are used in a blind identification test. That is stated openly because it favours the astrological hypothesis, not ours: the readings are written by someone with an incentive to make them identifiable, so a null result under those conditions is the stronger conclusion.
Blinding
Signs are stripped server-side and presentation order is shuffled per reader, so position carries no information.
Entry
One observation per participant per day, deduplicated by account and hashed IP.
Sealing
Same-day selections stay hidden until 00:00 UTC. Nobody should be able to see the crowd before choosing.
Registration
Filed before collection opens. Hypothesis, test, α, and milestones are locked; changes are appended and logged, never overwritten.
Open records
Anonymised raw data and the analysis code will be published and refreshed nightly from the day collection opens.

Known ways this could mislead

  1. 01
    Twelve signs means twelve comparisons. At the usual threshold, roughly one sign will look remarkable by luck alone — so per-sign results will carry a multiplicity correction, and the uncorrected figures will be shown beside them rather than quietly dropped.
  2. 02
    A single well-written reading tends to attract picks from everyone regardless of sign. That is the Barnum effect showing up as a crowd favourite, and it is measured separately rather than mistaken for a signal.
  3. 03
    Self-reported sun signs are unverified. Misreporting adds noise, which biases toward finding nothing — worth stating plainly, because it means a null result is the weaker of the two possible conclusions.
  4. 04
    Readers who arrive from astrology communities and readers who arrive from skeptic communities differ in ways we cannot control. Referral source is recorded so the split can be inspected rather than assumed away.

What a sample can and cannot see

Smallest detectable hit rate by sample sizep₀ = 8.33% · 80% POWER
Sample sizeSmallest detectable rateLift over chance
1,00010.8%+2.4%
5,0009.4%+1.1%
10,0009.1%+0.8%
25,0008.8%+0.5%
Sample size sets the smallest real effect an experiment could notice. These are computed for the blind horoscope design against its 8.33% chance baseline, at 80% power — not illustrations. — n ≈ 2,271 for a true rate of 10%.