# EXP-004 — Daily card draws against reported day quality

**Status:** filed, collection not yet open
**Protocol version:** v1.0
**Filed:** 2026-07-24
**Experiment slug:** `tarot-frequency-lab`

This document is the preregistration. Its SHA-256 is displayed on the
experiment page.

---

## 1. Question

Do the cards people draw relate to how their days actually go? Specifically: is
the mean reported day quality the same for every card, or do some cards — The
Tower, Death, The Sun — sit with better or worse days than chance allows?

## 2. The ordering decision, which is the whole experiment

A tarot reading normally runs: draw in the morning, then live the day. Testing
it that way is useless, because a bad card can spoil a day by being drawn. Any
correlation found would be expectancy, not divination, and the two are
indistinguishable in that design.

So this register reverses the order. **The participant rates the day that has
just ended, the rating is recorded, and only then is a card drawn.**

Under that order:

- the card cannot have influenced the rating — it did not exist yet;
- the rating cannot have been adjusted to fit the card — it was already stored.

Which means an association, if one appeared, could not be explained by
expectancy. It would require the draw to carry information about a day already
over. That is a strong claim, and it is the claim the tarot tradition
implicitly makes when it treats a card as meaningful about a period of time.

The cost of this design is that it does not test predictive tarot as practised.
It tests something narrower and cleaner, and says so.

## 3. Materials

The 22 major arcana, in canonical order (0 The Fool … XXI The World). The minor
arcana are excluded: 78 cards would divide the sample into groups too small to
say anything about, and the cards people attach meaning to are almost all major.

Each card is drawn uniformly and independently. A draw is **not** a shuffled
deck without replacement: every entry is an independent trial at p = 1/22.

## 4. Hypotheses

- **H₀ (null):** mean day rating is equal across all 22 cards.
- **H₁:** at least one card differs.

## 5. Primary outcome

Mean reported day quality (1–7) per card.

## 6. Statistical analysis

Fixed before collection, not modifiable while the register is open:

- **Omnibus test:** Kruskal–Wallis across the 22 cards, α = .05. Chosen over
  one-way ANOVA because a 1–7 rating is ordinal and its distribution is likely
  skewed toward the middle.
- **Prespecified contrasts:** four, named now so they cannot be chosen later to
  suit the data — **The Tower**, **Death**, **The Sun**, **The Star**, each
  against the pooled mean of the remaining cards.
- **Correction:** Holm–Bonferroni across the four contrasts, and separately
  across all 22 if the omnibus test is significant.
- **Interval estimates:** 95%, normal approximation for means.
- **Milestone reads:** at 5,000, 20,000 and 50,000 entries.

With 22 cards, roughly one will look remarkable at α = .05 by luck alone. The
correction is what stops that being reported as a finding, and the simulated
null (§7) is what makes it visible rather than merely stated.

## 7. The visible null

A simulation of the same number of entries, drawn from a genuinely random
process with no relationship between card and rating, runs alongside the live
figures and is displayed with them.

Its purpose is to show what an absence of effect *looks like* — a scatter of
card means, some of them apparently striking — so a reader can compare the real
scatter against a scatter known to contain nothing. A table of p-values makes
the same point and is easier to ignore.

The simulation uses the observed overall rating distribution, so it reproduces
the real spread rather than an idealised one.

## 8. Eligibility and deduplication

- One entry per participant per UTC day, enforced by a database constraint.
- Participants are anonymous, identified by a device-scoped random identifier.
- Entries are independent trials; no participant's history affects their draw.

## 9. Prespecified threats to validity

1. **Retrospective rating.** People rate a day after living it, and memory is
   reconstructive. This adds noise rather than bias toward any particular card,
   since the card is unknown at rating time.
2. **Mood-driven participation.** People may be likelier to record an entry on
   unusual days. That skews the overall distribution of ratings but not its
   relationship to a card drawn afterwards.
3. **Multiplicity across 22 cards.** Addressed in §6 and made visible in §7.
4. **Draw uniformity.** The generator is checked: card frequencies are
   published and should sit at 1/22 each. A skewed generator would produce
   apparent card effects, and is a defect rather than a discovery.
5. **Self-selection.** Participants arrive with priors about tarot. Referral
   source is recorded so the split can be inspected.

## 10. Data publication

Anonymised entries — date, card, rating, referral — published and refreshed
from the day collection opens, CC BY 4.0.

## 11. What would change our minds

A card whose mean rating differs from the rest under the preregistered test,
with correction applied and a clean card-frequency distribution, would be a
positive result and published as one. Given the ordering in §2, it would not be
explicable as expectancy — which is precisely why the bar should be, and is,
the preregistered test rather than an eye-catching chart.

## 12. Revision history

| Version | Date | Change |
|---|---|---|
| v1.0 | 2026-07-24 | Filed. |

Once collection opens, §§2–6 are frozen.
