The data
Everything this site computes is downloadable, licensed for reuse, and citable. If you want to check a number rather than take it on trust, start here.
DOI: 10.5281/zenodo.22037108 · CC BY 4.0 · v1.0.0
What is in it
| file | rows | what it holds |
|---|---|---|
| accounts.csv | 43,684 | One row per account: source, the id it carries in that archive, date, hour, coordinates, place, and which characteristics it contains. |
| characteristics.csv | 63 | The vocabulary, with how common each detail is and its inverse document frequency. |
| clusters.csv | 263 | Same-night, same-place groups of three or more accounts, scored for wording independence and coherence, and cross-referenced against ordinary causes. |
The column worth your attention
idf in characteristics.csv. Two accounts both mentioning a light means nothing, because nearly everything in a sky archive mentions a light. Two accounts both mentioning three knocks means a great deal. The inverse document frequency is that difference, expressed as a number, and it is what lets a match be explained rather than asserted.
What is deliberately not in it
The account text. The narratives belong to the archives that collected them and to the people who wrote them, so republishing tens of thousands of them as a bulk download is not ours to do. Every row carries its source and the id it holds there, so the original is one lookup away.
Before you quote the interesting number
184 of the 263 clusters have no ordinary cause on the public record. That figure is computed, not asserted: count the rows in clusters.csv where ordinary_causes_found is zero. But read the limit that matters first. Fireball data begins in 2015 and historical earthquake coverage is uneven, so a zero on an old night partly reflects a thinner catalogue rather than only a stranger event.
The rest of the limits are in the README inside the download, and they are there because a dataset that only advertises its strengths is not a dataset, it is a brochure. Characteristics are matched by curated regular expressions, which makes every hit explainable and every miss silent. The corpus is two archives, one about lights in the sky and one about something in the woods, so it holds effectively nothing about sleep paralysis, the hat man, or shared false memory, which are among the most searched and least collected experiences there are.
Citing it
Clouthier, K. (2026). The Anomaly Network characteristic layer (v1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22037108
The method behind the numbers
How characteristics are extracted, how clusters are detected, and how the engine was blind-tested against famous documented events.
How this works Add your own account