How it works, and how to catch us out
We do not tell you what happened to you. We tell you who else said something like it, and exactly why we put those accounts together. This page is the whole method, including where it fails.
The test we set ourselves
Anyone can build something that finds patterns. The hard part is proving it finds real ones. So we handed the engine 53,038 accounts stripped of any labels, told it nothing about history, and asked a simple question: would it find the famous nights on its own?
| the night | where it ranked | found |
|---|---|---|
| Tinley Park, Halloween | 2 | 25 accounts |
| Tinley Park, third night | 3 | 13 accounts |
| Tinley Park, first night | 4 | 16 accounts |
| The Phoenix Lights | 5 | 22 accounts |
| Tinley Park, fourth night | 6 | 11 accounts |
| Rockford, Illinois | 15 | 12 accounts |
| Southern Illinois, police witnesses | 66 | 4 accounts |
| O'Hare Airport | 85 | 3 accounts |
| Stephenville, Texas | 255 | 4 accounts |
| Yukon, Canada | not found | see below |
The point of that test is not the nine. It is that the same engine also flags nights nobody has ever heard of, and finding the famous ones is what earns the right to take those seriously. There are 24 of them, checked against the ordinary record and still standing.
Where it failed
It missed the Yukon entirely. Not because the engine is blind, but because the archives we started with are heavily American and barely record where Canadian accounts came from. That is a gap in the record, not a mystery, and we are filling it. We would rather show you the miss than a table of nine wins.
How two accounts get connected
Three things, all of them visible to you on every match.
Does it read the same?
Your words are compared against every account in the record (where every one of them comes from is here). Not keyword search, meaning. Someone writing “a tall dark shape crossed the road” and someone writing “a big black figure walked in front of my car” are saying the same thing.
Did you both notice something unusual?
Sharing something rare counts far more than sharing something common. Half the record mentions lights. Almost nobody mentions wood knocking. If you both mention wood knocking, that matters.
How far apart were you?
Closer counts for more, but distance never rules a match out. Some of the strangest connections in the record are between people on opposite sides of a continent.
Under the hood, for people who want it
This section is more technical than the rest of the page. Skip it if you would rather just read the accounts.
Finding candidates is the hard part, not ranking them
The obvious way to build this is to search for accounts that read like yours and put the best ones first. We measured that approach against a set of accounts we already knew belonged together, and it found a genuine match only 20% of the time. Worse, no amount of adjusting how results were ranked moved that number, because ranking cannot promote something the search never returned.
So the engine gathers candidates two ways at once. One is meaning: accounts that describe the same thing in different words. The other is structure: accounts from the same region, the same period, or sharing a characteristic rare enough to be worth noticing on its own. The two pools are combined before anything is scored. On the same measurement, that took the hit rate from 20% to 100%.
The record answers backwards
Adding a report to an ordinary database does nothing to the reports already in it. They sit where they were. Here, every account that arrives is compared against the whole archive, so a new one can answer an old one, and the record gets better in both directions.
That is measurable rather than a claim, so we measured it. Of the 6,456 published accounts that have both a date and a closest match, 3,303 found that match in an account written later than their own. That is 51%. The median gap was 6 years, 1,324 waited a decade or more, and the longest wait in the record is 88 years.
Which means roughly half the people here were answered by somebody who had not written their account yet when they wrote theirs. It is also the honest reason to write yours down even if nothing matches today. The measurement is retrospective and we say so: it runs the same comparison over history that the engine runs on arrivals. It is deterministic, and we will hand the exact method to anybody who wants to disagree with it.
Measuring strangeness instead of assigning it
The idea we are leaning on is not ours and it is not new. J. Allen Hynek, the astronomer the US Air Force hired to explain its UFO reports away, ended up proposing that every case be rated on two separate axes: how credible the witnesses were, and how strange the account was, meaning how much of it survives after every ordinary explanation has been applied. High strangeness is the phrase that came out of that, and it is a better description of what this record collects than any of the category names people usually reach for.
Hynek assigned his strangeness ratings by hand, one case at a time, using his own judgement. That is the part worth improving on. Here the strangeness of a detail is not a matter of opinion, it is measured: how rarely does this appear across every account in the record? A light in the sky is not strange, because a great many people report one. The woods going silent at the same moment is, because almost nobody does. The engine reads that number off the corpus rather than off anyone's intuition, and prints it on the page so you can disagree with the corpus rather than with us.
Rare details count for more
Every characteristic carries a weight set by how often it appears in the whole record. Hovering appears in about seven accounts in a hundred, so two people both mentioning it means very little. Wood knocking appears in well under one in a hundred. When two strangers both describe that, the engine treats it as real evidence and tells you the percentage on the page so you can judge it yourself.
Twenty-five accounts are not always twenty-five witnesses
Before any group of accounts is called significant, the engine compares them for overlapping runs of words. Accounts that echo each other, whether through news coverage or a story being retold, score low on independence and are marked down accordingly. The same comparison run across the whole record found duplicate accounts that had reached us through more than one route, and those never count twice. Both numbers are published on the pages they affect.
The measure has a second edition, adopted August 2026 before the record had readers to confuse. The first edition compared vocabularies, which quietly punished people for describing the same thing with the same words. The second compares runs of three words, because measurement on seventeen thousand real pairs showed that independent writers share almost no exact phrases while copying is unmistakable. Where both numbers appear, they are labelled. Nothing already published was silently changed, and the blind validation stands exactly as originally run.
Nothing is asked of a language model
No generative model reads your account and decides what it means, writes a word of this site, or judges any connection. The matching is done by a memory system built for the purpose, which is why the same question gives the same answer tomorrow, why any connection can be reconstructed afterwards from the record of how it was made, and why searching the entire archive costs nothing per query and can be offered to everybody rather than rationed behind a subscription. Where a measurement uses an embedding model to compare meanings, that model is named in the measurement's own script, so the result can be rerun and argued with like everything else here.
How the counts on the seeking pages are measured
A page like shadow people once counted every account that shared a surface detail, and a reader caught what that produced: animal eyes and treeline silhouettes wearing a bedroom phenomenon's name. Counts on those pages now come from a stricter instrument: a topic is defined by many example phrasings and, just as important, by counter-examples of what it is not; every account is read sentence by sentence against both; and the surviving candidates face a second, sharper comparison before they count. Then the result has to survive an audit: the accepted items are tested against known look-alike families (a light vanishing into the sky is not a glitch in reality; woods unease is not a bedroom presence), the mix of sources has to make sense for the phenomenon, and an independent check must agree with a large majority of the samples. A topic that fails any of that is refused automatically and its page falls back to a simple text-search count, labelled by its looseness. Refusal is a result here, not an embarrassment: it is what keeps every number on these pages meaning what it says.
The instrument also has to know what a sentence is not saying. Comparison by meaning has a documented blind spot: it barely registers negation, so “I have never seen a shadow figure in my bedroom” reads to it almost exactly like the sighting it denies. We probed our own classifier with denial sentences for every topic and eleven of fourteen were accepted as evidence. That gap is now closed: a sentence that denies the experience or explains it away is disqualified from counting, while negation that is the experience itself, “I could not move”, “nobody was there”, still counts. The denial probes are part of the standing audit, and any change to the instrument has to pass them clean before its numbers reach a page. We found this one because a measurement we ran in public forced us to look. That is the argument for doing them in public. The full post-mortem is at /notes/negation.
How the record is actually held
Underneath all of this the archive is not stored as rows in a table. Every account, every characteristic, every place and every date is held as a point in a space of a thousand dimensions, and the relationships between them are held in a hyperdimensional knowledge graph. In a space that large, two unrelated things are almost guaranteed to point in unrelated directions, which means agreement between two accounts is a signal rather than a coincidence. It also means a fact and its opposite can be folded into the same representation and pulled apart again cleanly later, so the record can carry contradictions instead of quietly resolving them.
The practical effect is the whole reason this site can exist. Asking the record a question is arithmetic on vectors rather than a scan through several hundred thousand documents, so it answers in the time it takes a page to load, it answers identically every time, and it costs so close to nothing per query that we can hand it to everybody instead of charging for it. This is not a technique bolted onto an off-the-shelf search product. It is the thing we built first, and this site is what happened when we pointed it at a subject that badly needed it.
What we could not find anywhere else
There are larger archives than this one and there are sites that will find you similar reports. What we went looking for, and did not find, was any of them showing the reasoning behind a connection, measuring whether witnesses were independent of each other, checking each cluster against the astronomical and seismic record for an ordinary cause, or working across kinds of experience instead of one category. If you know of one, please tell us, and we will link to it.
What is actually doing the searching
The connections are found by a memory system called Neruva, built by the same person who built this site. It is not a chatbot and there is no language model reading your account. That distinction matters more than it sounds, for four reasons.
It gives the same answer twice
Ask it the same question tomorrow and you get the same result, in the same order, for the same reasons. A system that quietly changes its mind cannot be checked by anybody, and a record that cannot be checked is just a rumour with a database behind it.
It can be replayed
Every connection can be reconstructed afterwards: what was searched, what came back, and how it was scored. When we publish a claim like the nine of ten result above, that claim is a re-runnable procedure rather than a screenshot.
It costs nothing to ask
Searching the whole record does not bill anybody by the question, so we never have to ration it, put it behind a subscription, or quietly search less of the archive to save money. One person running this on a modest budget can let everybody search everything.
It shows the reasons
Because the connection is assembled from parts we can name, we can hand you the parts. That is why every match on this site arrives with a sentence explaining itself instead of a confidence percentage you are asked to take on faith.
How the substrate represents and retrieves all this is the technology behind it, and we keep that under the hood. What we owe you is not the engine schematic, it is the reasoning behind every claim on your screen, and that is on the page every time.
When we have nothing
If the record holds nothing genuinely like your account, we say so and show you nothing. It would be easy to fill the page with five loose matches and let you draw a line between them. That is how these places go wrong. “You may be the first” is a real answer here.
The same account twice is still one account
If one report reaches us through two archives, or somebody submits theirs twice, our counts inflate and the site shows corroboration that never existed. That is the error most likely to discredit a record like this, and it is a known plague of every large report database.
So before anything is counted, the record is checked for accounts that are the same account written twice, by comparing overlapping runs of words rather than trusting dates and places. We found 38 clusters and set aside 39 accounts, 0.08% of the record. They stay readable. They just never count as a second witness.
Twenty-five people are not always twenty-five people
If a story reaches the news, accounts start echoing each other. So before we call a night significant we check whether people described it in their own words or in the same borrowed ones. Our own headline case gets marked down for this: the Halloween night in Tinley Park has the most accounts of any night in the record, and the weakest independence score. We publish that number next to it.
We check for the boring answer first
Most places in this field leave “could it have been something ordinary?” as an argument in a comment thread. We compute it. Every night the engine detects is checked against the public record of what was genuinely happening at that place and hour:
- Fireballs, from NASA's near-earth object survey. One honest limit: that catalogue begins in 2015, which is later than every clustered night currently in the record, so this check is armed for new nights rather than doing work on the old ones.
- Earthquakes, from the USGS catalogue, which explains a surprising number of booms, rumbles and shaking houses.
- The moon, because a full moon behind cloud has fooled a great many people, including sober ones.
- Venus, calculated for that place and that hour. It is the most misidentified object in the sky, by a distance.
- Aurora, from the geomagnetic Kp index for that day, which reaches back to 1932. A storm big enough to push the lights south is the largest single producer of nights when a whole region reports strange lights at once.
- The calendar, because orange lights rising silently on the fourth of July are usually sky lanterns.
Of the first 120 nights we checked, 79 had at least one ordinary candidate, and it is written on the page. The other 41 had none. That second number is the one worth caring about, and it only means anything because we went looking for the first.
A candidate is not a verdict. A fireball the same evening does not prove that is what somebody saw, and we do not close a case because a computer found a coincidence.
The meteor
One night the engine flagged as significant turned out to be a documented fireball over the Great Lakes in 1999. It had no idea what a meteor was. It simply noticed ten strangers describing the same thing within an hour. We keep that one on the site, marked answered. An instrument that cannot find ordinary explanations has no business pointing at strange ones. Read it
What we do with your account
- A person reads it before anything appears publicly.
- Locations are always blurred. We never publish an address.
- Nobody gets your contact details. If two people want to compare notes, it goes through us, both sides have to agree, and either can walk away.
- We never mock a witness and we never diagnose one.
- Ask us to remove your account and we remove it.
Still sceptical?
Good. Send us the account that should break this, or the ordinary explanation we have missed. Both make the record better.
Tell us what you saw