Windowsill Lab · Field Explainer · Astronomy From Open Archives

Every candidate carries
its own odds of being nothing

Search a thousand stars and some of them will show a convincing dip by pure luck. The only useful question is how likely it is that this one did.

1,174distinct stars searched
256shuffles per candidate
0 of 25placebos survived
MILESTONE IN PROGRESS · PROGRESS 0.6 · NOT PROMOTED
Everything on this page is reported, not graded. Nothing here is submitted anywhere.
A05

Making every candidate pay for itself

If you search a thousand stars, some of them will show a convincing dip by pure luck — so how do you work out the odds that this particular candidate is just noise wearing a costume?

AI-painted illustration: a wide windowsill crowded with rows of small glass specimen jars each holding one indistinct pale speck, one jar standing slightly forward with its lid off and a blank brass tag tied to its neck, a magnifying lens resting face-down beside it
Illustration (AI-painted) — one jar set forward for a second look

A04 measured a false-alarm floor for the survey as a whole: across the sample, junk topped out at a score of 6.6, so a threshold of 8.0 sat in a real gap. That is a good answer to the question "how often does this pipeline fool itself?" It is not an answer to "how likely is it that this star fooled it?" — and once you are searching thousands of stars rather than dozens, the second question is the only one that matters.

A05's answer is to make every candidate pay for itself. Take the star's own measurements and shuffle them, destroying any real signal while keeping the noise exactly as noisy as it truly was. Search the shuffled version. Do it 256 times. Now you have 256 examples of the best dip this specific star's noise can produce when there is definitely nothing there, and you can see directly where the real candidate falls among them. That fraction is its false-alarm probability — measured from its own data, not assumed from a table.

Panel one — the null, built in front of you

Live · shuffle, re-search, drop the peak onto the null SYNTHETIC LIGHT CURVE — INTUITION ONLY
0 draws

building the null…

left · the star, and the same star with its measurements scrambled · right · the best score each scrambled version managed. The two nulls largely overlap; the difference that matters is in the right-hand tail, which is the only side the probability is read from — watch the 95th percentiles in the line above rather than the peaks.

The shuffling is done two ways, because the naive way cheats. Scrambling every point independently destroys the slow correlated wiggles real stars have, making the noise look tamer than it is and the candidate look better than it is. So a second scheme shuffles in contiguous blocks, keeping those wiggles intact, and the grade is taken from whichever of the two is less flattering.

The panel is worth being honest about here, because it does not show what you would expect it to. The naive story is that the block null should sit visibly further right than the independent one. It does not — the two overlap almost completely, and their peaks land in the same place. The reason is that the score is already normalised against its own period grid, so a change in the overall noise character largely divides out before it can move the distribution. What does survive is a heavier right-hand tail under block shuffling, and since a false-alarm probability is read entirely from that tail, the take-the-worse-of-two rule still earns its place: it costs nothing on the draws where the two schemes agree, and it is the only thing standing there on the draws where they do not. The correction is real and it is smaller than the sales pitch — which is the kind of thing a page about not overclaiming ought to say about its own centrepiece.

Around that sits the rest of a survey. One target in ten is a deliberate placebo, its timing scrambled so that nothing real can survive — if a placebo ever produces a candidate, the calibration is broken and the run says so. A pre-search census kills about a third of every sector before a single period is tried, on the grounds that known binaries, known variables and puffed-up giant stars produce planet-shaped lies at industrial scale. And every survivor gets a disposition from a fixed vocabulary.

Panel two — the ladder, and where it stops

The disposition cascade · the 26 real rows, falling through it COUNTS FROM THE RECEIPTS
26 above-threshold detections

Every above-threshold detection falls through these gates in order. The gate that catches it lights up. What survives all of them reaches the last box — and the last box is the last box.

The vocabulary's best possible verdict is lead — awaiting human review. There is no higher state. The machine cannot output the word planet, and the count of discoveries cannot be raised by code at all. That ceiling is deliberate, it is enforced in the receipt contract rather than by convention, and it is the most important thing on this page.

The two pilot runs are why. One candidate held planet-candidate for forty minutes before its own frequency spectrum revealed it as a pulsating star. Another matched a catalogued object whose community verdict is false positive — which is not a recovery and not a lead. Two later leads survived the machine and then died within two minutes of a human checking external catalogues: one flagged as a spectroscopic binary, the others implying planets 2.3 and 2.8 times the radius of Jupiter, which is not a planet, it is a small star. The census exists because of that night.

Disposition, as it actually firedCount
known-planet7
recovery-or-known5
stellar-pulsation4
eclipsing-binary-secondary3
lead-awaiting-human-review3
harmonic-alias2
eclipsing-binary-odd-even1
insufficient-coverage1

Panel three — the floor that rises under you

The measured noise floor against sample size THREE MEASURED POINTS — HELD, NOT INTERPOLATED
n = 22

Search more stars and luck gets more chances, so the bar a candidate must clear has to climb. Three sample sizes were measured; the line is held flat between them rather than interpolated, because a value between two measurements is not a measurement.

The floor moved from 6.6 at n = 22 to 7.65 at n = 153 to 7.875 at n = 551. It is closing on the 8.0 that A04's much smaller sample made look comfortable. The two-point triage heuristic that was supposed to anticipate the n = 551 floor missed it by 0.47 SDE — in the conservative direction, but a miss, and never graded.

1,214target rows attempted, seven full hunts MEASURED
1,164searches that completed MEASURED
1,174distinct (sector, TIC) pairs touched MEASURED
570targets in the pilot phase — not 728 MEASURED
26above-threshold, all dispositioned, over 18 distinct pairs MEASURED
3 / 2lead rows over distinct stars MEASURED
0 of 25placebos that produced a candidate MEASURED
KS 0.236control probabilities uniform, as calibration requires MEASURED
37 % / 35 %of each sector killed before search MEASURED

The counting deserves a note of its own, because it is the kind of thing that goes wrong quietly. 1,214 target rows were attempted across the seven full hunts; 1,164 of those searches completed once 41 skips and 9 errors are removed; and they cover 1,174 distinct (sector, TIC) pairs, because 36 targets appear in more than one hunt. The pilot phase holds 570 targets and not 728 — the 570-target receipt declares that it supersedes the 158-target one, and the receipt contract excludes superseded receipts precisely so that cumulative runs cannot double-count themselves. Three sectors are involved: sector 2 on the Windows box, sectors 3 and 30 split to the Linux box.

The two surviving leads are TIC 234518605 at SDE 8.42, flagged in two consecutive sector-2 hunts, and TIC 272357134 at SDE 8.08. Both have dossiers on disk. Both are leads. Neither is a planet, and no code path in this survey can make either one into a planet.

Two named objects carry the lesson forward. TIC 140940493 held planet-candidate for forty minutes until its amplitude spectrum unmasked a δ Scuti-type pulsator. And TIC 278866211, at SDE 10.3, matched TOI 189.01, whose community disposition is false positive — so it carries toi-known-fp and becomes a test target for the blend gates rather than a result.

Per-target statistics over a pre-declared slice of one sector's two-minute targets: each graded number is an empirical false-alarm bound, each above-threshold detection carries a machine disposition from a closed vocabulary, and each surviving host carries its own measured depth limit. The survey claims completeness of disposition over that sample — every hit named, no hit promoted — not completeness of detection, not an occurrence rate, and no discovery: the terminal machine state is lead — awaiting human review, nothing is submitted anywhere on the strength of these receipts, and survey-level sums are reported, never graded. One further honesty: the shuffles are of the detrended flux, so the detrend has already absorbed some genuinely random low-frequency variance and every probability here is optimistic by some margin — the block scheme and the take-the-worse-of-two rule are partial compensation, not a fix. This milestone is in progress and has not been promoted.