Learn — Method

Calibrate Before You Search

Any search for rare things will find some. Five habits from a small planet survey for knowing which finds to believe, and they carry over to an A/B test, a model eval, or a log search at two in the morning.

The windowsill lab searches public data from NASA's TESS telescope for the small, regular dip in starlight that a passing planet makes. So far it has run 12,898 searches across seven patches of sky. The search itself is thirty-year-old math and a few hundred lines of code. Everything below is about the hard part: deciding what a hit means. The full account is in the survey paper; this is the part you can take home.

Price the noise first

The search tries 3,000 possible orbital periods on every star and reports the best one. Something is always best. Even a star with no planet at all will hand you a winning period, so the question is never "did it find a peak" and always "is this peak taller than the tallest peak noise makes?"

That last word matters. Comparing a hit against noise at the same period ignores the other 2,999 tries; with that many tickets, one is bound to look lucky. So for each promising star, the lab shuffles the star's own brightness measurements 256 times, reruns the full search on every shuffle, and records the tallest peak each time. The real peak has to beat that whole distribution.

Then it measured the threshold itself. Across 325,000 shuffled draws, the tallest noise peak reached 8.65, above the survey's cut of 8.0. Thirteen draws crossed the line. Scaled to 12,898 searches, that predicts roughly one false crossing in the whole survey. Of the 116 stars that crossed, expect about one to be pure noise, and expect the numbers alone to be unable to say which. Writing that sentence down before anyone asks is the whole habit.

The everyday version: if you test twenty button colors, one will win at the 5% level by chance. Shuffle your labels, rerun the test, and see how often "a winner" shows up anyway.

Plant fakes and see if you catch them

A search that cannot find a planet you put there cannot tell you anything about the planets it missed. So the lab plants them. It has injected 31,644 artificial transits of known size into real light curves and recovered 20,538, about 65%.

It does this star by star, which is the part people skip. For each of 1,565 stars it measured the shallowest dip it could still recover. On 384 of them, the search could not recover a 1% dip at any period it tried. Those stars are barred from every statement about absence: "we found nothing around this star" would be true and meaningless, because the instrument was blind there.

There is a free version of this control too. The search turned up nine planets other astronomers had already confirmed. Nobody told it where they were; each was matched to the catalogue only after the search had reported its period. Known answers that you did not hand the search are the cheapest calibration there is.

Feed it nothing and make sure it finds nothing

The opposite test: take 1,850 real light curves and scramble the timing of every measurement. Same brightness values, same gaps, same noise, and no repeating signal anywhere. Then run them through everything, the search and every vetting step after it.

They produced zero candidates. One would have been too many. The shuffling in the first habit checks the statistic; this checks the whole machine behind it, including every filter and judgment call, for its talent at manufacturing a discovery out of static.

Leave the word you want out of the vocabulary

Every star that crosses the threshold gets a verdict from a closed list of eighteen terms: eclipsing binary, stellar pulsation, harmonic alias, known planet, and so on, ending at lead awaiting human review. The list has no word for "planet". A run that tries to write one is thrown out. Promoting a lead to a planet is something a person does, in writing, and no code path can do it.

If the thing you are hoping to find is a value your code can emit, your code will eventually emit it. Make it a decision instead.

One more lesson came free. For two days in August the verdict list existed as two copies, one in the code that wrote verdicts and one in the code that graded them. The writer learned five new words and the grader did not, so the first time those words were used, the grader would have rejected the entire run and quietly sent its stars back into the pool. The fix was to make it one list read by both. A rule written down twice is two rules.

Count what you threw away

Three numbers belong next to every result, and most write-ups leave all three out.

  • The real denominator. 12,898 is a count of searches. Some stars were searched twice, so the count of stars is between 12,151 and 12,702. Work is counted in searches, sky in stars, and the two are never averaged.
  • The runs you refused. Five whole runs failed their own controls and are excluded from every total, by name. One of them was the only source of a lead, and the lead left with it.
  • The leads that died. The search raised seven leads. Its own checks killed five, one is parked on a named gap in the data, and one is still open: TIC 374861595, a 10.5% dip on a small red star. No planet is claimed.

A record of refusals is cheap to keep and almost nobody publishes one. It is also the most useful thing a small search can hand the next person.

Run it on your own search

before-you-announce.md — copy this

# Calibration checklist: [search name]

## Noise
- How many things did I try?                  [count]
- Best result from shuffled / random input?   [value]
- Expected false hits at my threshold:        [count]
- Can I tell which hit is the false one?      [yes / no, and say so]

## Plant fakes
- Fakes planted, fakes recovered:             [n / n]
- Where is my search blind?                   [segments, sizes, ranges]
- Known answers it found without being told:  [list]

## Feed it nothing
- Scrambled inputs run end to end:            [count]
- Hits on scrambled inputs:                   [should be 0]

## Vocabulary
- Can my code emit the word I am hoping for?  [it should not]
- Is the verdict list one copy, read by all?  [yes / no]

## Count
- Denominator, in the unit that matters:      [n]
- Runs refused, and why:                      [list]
- Leads raised / killed / open:               [n / n / n]

For a model eval, the fakes are cases where you know the right answer, and the scrambled input is the same test set with its labels shuffled. For an A/B test, the scrambled input is an A/A test. The habits stay the same; only the nouns change.

Next door

The survey's own record is in the paper, and the one open lead has a page of its own. For the stretch before a search returns anything, when you are stuck, read the Filip Method.

❦
← learn the windowsill the paper