The journey, step by step
Each step names what to do, the file in the kit that covers it, and how you know you’re done. Go in order. Step 0 is the one people skip, and step 10 is the one that turns a run into a survey.
Learn what a transit is
You don’t need a physics degree. You do need to know the three things that most often look like a planet and aren’t.
Done when you can explain why a dip that repeats isn’t automatically a planet. · LEARN.md
Set up
Clone the lab, install numpy, and run
doctor. It checks your Python, the lab code, and whether NASA’s archives answer from your machine.Done when
doctoris all green, or you know which line isn’t and why. · README.mdCalibrate on a star whose answer you know
Run WASP-18 (TIC 100100827), a confirmed hot Jupiter, first. If your setup can’t re-find a known planet, nothing it finds anywhere else means anything yet.
Done when WASP-18, sector 2, ends
already-known.Pick your star
Good first picks are bright, not crowded, and observed in two or more months, so a lead can be checked in a second month. “It’s the one from the viral post” is a fine reason, as long as you say so.
Done when you have a TIC number, its sectors, and a reason. · DATA.md
Write the question down, and commit it
The kit writes the form for you. You fill in the question and why this star, and commit it before you open the light curve.
Done when the form is committed and you haven’t looked yet. · PREREGISTRATION.md
Run the ladder
About three minutes per month of data on one core. The blind search, the chance check, the missed-dip floor, the tests, the catalog check and the placebo all run. You don’t pick which.
Done when you have a receipt that passes the kit’s own check. · PROTOCOL.md
Read the verdict
explainturns the receipt into one word and one sentence you are allowed to say.Done when you can say the word and its sentence out loud.
Decide what’s next
A lead in one month is provisional: the next step is the same search, written down again, on another month. Everything else gets written down as it is.
Done when a lead has a second search planned, or the result is recorded. · PROCEDURES.md
Share
A refuted star shared with its receipt helps the field more than a “candidate” with none. Put the no-token line first.
Done when your post uses only words from the table below and links the receipt and the commit. · BEFORE-YOU-POST.md
Find your people
ExoFOP, the TESS follow-up program, Planet Hunters TESS, AAVSO and NASA’s Exoplanet Watch each want different things from you.
Done when you know which venue fits your result and what it asks. · COMMUNITIES.md
Come back
Every receipt you keep joins your own survey.
ledgerregrades each star across all your receipts, so a one-month lead turns persistent, or doesn’t, when you search another month. TESS returns to most of the sky every couple of years.Done when your ledger lists every star you’ve searched and what each one needs next. · JOURNEY.md
Not ready to run code? NASA’s Exoplanet Watch and Planet Hunters TESS let you help with a browser, and the kit’s learning ladder goes from never having seen a light curve to knowing where results go.
What each word lets you say
This is the table to keep open when you write the post. It comes straight from the kit.
| The kit says | You can say | You can’t say |
|---|---|---|
nothing-above-threshold | “I searched TIC X in sectors Y and found no repeating dip. On this star the search could have seen transits down to Z %.” | “There are no planets here” (shallower ones would be missed) |
refuted | “My search found a signal and explained it: it’s [the plain reason from the receipt].” | anything about a planet |
already-known | “My search independently re-found [the catalogued signal].” A real result: it shows your pipeline works. | “I found it” (somebody else did, first) |
lead-awaiting-human-review | “I found a repeating dip that my tests couldn’t explain. It is a lead, not a planet.” | “I found a planet”, “a new world”, “NASA confirmed it” |
incomplete / control-failed | “The run didn’t finish / my own controls failed, so it says nothing.” | anything else |
From a dip to a confirmed planet is eight steps, and only the first five involve you: a signal, a lead, the same lead in a second month, a published paper, and filing it as a community candidate. After that come the TESS team’s own list, statistical validation and a measured mass, and those belong to astronomers. Getting telescope time to look again is a step toward those. It isn’t any of them. The full ladder is in the kit.
Try it
Start with the star whose answer you already know. The kit should end already-known. If it doesn’t, fix your setup before you go looking for anything new.
git clone https://github.com/benskamps/windowsill-lab && cd windowsill-lab
pip install numpy # the kit needs nothing else
export PYTHONPATH=src
# 1. Write the question down first, then commit it.
python -m lab.planetkit prereg 100100827 --sectors 2 --out my-prereg.md
git add my-prereg.md && git commit -m "prereg: TIC 100100827 s2"
# 2. Run the ladder. --download fetches the declared sectors from MAST.
python -m lab.planetkit run --prereg my-prereg.md --download
# 3. Read what you're allowed to say.
python -m lab.planetkit explain receipt-TIC100100827-*.json
# 4. Later: every star you've searched, and what each one needs next.
python -m lab.planetkit ledger
Have NASA’s light-curve files already? Use --fits. Only a time,flux table? Use --csv, and the receipt will say which tests couldn’t run without the extra columns. --quick checks the plumbing in seconds, and its receipts are marked so you won’t post one by accident.
What a result looks like
TIC 100000001: lead-awaiting-human-review
My search found a repeating dip that every automatic test failed to explain.
It is a lead for a human to review, not a planet. It is in one sector only,
so it is provisional: the next step is the same search on another sector.
sector 20: SDE 9.5, P = 3.1791 d, word = lead-awaiting-human-review, placebo passed
! The preregistration was never committed, so nobody can tell it was
written before the data were seen.
Do not say: I found a planet; I discovered a planet; NASA confirmed it; a new world
That is a real full-size run of the kit on a synthetic star: a 0.4 % dip every 3.18 days planted in one simulated month of data. The star is fake. The output is not.
Every run also draws a folded light curve for each month it searched: every dip lined up on top of each other, captioned “A picture, not a verdict.” python -m lab.planetkit_dashboard turns a receipt into one offline page with the system, the star and the candidate side by side. It draws nothing the verdict doesn’t: a refuted star shows why it was refuted, and nothing gets an orbit it hasn’t earned. DASHBOARD.md has the details.
With Claude Code, or any agent
The kit is a Claude Code plugin. Add the lab as a plugin marketplace, install it, and ask Claude to “find a planet the windowsill way”:
/plugin marketplace add benskamps/windowsill-lab
/plugin install find-your-own-planet@windowsill-lab
The plugin installs the rules, not the code: if your folder isn’t a clone of the lab, it tells Claude to clone it first. Not on Claude Code? The skill folder is plain markdown any agent can read. It holds the agent to the same rules as hard lines: write the question first, use the kit’s runner instead of a hand-rolled search, report the verdict word for word, never fix a failed control, and never write that anyone found a planet. Ask it to draft a “we found a planet” post and it will draft the honest one instead.
The ten rules underneath
The kit is these rules, enforced in code. PROTOCOL.md gives each one its reason and where it came from.
Does the kit follow them? CONFORMANCE.md checks the runner against all ten, row by row, naming the line of code that enforces each one. Nine are enforced. The tenth is only partly: the kit prints nothing but receipt values, but it can’t check a post you write yourself. Where the kit bends a rule on purpose, it says so there: an uncommitted question is allowed for a calibration run, and every receipt from that run says so.
- Write the question down first, and date it publicly.
- One threshold for every star.
- Give every signal its own null, in two schemes, and take the worse one.
- Measure what this star could have shown you.
- Run a placebo through the whole ladder.
- Walk a fixed ladder of refutations, and record every rung.
- Speak only in a closed vocabulary, with no word for “planet”.
- Refuse, don’t repair.
- A lead is a property of a star, not of one sector.
- Every number in the write-up comes from the receipts.
Is any of this new? The individual tests aren’t. Shuffled-data nulls, planted-dip recovery and automated vetting are how NASA’s Kepler mission vetted its catalogue. What is unusual is putting them in one package for one person with an AI agent, and governing the words: the machine is only allowed to refute, and only a human can move a lead forward.
What the kit doesn’t do (yet)
The runner is the lab’s survey pipeline, and that pipeline stops short of the professional field in a few places. They are listed so you don’t mistake a gap for a pass.
| Gap | What it means for your result |
|---|---|
| No statistical validation (TRICERATOPS) | Nothing here can move a small-planet lead past “candidate”, even in principle |
| No difference imaging of its own | The on-target check uses NASA’s centroids, which crowding can bias. CSV input has no on-target check at all |
| Fixed 0.5-day detrend | Transits longer than about 5.5 hours lose depth |
| BLS only, no TLS | TLS recovers more Earth-size signals at the same false-alarm rate |
| One sector searched at a time | Periods beyond about 9 days, and signals too shallow for one month, are missed |
| 2-minute targets only | Most of the sky’s faint stars are out of reach |
These are known, written down, and not quietly patched: changing the pipeline changes every receipt the lab has already published. PROCEDURES.md has the full comparison with sources.
Keep going
- TIC 374861595: the lab’s open lead, the whole method on one real star.
- The survey paper: the draft every number on this page comes from.
- Fold a known planet and search stars nobody pointed you at: the method as small hands-on rooms.
- Calibrate before you search: the same habits, written so they carry over to A/B tests and model evals.
- Cite the kit: if you use it in a post, a class or a paper,
CITATION.cffhas the reference, and GitHub’s “Cite this repository” button reads it. - The live feed: what the lab’s own computers are searching right now.