Windowsill Lab · a patient home instrument

windowsill

day · spring · feels

The goal is a measurement nobody has made yet, given away free.

The measurement

Tc = 2.269, solved by hand in 1944. This run measured 2.30.

128×128 magnets, 21 temperatures from 1.5 to 3.5, 21.0 seconds on a consumer GPU — in the model’s own units, not degrees. Every mark below is plotted straight from that run’s committed receipt — nothing is redrawn by hand, and you can diff it yourself. A calibration, not a discovery: nothing on this page is claimed as a new result. (The plant below follows the newest turn; this panel holds the newest calibration. Its three lattice frames name the run they came from, which is not always this one — see the line beneath them, and the provenance at the panel’s foot.)

T = —ordered · one domain wins
T = —nearest saved frame to Tc
T = —disordered · thermal noise

Left, the sheet is cold and almost everything agrees. Right, it is hot and nothing does. The middle frame is the nearest of the three to the exact tipping point Lars Onsager solved by hand in 1944 — Tc = 2.269 — and it is the only one that looks like neither: patches of both colours at mixed sizes, instead of one colour winning or none. How close that middle frame really sits to the tipping point depends on which temperatures the run saved a picture of; its caption gives the number, so you can judge the gap yourself. Each frame is magnets, straight from the run.

Susceptibility χ(T) — how twitchy the sheet is. The peak is where the smallest nudge in temperature causes the biggest change in behaviour. The equilibrium-qualified peak is T ≈ —, against Onsager’s exact 2.269. The sweep steps in 0.1, so its peak can only be reported to the nearest grid point — a spacing coarser than the upward shift a finite 128×128 sheet produces. This run calibrates against Onsager to that resolution; it does not resolve that shift. Reading the quality guard…
Magnetization |m|(T) — how much the sheet agrees with itself. The measured points (dots) fall on the exact curve (line), computed from the formula rather than fitted to these dots. The exact result is C. N. Yang’s, published in 1952, building on Onsager’s 1944 solution of the lattice. Their agreement is the calibration. Per-point uncertainties are the receipt’s own abs_mag_err — the largest, 0.004, is smaller than these dots.
Specific heat C(T) — how much heat the sheet drinks. A different measurement entirely — heat, not magnetism — and its peak has to land near the same temperature. Two different measurements landing in the same place, to the resolution of this grid, is the instrument checking itself.

Provenance arrives with the feed: run date, lattice size, seed, device, wall clock, commit, and the equilibrium guard’s verdict.

Two computers in a house work the same experiment ladder, a turn scheduled every three hours, around the clock199 turns so far. linux-rocm took the last turn; windows-cuda has been quiet since 20 Aug.

AI agents wrote this instrument and keep it running; a human signs every promotion, and the receipt names who read the evidence — an agent cannot decide that a result is true.

The experiment shelf · the rooms behind the runs

28 experiments on the record, and most open into a live room — the real update rule running in your browser, receipt beside it. 23 reproduced their targets, 3 awaiting a human read, 2 came back null.

walk the shelf →

Below: seven pots on one windowsill, one for each science. Every leaf is one real experiment — green means a person checked it, amber is waiting for one, grey is a miss and it stays on the plant. And the pots wear the work: the glaze climbs as a track climbs, the scratches count its runs, the chips are its failures.

milestones watered days tended
local scene feed source snapshot.json live explainer
The curriculum

Trust first. Then harder questions.

The point of the ladder is the top of it: a measurement nobody has made yet — a small, real addition to physics, published free for anyone to check, reuse, or run again on their own machine. That is the destination, not the status: nothing on this page is claimed as a new result. The lab earns the top of the ladder by climbing the first three rungs in public, misses included — which is the only reason anyone would believe the fourth.

  1. 01Verify the instrumentM01–M05 · exact answers
  2. 02Map known territoryM06–M10 · new models
  3. 03Push the edgeM11–M14 · disorder
  4. 04Ask open questionsM15–M18 · non-equilibrium
The conservatory

Seven instruments. One standard of proof.

Height tracks how far each track has climbed. Amber waits for a human read; grey is a null.

What am I looking at?

A small garden on a windowsill, wired to real work. Two quiet computers take turns simulating physics, checking arithmetic, reading public telescope archives, or listening through hardware sensors — on a three-hour schedule, around the clock, 199 turns in. Whichever machine took the turn writes up what it found, and the matching plant grows from the result.

Every layer underneath was written by AI agents: the GPU physics, the checker that independently re-derives each result, the scheduler, the report format, the published feed, the tests — and this page you're reading. Agents also design the experiments and draft the write-ups. One person set the direction and holds the gate.

The gate is the part worth understanding. An agent cannot decide that a result is true. A deterministic check has to pass first — that earns an amber leaf — and then a promotion signed by a person turns it green. Reading the evidence is sometimes delegated to another agent, and every receipt names who did it. When a run misses, the miss stays on the plant as a grey folded leaf, with its receipt. The milestone counts under the sill are the whole lab's, across every science, nulls included.

The lower leaves are calibration: phase transitions and magnetic models whose answers are known well enough to expose a bad instrument. Higher leaves move into disorder and dynamics, where the claims must get narrower as the questions get harder. Tap any leaf—or use the milestone rail—to see its question, finding, and technical receipt. The growing tip says whether the next question has a runner ready or is still only on the bench.

How to read the plant

green leaf
a machine-checked result promoted onto the permanent record by a person, with the receipt naming who read the evidence.
amber leaf
a measurement whose checker passed, waiting for a human to read the evidence before it can turn green.
grey folded leaf
an experiment that missed. It stays on the plant, and its field note says why.
the soft growing tip
the question at the front of the curriculum. Its field note says whether code exists to run it yet.
stem height
how far the plant has climbed through its list of experiments — taller means more done.
dark, damp soil
a fresh run just finished and watered it. The soil slowly dries until the next turn.
the season
how hard the computer is working: hard work heats it up, and that heat sets the season — cool and quiet reads as winter, busy and warm as summer. It shows spring when a run doesn't report a temperature.
the sky & light
your own time of day, in real time — dawn, noon, dusk, night. The plant keeps the same hours you do, wherever you are.

Why do this?

A pair of home computers has something a busy university supercomputer doesn't: slow, patient time, and no line of people waiting. So they take the unglamorous jobs nobody's in a hurry to run — one small experiment at a time, turn after turn, filling a real notebook over months. First they re-check answers we already know, to earn their trust; once proven steady, they can wander toward corners of science nobody has gotten around to mapping yet.

The reason this can exist at all is newer than the hardware. A patient instrument like this was never blocked on flops — it was blocked on the months of human evenings it takes to write the simulation, the checker, the provenance, the scheduler and the write-up, for a question nobody is paying you to answer. Agents absorb that cost. What one person still has to supply is the judgment about which answers count, and that has not been handed over.

Whatever comes out is meant to be given away. Every run's numbers, every receipt, every kept null, the whole instrument and this page are public and MIT-licensed: you can read how a result was reached, re-derive it from the saved measurements, fork the lab onto your own GPU, or point it at a question you care about instead. A small new result anyone can check and build on is worth more than a private one. The commons is the destination, not a byproduct.

For the curious — the real names and numbers

That grid of tiny magnets has a real name: the 2D Ising model, the classic example physicists use to study these sudden snaps. The exact tip-over temperature (written T_c, equal to 2.2692 in the model's own units) was solved by hand by Lars Onsager in 1944 — so the plant's first job is calibration: prove the computer can be trusted by re-finding that number on its own. Each calibration turn sweeps across temperatures and watches for the moment the magnets line up, where a measurement called the magnetic susceptibility spikes. In the calibration plotted above it landed at 2.30 ± 0.05 against Onsager’s 2.2692, in 23.2 seconds on the GPU (the graphics chip that also runs video games). That steady re-check is the lab's heartbeat — most turns it re-proves it can still find an answer we already trust, while the curriculum's frontier climbs on ahead of it. The counter under the plant counts turns — one scheduled pass, whichever machine took it; before July 2026 it counted nights. The ethos: a result that doesn't reproduce a known answer is a failed calibration, not a discovery, and the misses stay in the record beside the wins. So far several milestones are verified — beginning with the 2D Ising critical point and climbing through finite-size scaling, critical exponents, other lattices and models, and on toward spin glasses — out of a 40-step curriculum whose physics ladder climbs in four phases: verify famous exact results, map textbook-known territory, push corners last explored in the 1990s on big computer clusters a single modern chip now out-muscles, then reach genuinely open questions. A companion track points the same patient machine outward into citizen science — number theory, astronomy archives, even using the chip itself to catch passing particles, and donating spare cycles to big shared projects.

How the machine works — the engineering under the calm

The calm surface is the last mile of a real instrument. Underneath, one automated pipeline runs itself, and two machines take alternating turns running it:

schedule → simulate → verify → report → publish → grow

A scheduled task wakes whichever machine has the turn and asks the curriculum for a runnable next step. If the frontier is still only a design, the lab records that fact and runs its M01 calibration heartbeat instead. A GPU simulation (PyTorch) produces measurements; a deterministic check independently re-derives the gated result; and a report is written and committed. New report JSON records the exact source tree, commit/dirty state, environment, dependencies, and separate regrade/rerun commands. Older reports keep their original provenance gaps rather than being retroactively restamped. A small versioned JSON feed (pot.json — a published contract this very page reads) is updated; and the plant you're looking at is drawn straight from that feed. Continuous integration re-runs the whole test suite on every change.

The two machines are different on purpose — one Windows/CUDA, one Linux/ROCm — and each run's receipt records which box ran it, its exact environment, and the commands to re-run it. The lab began as one machine's night shift; a second machine joined in July 2026, and the two now alternate.

Every layer above was agent-written — the GPU kernels, the verification gate, the scheduler, the provenance, the JSON contract, the tests and CI, and the live drawing that renders the plant — from the silicon to the seedling. The commits are in the public history if you want to read how it was actually built, misses included.

The live feed, the source, and the full curriculum live at windowsill-lab; the rest of the lab is at the lab.