The undeletable copy
Can you make a quantum copy disappear? Can you make an AI truly forget?
Known, reproduced
Hand someone two copies of an unknown quantum state and ask them to blank one. The no-deleting theorem (Pati and Braunstein, 2000) says any reversible process can only move the copy, never destroy it. Quantum rules are reversible, so whatever wipes copy 2 must leave its information in a helper qubit (the "trash bin").
The notebook asked an optimizer for the best possible deleter. It rediscovered a swap: the copy goes into the helper. Then it tried the AI version. Tiny networks told to "unlearn" a secret stop saying it, yet the secret is still sitting in their weights.
Run the deleter
starting…
Each trial draws a random unknown state ψ (seeded), loads it into copy 1 and copy 2, and leaves the helper blank. The gate is cos β · 1 − i sin β · SWAP on copy 2 and the helper. It is reversible at every β, and β = 90° is the full swap the optimizer found. Discs show each qubit's Bloch arrow (side view); a shorter arrow means that qubit is entangled with the others. Right · each new trial is paired with the last few: x = how alike their inputs were, y = how alike their trash (copy 2 + helper) ended up. A true erasure would make all trash identical and put every dot on the top line. Every dot lands on the diagonal at every β.
In plain words
A reversible process keeps "how alike are these two states" fixed. If copy 2 ends up blank for every input, that likeness has to live somewhere else, and the only other place is the helper. Blanking a copy is therefore a move. The optimizer, given free rein over every 3-qubit gate, found exactly that. Its deleter blanked copy 2 almost perfectly, and the trash pairs matched the input pairs. Forbid the helper from learning anything and the best machine blanks the copy only about 66% of the way.
The AI side is an analogy. A small network memorized 60 secret labels. Gradient-ascent unlearning (pushing the trained weights away from the secrets) got recall to 0%. Yet the secret still sat in the net's top 3 guesses 48% of the time, against 21% for a net retrained without the secrets. In a separate local rerun, a net "forgotten" by fine-tuning on everything else relearned its old secrets about 8 points faster than brand-new ones, in 31 of 32 runs. Small edits from weights that already hold a record tend to hide it, much like the swap. Only retraining from scratch and discarding the old model forgets for sure.
Prior artPati and Braunstein 2000 (the theorem). Thudi et al. 2022 showed unlearning can't be proven from the weights alone. The unlearning part is an analogy, and the theorem does not bind classical nets.
Next clickThe Landauer cost of forgetting, against the compute floor for exact unlearning.
- The live panel runs the qubit half only. Its gate is a hand-picked partial swap that shows the rule; the notebook's deleter came from an optimizer.
- Neural networks are classical. Ordinary bits can be deleted when you know what they are, so no-deleting is not a law for AI. The unlearning results are a pattern seen in tiny nets (5 seeds, 60 secrets), and they are not a theorem.
- The relearning gap comes from a separate local rerun, not from the 5-seed run.