MachineSex/results/collapse_null/README.md
Giorgio Gilestro ab3dc10587 Restructure: descriptive tier and experiment names, paper/manuscript
- paper/pnas -> paper/manuscript (venue-neutral)
- configs/layer1 -> configs/inheritance, src/knowledge -> src/inheritance
  (imported as `inheritance`), make layer1 -> make inheritance; layer2 alias dropped
- inheritance and trained-network bundles named after the manuscript figure
  they feed (fig2_grounding_sweep, figS3_rebaselining, ...), or descriptively
  where they feed none; configs keep their `experiment:` value so parquet
  hashes are unchanged, only output.dir moves
- figure scripts, SI figure sources, notebooks, REPRODUCING.md, README and the
  SI Methods/tables updated; make clean no longer deletes tracked manifests;
  reproduce.sh hashes the s{seed}/ layouts too

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 17:00:40 +01:00

2.2 KiB
Raw Blame History

E1 — Distillation without grounding collapses, tail-first

Claim tested: if a model is trained only on the previous model's output, generation after generation, does it lose knowledge — and does the rare knowledge go first?

Setup (Layer 1, pure math). A "population" of K = 500 items with a fixed true frequency p* shaped like a Zipf curve (a few common items, a long tail of rare ones). Each generation we draw n = 100 samples from the current model and refit — no real data is ever added (g = 0). Run for 600 generations, averaged over 100 independent repeats.

Symbols

  • p* — the true frequencies (fixed reality). p_t — the model's frequencies at generation t (drifts).
  • H heterozygosity = diversity (1 = everything equally likely, 0 = one item left). H* = diversity of the truth.
  • forward-KL D(p*‖p_t) — how far the model has drifted from truth (0 = perfect, grows without bound as the tail is forgotten).
  • support = how many items still have any probability. head/tail = common/rare items.

The three panels

  1. Geometric decay. Blue = the simulated diversity H; black dashed = the exact textbook law H₀·(1 1/n)^t. They sit on top of each other — the loss of diversity is exactly the population-genetics drift law, not an approximation. (This is the validation gate: if these two curves disagreed, the simulator would be wrong.)
  2. Tail dies first (log axis). Red = fraction of rare (tail) items still alive; green = fraction of common (head) items still alive. The red curve plunges far faster — rare knowledge is lost roughly an order of magnitude sooner than common knowledge.
  3. Collapse. Purple (left axis, log) = number of distinct items surviving, falling from 500 toward ~1 (everything collapses onto a single dominant item). Orange (right axis) = forward-KL to truth, diverging as the tail vanishes.

Takeaway

Unchecked model-on-model training is a ratchet: diversity decays on a precise mathematical schedule, and the rare tail is destroyed first. Falsifier (not triggered): if H had stayed flat, the whole thesis would fail. It didn't.