MachineSex/results/collapse_null/README.md
Giorgio Gilestro ab3dc10587 Restructure: descriptive tier and experiment names, paper/manuscript
- paper/pnas -> paper/manuscript (venue-neutral)
- configs/layer1 -> configs/inheritance, src/knowledge -> src/inheritance
  (imported as `inheritance`), make layer1 -> make inheritance; layer2 alias dropped
- inheritance and trained-network bundles named after the manuscript figure
  they feed (fig2_grounding_sweep, figS3_rebaselining, ...), or descriptively
  where they feed none; configs keep their `experiment:` value so parquet
  hashes are unchanged, only output.dir moves
- figure scripts, SI figure sources, notebooks, REPRODUCING.md, README and the
  SI Methods/tables updated; make clean no longer deletes tracked manifests;
  reproduce.sh hashes the s{seed}/ layouts too

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 17:00:40 +01:00

32 lines
2.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# E1 — Distillation without grounding collapses, tail-first
**Claim tested:** if a model is trained only on the previous model's output, generation after
generation, does it lose knowledge — and does the *rare* knowledge go first?
**Setup (Layer 1, pure math).** A "population" of `K = 500` items with a fixed true frequency
`p*` shaped like a Zipf curve (a few common items, a long tail of rare ones). Each generation we
draw `n = 100` samples from the current model and refit — **no real data is ever added** (`g = 0`).
Run for 600 generations, averaged over 100 independent repeats.
### Symbols
- **`p*`** — the true frequencies (fixed reality). **`p_t`** — the model's frequencies at generation *t* (drifts).
- **`H`** heterozygosity = diversity (1 = everything equally likely, 0 = one item left). **`H*`** = diversity of the truth.
- **forward-KL** `D(p*‖p_t)` — how far the model has drifted from truth (0 = perfect, grows without bound as the tail is forgotten).
- **support** = how many items still have any probability. **head/tail** = common/rare items.
### The three panels
1. **Geometric decay.** Blue = the simulated diversity `H`; black dashed = the exact textbook law
`H₀·(1 1/n)^t`. They sit on top of each other — the loss of diversity is *exactly* the
population-genetics drift law, not an approximation. (This is the validation gate: if these two
curves disagreed, the simulator would be wrong.)
2. **Tail dies first** (log axis). Red = fraction of *rare* (tail) items still alive; green =
fraction of *common* (head) items still alive. The red curve plunges far faster — rare knowledge
is lost roughly an order of magnitude sooner than common knowledge.
3. **Collapse.** Purple (left axis, log) = number of distinct items surviving, falling from 500
toward ~1 (everything collapses onto a single dominant item). Orange (right axis) = forward-KL to
truth, diverging as the tail vanishes.
### Takeaway
Unchecked model-on-model training is a ratchet: diversity decays on a precise mathematical schedule,
and the rare tail is destroyed first. **Falsifier (not triggered):** if `H` had stayed flat, the
whole thesis would fail. It didn't.