One self-contained README.md per results/ figure folder (Layer 1 E1-E6 and Layer 1.5 bridge/collapse/grounding/architectures/recombination): plain-language claim, setup, a compact symbol glossary, a panel-by-panel walkthrough, and the takeaway + falsifier. Auto-renders when browsing the folder; carries the honest caveats (grounding's ruler reframing, the excluded VAE). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2.2 KiB
2.2 KiB
E1 — Distillation without grounding collapses, tail-first
Claim tested: if a model is trained only on the previous model's output, generation after generation, does it lose knowledge — and does the rare knowledge go first?
Setup (Layer 1, pure math). A "population" of K = 500 items with a fixed true frequency
p* shaped like a Zipf curve (a few common items, a long tail of rare ones). Each generation we
draw n = 100 samples from the current model and refit — no real data is ever added (g = 0).
Run for 600 generations, averaged over 100 independent repeats.
Symbols
p*— the true frequencies (fixed reality).p_t— the model's frequencies at generation t (drifts).Hheterozygosity = diversity (1 = everything equally likely, 0 = one item left).H*= diversity of the truth.- forward-KL
D(p*‖p_t)— how far the model has drifted from truth (0 = perfect, grows without bound as the tail is forgotten). - support = how many items still have any probability. head/tail = common/rare items.
The three panels
- Geometric decay. Blue = the simulated diversity
H; black dashed = the exact textbook lawH₀·(1 − 1/n)^t. They sit on top of each other — the loss of diversity is exactly the population-genetics drift law, not an approximation. (This is the validation gate: if these two curves disagreed, the simulator would be wrong.) - Tail dies first (log axis). Red = fraction of rare (tail) items still alive; green = fraction of common (head) items still alive. The red curve plunges far faster — rare knowledge is lost roughly an order of magnitude sooner than common knowledge.
- Collapse. Purple (left axis, log) = number of distinct items surviving, falling from 500 toward ~1 (everything collapses onto a single dominant item). Orange (right axis) = forward-KL to truth, diverging as the tail vanishes.
Takeaway
Unchecked model-on-model training is a ratchet: diversity decays on a precise mathematical schedule,
and the rare tail is destroyed first. Falsifier (not triggered): if H had stayed flat, the
whole thesis would fail. It didn't.