Removed from main (all preserved on the dev branch): the arXiv build and
its sources, design documents (blueprint, results summary, review responses,
essay drafts), tasks/ and CLAUDE.md, the cover letter and reference tooling,
two unused manuscript figures, and every experiment that feeds no figure or
number in the paper: the collapse null, the sexual-vs-asexual lineage, the
NK speciation variant, the 0.5B single-seed LLM prototypes, the compose and
society experiments with their calibration and pilot runs, and their
configs, runners, tests, figure scripts and PBS jobs. Their result bundles
are moved to results/_archive/ (ignored) so the parquets stay on disk.
Also: plot_llm_speciation reads the s{seed}/ layout; the mating-breadth
plot writes under its bundle name; Makefile targets reduced to the kept
experiments; REPRODUCING.md and README point to dev for the rest.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
|
||
|---|---|---|
| .. | ||
| figS6_grounding_rnn.pdf | ||
| figS6_grounding_rnn.png | ||
| manifest.json | ||
| README.md | ||
| resolved_config.yaml | ||
grounding — the grounding response in real weights (and why the ruler matters)
Claim tested: does the E2 result — a small dose of real data (g* ≈ 0.05) rescues diversity —
reproduce in a trained GRU? The honest answer reframes the question.
Setup (Layer 1.5). Autoregressive GRU, K = 256 modes, n = 200, 30 generations, grounding
swept over 9 values g ∈ {0, 0.005, …, 0.2}, 18 repeats (many repeats are needed because each
lineage's fate is genuinely noisy under n = 200 drift). The falsifier was pinned in the config
before running.
Symbols
ggrounding fraction (share of real data),g*its critical value.- forward-KL distance from truth (the operative neural collapse metric here).
- tail survival
tail_truth_mass_alive— truth-weighted fraction of the rare tail retained.Hdiversity. - recovery fraction — how much of the achievable forward-KL reduction a given
ghas bought (0 = dry, 1 = best observed).
The four panels
- Trajectories. Forward-KL over generations per
g: grounding suppresses the climb. - Phase boundary. Stationary forward-KL vs
gfalls monotonically (dry ≈ 2.08 →g = 0.2≈ 0.75); the effect is statistically significant (paired t up to 3.3; 89% of lineages improve atg = 0.2). The SIGN is confirmed. - Recovery curve. Fraction of the divergence gap closed vs
g. Half the gap closes by a median-recovery grounding ofg ≈ 0.04(CI [0.004, 0.116]) — a striking echo of Layer-1's 0.048 (black dashed) — but full recovery needsg ≈ 0.19, far more than the exact histogram: the GRU's smoothing both caps the collapse and slows the rescue. - Why forward-KL (the key methodological panel). Normalised responses of three rulers vs
g:H/H*(flat ~0.8) and tail survival (flat / non-monotone — dry is as high as grounded!) both fail to register the effect, while forward-KL recovery responds cleanly. A smoothing model keeps spurious tail support alive, so counting surviving modes is misleading; only distance-from-truth is honest.
Takeaway (an honest reframing)
Two results: (1) the operative neural collapse metric is forward-KL, not H or tail-survival —
smoothing decouples "modes alive" from "close to truth." (2) The sharp threshold g* ≪ 1 is a
property of the exact operator, carried quantitatively by the histogram bridge (g* = 0.047); the
trained GRU confirms grounding's direction and softens its sharpness. The pre-registered
95%-of-H*/tail falsifier is not met — but because those are the wrong rulers for a smoothing
model, not because grounding fails; the blueprint §3.5 directional claim holds robustly.