Clarity pass over the main text (36-item audit), Discussion rewrite and cut, acknowledgements, Souly et al. as ref 62, lettered SI panels, model section moved under Results; plus the untracked curriculum/society/compose/smol configs, runners, figures, stats and tests that the SI already cites. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm |
||
|---|---|---|
| .. | ||
| grounding.pdf | ||
| grounding.png | ||
| manifest.json | ||
| README.md | ||
| resolved_config.yaml | ||
grounding — the grounding response in real weights (and why the ruler matters)
Claim tested: does the E2 result — a small dose of real data (g* ≈ 0.05) rescues diversity —
reproduce in a trained GRU? The honest answer reframes the question.
Setup (Layer 1.5). Autoregressive GRU, K = 256 modes, n = 200, 30 generations, grounding
swept over 9 values g ∈ {0, 0.005, …, 0.2}, 18 repeats (many repeats are needed because each
lineage's fate is genuinely noisy under n = 200 drift). The falsifier was pinned in the config
before running.
Symbols
ggrounding fraction (share of real data),g*its critical value.- forward-KL distance from truth (the operative neural collapse metric here).
- tail survival
tail_truth_mass_alive— truth-weighted fraction of the rare tail retained.Hdiversity. - recovery fraction — how much of the achievable forward-KL reduction a given
ghas bought (0 = dry, 1 = best observed).
The four panels
- Trajectories. Forward-KL over generations per
g: grounding suppresses the climb. - Phase boundary. Stationary forward-KL vs
gfalls monotonically (dry ≈ 2.08 →g = 0.2≈ 0.75); the effect is statistically significant (paired t up to 3.3; 89% of lineages improve atg = 0.2). The SIGN is confirmed. - Recovery curve. Fraction of the divergence gap closed vs
g. Half the gap closes by a median-recovery grounding ofg ≈ 0.04(CI [0.004, 0.116]) — a striking echo of Layer-1's 0.048 (black dashed) — but full recovery needsg ≈ 0.19, far more than the exact histogram: the GRU's smoothing both caps the collapse and slows the rescue. - Why forward-KL (the key methodological panel). Normalised responses of three rulers vs
g:H/H*(flat ~0.8) and tail survival (flat / non-monotone — dry is as high as grounded!) both fail to register the effect, while forward-KL recovery responds cleanly. A smoothing model keeps spurious tail support alive, so counting surviving modes is misleading; only distance-from-truth is honest.
Takeaway (an honest reframing)
Two results: (1) the operative neural collapse metric is forward-KL, not H or tail-survival —
smoothing decouples "modes alive" from "close to truth." (2) The sharp threshold g* ≪ 1 is a
property of the exact operator, carried quantitatively by the histogram bridge (g* = 0.047); the
trained GRU confirms grounding's direction and softens its sharpness. The pre-registered
95%-of-H*/tail falsifier is not met — but because those are the wrong rulers for a smoothing
model, not because grounding fails; the blueprint §3.5 directional claim holds robustly.