MachineSex/results/figS6_grounding_rnn
Giorgio Gilestro ab3dc10587 Restructure: descriptive tier and experiment names, paper/manuscript
- paper/pnas -> paper/manuscript (venue-neutral)
- configs/layer1 -> configs/inheritance, src/knowledge -> src/inheritance
  (imported as `inheritance`), make layer1 -> make inheritance; layer2 alias dropped
- inheritance and trained-network bundles named after the manuscript figure
  they feed (fig2_grounding_sweep, figS3_rebaselining, ...), or descriptively
  where they feed none; configs keep their `experiment:` value so parquet
  hashes are unchanged, only output.dir moves
- figure scripts, SI figure sources, notebooks, REPRODUCING.md, README and the
  SI Methods/tables updated; make clean no longer deletes tracked manifests;
  reproduce.sh hashes the s{seed}/ layouts too

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 17:00:40 +01:00
..
figS6_grounding_rnn.pdf Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
figS6_grounding_rnn.png Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
manifest.json Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
README.md Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
resolved_config.yaml Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00

grounding — the grounding response in real weights (and why the ruler matters)

Claim tested: does the E2 result — a small dose of real data (g* ≈ 0.05) rescues diversity — reproduce in a trained GRU? The honest answer reframes the question.

Setup (Layer 1.5). Autoregressive GRU, K = 256 modes, n = 200, 30 generations, grounding swept over 9 values g ∈ {0, 0.005, …, 0.2}, 18 repeats (many repeats are needed because each lineage's fate is genuinely noisy under n = 200 drift). The falsifier was pinned in the config before running.

Symbols

  • g grounding fraction (share of real data), g* its critical value.
  • forward-KL distance from truth (the operative neural collapse metric here).
  • tail survival tail_truth_mass_alive — truth-weighted fraction of the rare tail retained. H diversity.
  • recovery fraction — how much of the achievable forward-KL reduction a given g has bought (0 = dry, 1 = best observed).

The four panels

  1. Trajectories. Forward-KL over generations per g: grounding suppresses the climb.
  2. Phase boundary. Stationary forward-KL vs g falls monotonically (dry ≈ 2.08 → g = 0.2 ≈ 0.75); the effect is statistically significant (paired t up to 3.3; 89% of lineages improve at g = 0.2). The SIGN is confirmed.
  3. Recovery curve. Fraction of the divergence gap closed vs g. Half the gap closes by a median-recovery grounding of g ≈ 0.04 (CI [0.004, 0.116]) — a striking echo of Layer-1's 0.048 (black dashed) — but full recovery needs g ≈ 0.19, far more than the exact histogram: the GRU's smoothing both caps the collapse and slows the rescue.
  4. Why forward-KL (the key methodological panel). Normalised responses of three rulers vs g: H/H* (flat ~0.8) and tail survival (flat / non-monotone — dry is as high as grounded!) both fail to register the effect, while forward-KL recovery responds cleanly. A smoothing model keeps spurious tail support alive, so counting surviving modes is misleading; only distance-from-truth is honest.

Takeaway (an honest reframing)

Two results: (1) the operative neural collapse metric is forward-KL, not H or tail-survival — smoothing decouples "modes alive" from "close to truth." (2) The sharp threshold g* ≪ 1 is a property of the exact operator, carried quantitatively by the histogram bridge (g* = 0.047); the trained GRU confirms grounding's direction and softens its sharpness. The pre-registered 95%-of-H*/tail falsifier is not met — but because those are the wrong rulers for a smoothing model, not because grounding fails; the blueprint §3.5 directional claim holds robustly.