MachineSex/results/figS6_grounding_rnn
Giorgio Gilestro 6f8cef1ac5 main: keep only what reproduces the manuscript; everything else lives on dev
Removed from main (all preserved on the dev branch): the arXiv build and
its sources, design documents (blueprint, results summary, review responses,
essay drafts), tasks/ and CLAUDE.md, the cover letter and reference tooling,
two unused manuscript figures, and every experiment that feeds no figure or
number in the paper: the collapse null, the sexual-vs-asexual lineage, the
NK speciation variant, the 0.5B single-seed LLM prototypes, the compose and
society experiments with their calibration and pilot runs, and their
configs, runners, tests, figure scripts and PBS jobs. Their result bundles
are moved to results/_archive/ (ignored) so the parquets stay on disk.

Also: plot_llm_speciation reads the s{seed}/ layout; the mating-breadth
plot writes under its bundle name; Makefile targets reduced to the kept
experiments; REPRODUCING.md and README point to dev for the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 17:07:23 +01:00
..
figS6_grounding_rnn.pdf main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
figS6_grounding_rnn.png Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
manifest.json Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
README.md Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
resolved_config.yaml Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00

grounding — the grounding response in real weights (and why the ruler matters)

Claim tested: does the E2 result — a small dose of real data (g* ≈ 0.05) rescues diversity — reproduce in a trained GRU? The honest answer reframes the question.

Setup (Layer 1.5). Autoregressive GRU, K = 256 modes, n = 200, 30 generations, grounding swept over 9 values g ∈ {0, 0.005, …, 0.2}, 18 repeats (many repeats are needed because each lineage's fate is genuinely noisy under n = 200 drift). The falsifier was pinned in the config before running.

Symbols

  • g grounding fraction (share of real data), g* its critical value.
  • forward-KL distance from truth (the operative neural collapse metric here).
  • tail survival tail_truth_mass_alive — truth-weighted fraction of the rare tail retained. H diversity.
  • recovery fraction — how much of the achievable forward-KL reduction a given g has bought (0 = dry, 1 = best observed).

The four panels

  1. Trajectories. Forward-KL over generations per g: grounding suppresses the climb.
  2. Phase boundary. Stationary forward-KL vs g falls monotonically (dry ≈ 2.08 → g = 0.2 ≈ 0.75); the effect is statistically significant (paired t up to 3.3; 89% of lineages improve at g = 0.2). The SIGN is confirmed.
  3. Recovery curve. Fraction of the divergence gap closed vs g. Half the gap closes by a median-recovery grounding of g ≈ 0.04 (CI [0.004, 0.116]) — a striking echo of Layer-1's 0.048 (black dashed) — but full recovery needs g ≈ 0.19, far more than the exact histogram: the GRU's smoothing both caps the collapse and slows the rescue.
  4. Why forward-KL (the key methodological panel). Normalised responses of three rulers vs g: H/H* (flat ~0.8) and tail survival (flat / non-monotone — dry is as high as grounded!) both fail to register the effect, while forward-KL recovery responds cleanly. A smoothing model keeps spurious tail support alive, so counting surviving modes is misleading; only distance-from-truth is honest.

Takeaway (an honest reframing)

Two results: (1) the operative neural collapse metric is forward-KL, not H or tail-survival — smoothing decouples "modes alive" from "close to truth." (2) The sharp threshold g* ≪ 1 is a property of the exact operator, carried quantitatively by the histogram bridge (g* = 0.047); the trained GRU confirms grounding's direction and softens its sharpness. The pre-registered 95%-of-H*/tail falsifier is not met — but because those are the wrong rulers for a smoothing model, not because grounding fails; the blueprint §3.5 directional claim holds robustly.