New Fig. 1 (experimental-programme schematic); Table 2 to SI; figures in citation order; Fig. 2B legible labels
Replaces the results table with a pipeline figure: five questions x three architecture tiers (exact Wright-Fisher simulator, trained networks, language models), filled cells naming the experiments, dashed cells the honest gaps. Table 1 (the dictionary) stays; Table 2 moves to SI Appendix Table S2. The renumber surfaced a pre-existing citation-order violation (the LLM figure was cited in the recombination section before Figs. 3-6), so figures are renumbered to strict first-citation order (LLM tier is now Fig. 3). Fig. 2B: the montage's baked-in raster labels are cropped away and replaced with vector row numbers under a rotated "generation" header. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
parent
073fc33509
commit
0159e2839a
13 changed files with 208 additions and 109 deletions
|
|
@ -96,6 +96,22 @@ long-horizon over-specialisation erodes mergeability at LLM scale (cf. arXiv:260
|
|||
| Emergent speciation without conflict | **Not observed** (pre-registered) | Shared ancestry, compatible tasks, tested divergences | E13b: residual 0.000; merge rescues specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
|
||||
| Grounding + sex + diversity jointly necessary | Exact-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |
|
||||
|
||||
## SI Table S2: headline quantitative results
|
||||
|
||||
Headline quantitative results with sample sizes, uncertainty, and outcome definitions (full
|
||||
per-experiment tables and falsifier status in the per-experiment documentation).
|
||||
|
||||
| Result | Setting / n | Outcome definition | Headline |
|
||||
|---|---|---|---|
|
||||
| Closed-form validation | Analytic tier; standing tests | Simulated vs closed-form H-decay, immigration equilibrium, multi-teacher union | Agreement < 0.5% |
|
||||
| Grounding retention | Minimal model; 18+ replicates per point | Fraction of equilibrium diversity retained at grounding g (operational threshold) | g ≈ 0.05 retained ≥95% (tested setting); smooth in g |
|
||||
| MNIST collapse & rescue | Conv-VAE, 4 replicates; frozen oracle (98.5% mode acc.) | Mode support / forward-KL over generations | Dry: 30→1 modes; 10% grounding: 30/30 held |
|
||||
| Fisher–Muller in LLMs | 5 seeds (0.5B), fixed tests; single 7B run | Merged vs best-specialist accuracy (overall; worst family) | Ties 0.647±0.027 vs 0.592±0.009; 7B 0.87 vs 0.77 |
|
||||
| Union vs blend (headroom) | 3 seeds (0.5B hard); single 7B-hard run | Paired per-seed ordering, routing vs weight-average | Routing > blend in 3/3 seeds; one catastrophic blend failure avoided |
|
||||
| Speciation decomposition | MLPs, 3 replicates | LMC error barrier residual after permutation+rescaling alignment | Same-task 0.001; conflict 0.497 (naive 0.502) |
|
||||
| Emergent isolation | MLPs 4 reps to 6.4× base training; LLM 1→12 epochs | Residual barrier; merged vs parent accuracy | 0.000 everywhere; merge rescues parents (≈0.955 vs ≈0.50) |
|
||||
| Predictive test | 13 conditions × 3 seeds (0.5B) | Merge penalty vs oracle parent potential (pre-registered; ±: clustered 95% CI) | Functional ρ +0.45/+0.46, CI excl. 0; LOCO ρ ≈ 0.4; geometry n.s.; paired differences n.s. |
|
||||
|
||||
## SI Methods (per tier — full details in the per-experiment READMEs and configs)
|
||||
|
||||
**Analytic tier (E1–E14).** Wright–Fisher simulator over K-item distributions; closed-form validation
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue