- paper/pnas -> paper/manuscript (venue-neutral)
- configs/layer1 -> configs/inheritance, src/knowledge -> src/inheritance
(imported as `inheritance`), make layer1 -> make inheritance; layer2 alias dropped
- inheritance and trained-network bundles named after the manuscript figure
they feed (fig2_grounding_sweep, figS3_rebaselining, ...), or descriptively
where they feed none; configs keep their `experiment:` value so parquet
hashes are unchanged, only output.dir moves
- figure scripts, SI figure sources, notebooks, REPRODUCING.md, README and the
SI Methods/tables updated; make clean no longer deletes tracked manifests;
reproduce.sh hashes the s{seed}/ layouts too
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
604 lines
52 KiB
Markdown
604 lines
52 KiB
Markdown
# Supplementary Information — The evolution of sex for artificial intelligence
|
||
|
||
## Contents
|
||
|
||
SI Text S1–S4, SI Tables S1–S2, SI Methods M1–M7, SI Statistics, SI Figures S1–S16, and a separate
|
||
Appendix 1, *The figures explained* (`figure_legends_for_students.pdf`), which restates every main and
|
||
supplementary figure with a plain-language account of the experiment behind it, for readers from biology.
|
||
|
||
## Reproducibility
|
||
|
||
Every experiment in this paper is defined by one committed configuration file under `configs/`.
|
||
Running it produces three artifacts under `results/<name>/`: the results table (`results.parquet`),
|
||
the fully resolved configuration, and a manifest recording content hashes, the master seed, and the
|
||
git commit. Each experiment directory also contains a README with the figure legend and the current
|
||
status of the experiment's falsifier — the outcome that would refute its claim (see Methods M1) —
|
||
plus a figure that regenerates from the parquet file alone. The script `reproduce.sh` re-runs the
|
||
entire study from the master seeds, and `REPRODUCING.md` maps every panel of the manuscript to the
|
||
configuration and seed behind it.
|
||
|
||
## SI Text S1. The incompatibility floor: what no alignment can remove
|
||
|
||
**Setting.** Two models, A and B, are trained on the same input distribution. Their label functions
|
||
`f_A` and `f_B` agree everywhere except on a *conflict set* `S`, whose size is its probability mass
|
||
`μ(S)`. In the conflict condition of the trained-network speciation experiment, `S` consists of the
|
||
cyclically relabelled classes, so `μ(S)` is approximately the configured conflict fraction, up to
|
||
class-balance corrections.
|
||
|
||
A *function-preserving transformation* `T` is any change to a network's weights that leaves its
|
||
outputs untouched. For a plain ReLU multilayer perceptron these transformations are exactly the
|
||
permutations of hidden units and the positive rescalings of individual units: scaling a unit's
|
||
incoming weights up and its outgoing weights down by the same factor does not change what the network
|
||
computes. Together they form the *unit symmetry group* of the architecture. By construction `T(B)`
|
||
computes the same function as B, that is `T(B)(x) = B(x)` for every input `x`.
|
||
|
||
**Proposition 1 (endpoint invariance).** Define the *chord* as the straight line connecting the two
|
||
endpoint loss values, `(1−α)·L(A) + α·L(B)`. It depends only on the endpoints and is the baseline used
|
||
in the definition of the interpolation barrier; it is not the loss along the interpolation path in
|
||
weight space. For every function-preserving `T`, the pair `(A, T(B))` has the same endpoint losses as
|
||
the pair `(A, B)`, and therefore the same chord. The interpolation path itself is generally not
|
||
invariant: the losses along `(1−α)·A + α·T(B)` change with `T`. This is exactly the room an alignment
|
||
has to lower a barrier. The proof is immediate from the definition of function-preserving.
|
||
|
||
*Scope of the alignment guarantee.* The aligner used here is guaranteed to recover a
|
||
permuted-and-rescaled copy of a network exactly. That is an important special case, but it does not
|
||
prove that the alignment is optimal over the whole symmetry group for independently trained networks.
|
||
Consequently the share of the barrier attributed to removable coordinate mismatch is a lower bound,
|
||
and the residual share an upper bound, on their true values.
|
||
|
||
**Proposition 2 (no merged model can serve both parents).** Let `h` be any single classifier; in
|
||
particular, any interpolated or merged model, under any alignment. On every input `x ∈ S` the two
|
||
parents disagree, `f_A(x) ≠ f_B(x)`, so `h` must disagree with at least one of them. Writing `ε_P(h)`
|
||
for `h`'s error rate against parent `P`'s labels,
|
||
|
||
`ε_A(h) + ε_B(h) ≥ μ(S)`, hence `max(ε_A(h), ε_B(h)) ≥ μ(S)/2`.
|
||
|
||
When two models' conventions conflict on a set of mass `μ(S)`, any hybrid of the two is wrong on at
|
||
least one parent's task at least `μ(S)/2` of the time. This floor is information-theoretic, holding
|
||
regardless of the alignment group, the architecture, or the merging operator. In the fitness sense it
|
||
is reproductive isolation: beyond a given functional conflict, no recombination operator can produce
|
||
an offspring faithful to both lineages.
|
||
|
||
**What remains empirical, and how the experiment is designed.** Propositions 1 and 2 do not bound the
|
||
single-task path barrier: the loss along the interpolation between A and `T(B)`, evaluated on one
|
||
parent's task alone. In principle such a path could dip toward one parent's function and yield a low
|
||
barrier even under conflict. Whether it does is an empirical question, and it is precisely what the
|
||
experiment measures. The measured answer is that it does not. In the conflict condition the barrier is
|
||
unchanged by permutation alignment (the `residual` readout) and by alignment modulo the full
|
||
permutation-and-positive-rescaling group (the `residual_scale` readout), while the very same aligner
|
||
removes almost all of the barrier between independently initialised networks, the positive control.
|
||
Work on richer symmetry groups for transformers (83) strengthens the removable side of the
|
||
decomposition and is therefore complementary to this result: the more barrier a larger group can
|
||
remove for *compatible* models, the sharper the meaning of the barrier that survives for
|
||
*incompatible* ones. Proposition 2 caps what any of these methods could ever achieve on the conflict
|
||
set.
|
||
|
||
**Terminology used in the paper.** "Residual (after alignment)" denotes the estimated functional
|
||
incompatibility: the part of the merge barrier that remains after the architecture's unit symmetries
|
||
have been divided out. For ReLU MLPs I align modulo the full unit symmetry group, so the estimate is
|
||
not confounded by symmetries of that architecture class that the aligner might have missed.
|
||
|
||
## SI Text S2. Emergent versus imposed incompatibility
|
||
|
||
The conflict condition *imposes* contradiction: the two label maps disagree on `S` by construction,
|
||
which pins `μ(S) > 0` and activates Proposition 2. A genuine Bateson–Dobzhansky–Muller
|
||
incompatibility is instead *emergent*. Each lineage's substitutions are harmless on their own
|
||
background, so the training signals never contradict and `μ(S) = 0`; any incompatibility appears only
|
||
when the two lineages are combined.
|
||
|
||
Two conditions realise this emergent setting. In `disjoint`, the parents are specialists on
|
||
complementary classes. In `augment`, they learn divergent input conventions on the same task. Neither
|
||
condition contains label conflict, so any barrier that survives alignment cannot be attributed to
|
||
label conflict. Such a barrier would be the emergent-speciation signal proper.
|
||
|
||
Both readings were registered before the run. If the residual barrier grows with divergence, then
|
||
model speciation is emergent in real weights, and the trajectory seen in the analytic speciation model
|
||
is realised. If the residual stays at the level of the `shared` control, then within this regime
|
||
trained networks are more merge-compatible than the biological analogy predicts. The second reading
|
||
would bound the analogy, and be a useful design result in its own right: merging is safe
|
||
whenever there is no functional conflict.
|
||
|
||
**Outcome.** Four replicates, with divergence up to 3,200 steps — up to 6.4× the shared base training
|
||
— returned the second reading. The residual barrier was 0.000 at every divergence in both emergent
|
||
conditions. Merging moreover *rescued* the `disjoint` specialists, which had forgotten the classes
|
||
outside their specialty: at the longest divergence the parents score 0.535 and 0.474 on the full task,
|
||
while the merged model holds approximately 0.955 at every divergence tested. This is a sustained
|
||
Fisher–Muller rescue at zero barrier. Within this regime, reproductive isolation in real weights
|
||
required functional conflict. The same question at language-model scale is answered by the duration
|
||
arm of the language-model speciation experiment, which likewise found no isolation from over-training
|
||
alone (1 to 12 epochs); whether still longer horizons erode mergeability (cf. 83) remains open.
|
||
|
||
## SI Text S3. Compatible loci and conflicting alleles in a multigenerational population
|
||
|
||
**The two kinds of new knowledge.** A *locus* is a position in the genome, and *alleles* are the
|
||
alternative versions that can occupy it: one blood-group locus, three alleles A, B and O, of which any
|
||
one chromosome carries exactly one. In a model population a locus is a slot for a capability ("how to
|
||
answer a two-way question") and alleles are the incompatible conventions that could fill it ("yes/no",
|
||
"true/false", "1/2"). A skill that conflicts with nothing a lineage already holds occupies a new locus
|
||
and is simply added; a skill that demands a different convention for a question shape the lineage
|
||
already answers is a competing allele, and a single model, like a single chromosome, carries one.
|
||
Proposition S2 gives the cost: when two parents' conventions disagree on a share `μ(S)` of inputs, any
|
||
merged child errs against at least one parent on at least `μ(S)/2` of them. In the six-generation
|
||
population a lineage obliged to merge at generation `t` pays that floor against its partner's
|
||
conflicting conventions; because the child continues the lineage, the loss is inherited, and the next
|
||
generation's conflict adds to it. Under the Latin-square curriculum `μ_t(S)` is zero while partners are
|
||
complementary (their skills occupy disjoint loci) and becomes positive once a partner carries a
|
||
differently conventioned version of a skill the lineage already holds. Two of the six families —
|
||
yes/no questions and two-way pronoun resolution — have the most idiosyncratic conventions and were
|
||
measured in calibration at 0.00–0.04 accuracy on every other family, so they carry the largest `μ(S)`
|
||
against every partner; the generation at which the curriculum hands them to a lineage's partner fixes
|
||
when that lineage's collapse begins.
|
||
|
||
**Negative controls that isolate convention conflict.** Three alternative explanations of the
|
||
obligate arm's collapse were tested directly and refuted. (i) *A destructive skill spreading through
|
||
merges.* A single 50/50 merge of two clean single-skill adapters (science questions 0.838 / yes-no
|
||
0.000; yes-no 0.800 / science 0.300) scored 0.863 and 0.787, mean 0.825 against 0.550 for the better
|
||
parent: one merge is protective, not destructive. (ii) *Geometric dilution of an adapter's signal
|
||
under repeated averaging.* Five chained convex merges left the first skill's accuracy unchanged even
|
||
though its nominal weight fell to 1/32; but a scaling control showed the adapter alone delivers
|
||
nothing at 1/32 (0.000; full effect down to 1/8), so what propagated through the chain was the answer
|
||
format supplied by whichever partner carried enough weight, not the skill. Dilution is refuted, and
|
||
the transmitted quantity is identified as the convention. (iii) *Continued training on merged
|
||
weights.* Merging then training on the incoming family beat merging alone on the tracked skill in four
|
||
of five rounds and on the incoming skill in all five, and absorbed the one format shock that dropped
|
||
the merge-only chain (0.567 → 0.883). With capacity ruled out by the lifelong-editing benchmark (80)
|
||
at three orders of magnitude more content, convention conflict is the mechanism that remains — the one
|
||
the framework predicts, and the one single-model studies report (81, 82).
|
||
|
||
**Neutral and functional variation.** Three adapters trained on the same family, differing only in
|
||
seed and data draw, were near-orthogonal in weight space (pairwise cosine +0.006) and disagreed on 24%
|
||
of answers, yet merging two of them gave 0.887 against 0.800 for the better one — exactly the fraction
|
||
of questions on which either was right (0.887). Decomposing the weight change across seeds, roughly
|
||
85% of a LoRA delta is run-specific: shared signal power 6.2 (after correcting the finite-sample mean
|
||
for its own noise) against noise power 35. That is why raw weight distance predicted nothing in the
|
||
main text's controlled test: most of what it measures is the counterpart of *synonymous substitution*
|
||
— sequence change without functional change — which averages out when adapters for the same skill
|
||
are combined, while the fraction that conflicts lives in the answer conventions. Averaging same-skill
|
||
adapters before crossing them with a different skill improved the cross modestly (0.825 → 0.850) while
|
||
leaving each single skill unchanged, the inbred-line pattern: averaging within a line does not improve
|
||
the line, it makes it cleaner to cross.
|
||
|
||
**Attenuation and the effectiveness cliff.** Scaling an adapter's weights down does not degrade its
|
||
skill gracefully. Each skill holds full accuracy to a skill-specific fraction (1/8 for science
|
||
questions, 1/4 for reading-comprehension spans, 1/2 for commonsense completion, 1/4 for yes/no) and
|
||
then loses nearly everything within one further halving. Four of six adapters scored higher when
|
||
attenuated (inference 0.40 → 0.68 at 1/4; completion 0.75 → 0.82 at 1/2; spans 0.72 → 0.78 at 1/4;
|
||
science 0.87 → 0.92 at 1/8): they were over-trained at full strength — the effect reported for
|
||
merging experts (84, 85) — and recoverable here by one scalar per adapter with no retraining (six
|
||
separate specialists 0.678 → 0.755). Denoising across seeds does not move the cliff, so the limit is
|
||
signal magnitude rather than signal-to-noise. Choosing per-skill merge weights from these solo curves
|
||
failed (0.686–0.689 against 0.708 for uniform weights): in a six-way merge a skill's effective
|
||
strength is its weight relative to the others — six conventions competing for one output — so raising
|
||
one starves the rest.
|
||
|
||
## SI Text S4. Proof of the blending-inheritance proposition
|
||
|
||
**Setting.** `K` parents; each independently retains a given rare item with probability `q`, and a
|
||
parent that retains it assigns it mass `p`. The child draws `n` samples from a *source distribution*
|
||
and keeps the item if at least one draw is that item. Two sources are compared: (A) one parent chosen
|
||
uniformly at random; (B) the mean of the `K` parents' distributions.
|
||
|
||
**Expected mass is conserved.** Let `J ~ Binomial(K, q)` be the number of parents retaining the
|
||
item. Under (A) the source mass of the item is `p` with probability `q` and 0 otherwise, so its
|
||
expectation is `pq`. Under (B) the source mass is `pJ/K`, whose expectation is `p·E[J]/K = pq`. The
|
||
expected number of copies in the child's sample, `n` times the source mass, is therefore `npq` under
|
||
both schemes (linearity of expectation).
|
||
|
||
**Survival agrees to first order.** Write `f(x) = 1 − (1 − x)^n` for the probability that at least one
|
||
of `n` draws hits an item of source mass `x`; `f` is increasing and concave, with `f(x) = nx + O((nx)²)`.
|
||
Survival is `E[f(M)]` with `M` the (random) source mass. Under (A), `E[f(M)] = q·f(p)`; under (B),
|
||
`E[f(M)] = E[f(pJ/K)]`. When `n·p ≪ 1`, every realised mass satisfies `nM ≤ np ≪ 1`, so `f(M) ≈ nM`
|
||
and both expectations reduce to `n·E[M] = npq`: the `1/K` dilution of scheme (B) is cancelled exactly
|
||
by the item being present in the mixture whenever any of the `K` parents holds it. (Equivalently, in
|
||
this regime the child's copy count is approximately Poisson with mean `nM`, and Poisson thinning by
|
||
`1/K` composed with a `K`-fold union preserves the mean.)
|
||
|
||
**Boundary 1 (common items).** Away from the first-order regime the comparison is settled by
|
||
Jensen's inequality. Both schemes give `M` the same mean `pq`; scheme (A) puts all its variance in
|
||
the two-point distribution `{0, p}`, and scheme (B) has strictly smaller variance for `K > 1`. Since
|
||
`f` is concave, `E[f(M)]` is larger for the less variable `M`, so averaging never lowers expected
|
||
survival, and raises it once `np` is not small. The extinction probability `1 − f` is convex, which is
|
||
the form in which the main text states this boundary. The proposition is thus a statement about rare
|
||
items, where survival is linear in mass; it does not claim averaging is harmful in general.
|
||
|
||
**Boundary 2 (union operator).** Let the child instead draw from the distribution that assigns each
|
||
item the largest mass any parent gives it, renormalised. The item's source mass is then `p` whenever
|
||
`J ≥ 1`, an event of probability `1 − (1 − q)^K`, increasing in `K` for every `q ∈ (0, 1)`. Expected
|
||
survival `(1 − (1 − q)^K)·f(p)` therefore rises with `K` in every regime, without a first-order
|
||
restriction. The operator needs an oracle (a verifier) to say which parent holds each item most
|
||
strongly, which is what routing supplies in the language-model tier.
|
||
|
||
Both statements are confirmed by simulation in Fig. S8, where mean-mixture survival is flat in `K`
|
||
and the item-wise maximum rises with it.
|
||
|
||
## SI Table S1: the claims ledger (status / assumptions / evidence / limits)
|
||
|
||
| Claim | Status | Key assumptions | Evidence | Known limits |
|
||
|---|---|---|---|---|
|
||
| Population collapse in the inheritance model is Wright–Fisher drift | Closed form; the diagnosis itself is due to prior work | Knowledge is a categorical distribution; refitting means resampling | Closed forms reproduced to <0.5% | Real learners add a signed, architecture-specific estimator bias (measured) |
|
||
| Grounding behaves like immigration, and the critical real-data fraction is far below one | Closed form, plus the sign confirmed empirically | Fresh samples from a fixed, non-drifting truth | Exact `H_eq`; `g*≈0.048`; sign holds in RNN/MLP/VAE and on MNIST | Deepest tail unrescuable at feasible budgets (`m ∼ 1/p`); sharp threshold softens in trained nets |
|
||
| "Merge, don't average" conservation | Exact **for the output-mean operator** | Rare-item regime; an oracle/verifier identifies the strongest source | `figS8_multiparent_union` closed form + simulation; neural reproduction | Weight-averaging and routing are empirical cousins, not instances; budgets differ; bridge = the headroom rule |
|
||
| Offspring exceed every parent (Fisher–Muller) | Interpretation + empirical | Complementary (decorrelated) parents; verifiable fitness | `figS9_specialist_superparent` (inheritance model); LoRA merges beat the best specialist overall in every seed at 0.5B (5 seeds) and 7B (3 seeds) | LLM tier: 3 lexically-distinct families |
|
||
| Outbreeding depression on rugged landscapes; operator design rule | Biological-model result; hypothesis at LLM scale | NK epistasis stands in for skill entanglement | `figS10_rugged_landscapes`, `figS11_directed_recombination`; directed selection rescues | Not yet mapped onto a real task-entanglement measure |
|
||
| Optimal mate-pool breadth shrinks with ruggedness | Biological-model result; hypothesis for merging populations | Ring population, local selection | `figS13_mating_breadth` | Phenomenon known to island-model evolutionary computation; the contribution here is the mapping and the diversity/mean decomposition |
|
||
| Merge failure decomposes into a coordinate artefact plus a functional residual | Empirical at the trained-network and language-model tiers | Alignment enumerates the architecture's unit symmetries | Full-symmetry residual ≈ 0 for compatible parents versus ≈ the naive barrier under conflict; a cliff in hybrid fitness; function-specific breakdown at the LLM tier | Scoped to aligned linear interpolation; conflict floor is information-theoretic, not genetic |
|
||
| Epistasis (not divergence) sets the cliff; snowball onset | Biological-model result; **hypothesis** at the neural tier | BDM incompatibility structure | `fig5_speciation_bdm` | Snowball count ≠ performance cliff without the effect-size link; neural test outstanding |
|
||
| Pre-merge functional disagreement predicts merge penalty | Empirical, within a controlled grid (0.5B, 13 conditions × 3 seeds) | Constructed conflict/overlap/duration axes; oracle-potential outcome (pre-registered; ordering sensitive to reference) | Clustered CIs exclude 0; held-out LOCO ρ≈0.4; selected geometry baselines ≈ 0 | Head-to-head predictor differences not individually significant; only selected baselines; generalisation to real task pairs open |
|
||
| Confidence weighting improves rank prediction over raw disagreement | Not supported (pre-registered internal prediction) | — | Paired contrast over the same bootstrap resamples: Δ\|ρ\| = −0.021, CI [−0.130, +0.059] | The weighting does sharpen the conflict-versus-compatible level contrast, so it is not useless — only no better as a rank predictor |
|
||
| The predictor improves budget-matched operator choice | **Open** | — | Soup-vs-route gap readout noise-dominated at 0.5B | The practical payoff; untested |
|
||
| Emergent speciation without label conflict | Not observed (pre-registered) | Shared ancestry; compatible tasks; the divergences tested | Residual 0.000 to 6.4× base training; the merge rescues the specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
|
||
| Grounding, recombination, and diversity preservation make complementary contributions | Biological-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | `fig4_society_ablation` four-arm ablation; each arm fails in a distinct way | General joint necessity is not established; the language-model population (Fig. 4B–C) lacks differential reproduction between lineages |
|
||
| Obligate recombination collapses once partners carry conflicting conventions | Empirical (1.5B base, 3 lineages × 6 generations, 3 seeds) | Latin-square curriculum; replay present; linear merge; no culling of lineages | Best lineage 0.269 vs 0.796 never merging; onset at complementarity < 0.8; own-ancestor merge 0.663; three alternative mechanisms refuted (SI Text S3) | Six generations; one base; the arrival order of conflicting families is set by the curriculum |
|
||
| A declinable merge reverts the population to asexual accumulation without advance knowledge of when to stop | Empirical (same population, plus two controls, 3 seeds each) | "Keep the parent" scored as one candidate on validation data | Fraction declined 0.44 → 1.00 across generations; finishes 0.792 vs 0.796 never merging. Forced stop after generation 2 finishes 0.793 (veto − stop3 per seed −0.008/−0.006/+0.011). Under a decorrelated curriculum (complementarity 0.00 → 0.70 → 0.00) declines still rise 0.44 → 0.89; pooled partial ρ(declined, complementarity \| generation) = −0.07, CI (−0.21, +0.09); partial ρ with generation +0.31 | The reduction-principle reading (declines track complementarity) is **not supported**; declines track generation, which here confounds adapter age, skill count and the arrival of conflicting conventions. Modifier set by evaluation, not evolved |
|
||
| Recombination's net benefit across six generations is an early lead, not a final gain | Empirical (same population); consistent with the inheritance model's speed advantage | Every skill reaches every lineage by the curriculum regardless | +0.08 at generation 0; −0.005 at generation 5 (per-seed −0.03/+0.01/+0.01) | Replay present, so forgetting was not a live pressure; a curriculum that withholds skills from some lineages is untested |
|
||
|
||
## SI Table S2: headline quantitative results
|
||
|
||
Headline quantitative results with sample sizes, uncertainty, and outcome definitions (full
|
||
per-experiment tables and falsifier status in the per-experiment documentation).
|
||
|
||
| Result | Setting / n | Outcome definition | Headline |
|
||
|---|---|---|---|
|
||
| Closed-form validation | Inheritance model; standing tests | Simulated vs closed-form H-decay, immigration equilibrium, multi-parent union | Agreement < 0.5% |
|
||
| Grounding retention | Inheritance model (`fig2_grounding_sweep`); 100 lineages per grounding level | Fraction of equilibrium diversity retained at grounding `g` (operational threshold) | `g ≈ 0.05` retains ≥95% in the tested setting; smooth in `g` |
|
||
| MNIST collapse & rescue | Conv-VAE, 4 replicates; frozen oracle (98.5% mode acc.) | Mode support / forward-KL over generations | Dry: 30→1 modes; 10% grounding: 30/30 held |
|
||
| Fisher–Muller in LLMs | 5 seeds (0.5B) and 3 seeds (7B), fixed tests | Merged vs best-specialist accuracy (overall; worst family); ±: 95% CI over seeds | 0.5B ties 0.647±0.027 vs 0.592±0.009; 7B soup 0.873±0.004 vs 0.807±0.038 (soup − best +0.066±0.036, 3/3 seeds) |
|
||
| Union vs blend (headroom) | 3 seeds (0.5B hard); 3 seeds (7B hard) | Paired per-seed ordering, routing vs weight-average | 0.5B: routing > blend in 3/3 seeds, one catastrophic blend failure avoided. 7B: routing 0.503±0.007 vs soup 0.408±0.021 (+0.094±0.015, 3/3); soup vs best specialist +0.001±0.041 (the seed-1 'soup below best parent' did not replicate). Directed − soup +0.073±0.031 (3/3) |
|
||
| Speciation decomposition | MLPs, 3 replicates | LMC error barrier residual after permutation+rescaling alignment | Same-task 0.001; conflict 0.497 (naive 0.502) |
|
||
| Emergent isolation | MLPs 4 reps to 6.4× base training; LLM 1→12 epochs | Residual barrier; merged vs parent accuracy | 0.000 everywhere; merge rescues parents (≈0.955 vs ≈0.50) |
|
||
| LLM speciation, seeds | 0.5B; 3 training seeds; fixed test prompts | Conflict cliff: merge best-convention accuracy vs parents' own at full conflict. Duration null: merged private-task accuracy, 1 → 12 epochs | Cliff in 3/3 seeds (merge 0.02/0.12/0.16 vs parents 0.23–0.25); merged coherence over the sweep 0.147±0.013 → 0.100±0.082. No isolation in 3/3 (0.760±0.075 → 0.950±0.010) |
|
||
| Predictive test | 13 conditions × 3 seeds (0.5B) | Merge penalty vs oracle parent potential (pre-registered; ±: clustered 95% CI) | Functional ρ +0.45/+0.46, CI excl. 0; LOCO ρ ≈ 0.4; geometry n.s.; paired differences n.s. |
|
||
| Predictive test, seed sensitivity | Same; per-seed and leave-one-seed-out | Spearman ρ vs merge penalty within each seed alone (n = 13 conditions) | Functional +0.37 to +0.53 in every seed; weight geometry ≈ 0 in every seed; gradient alignment seed-unstable (−0.11 to −0.55) |
|
||
| Six-generation population | 1.5B base; 3 lineages × 6 generations; 3 training seeds; fixed tests (60 per family) | Best-lineage accuracy over six families at the final generation (mean of seeds; per-seed contrasts) | Never merge 0.796; declinable merge 0.792 (Δ −0.03/+0.01/+0.01); forced stop after generation 2: 0.793 (declinable − stop −0.008/−0.006/+0.011); obligate merge 0.269 (declinable − obligate +0.57/+0.54/+0.45); own-ancestor merge 0.663; single model 0.802 |
|
||
| Conflict-arrival curricula | Conflict-early / conflict-late (boolq + winogrande in generations 1–2 or 5–6); isolated, declinable and obligate arms; 3 seeds each | Decline rate and obligate-arm accuracy per generation; partial Spearman of declines on a conflict-present indicator controlling for generation (seed-clustered bootstrap) | Declines 0.56 → 0.78 (early), 0.44 → 0.89 (late); partial ρ(conflict | generation) −0.09, CI (−0.45, 0.15); ρ(generation | conflict) +0.45; pooled over four curricula −0.04, CI (−0.30, 0.15). Obligate final 0.280 / 0.386 vs isolated 0.796 / 0.781 (3/3 seeds); declinable 0.777 / 0.791 |
|
||
| Differential reproduction | Latin square; truncation selection (worst lineage re-founded from the best each generation); isolated and declinable arms; 3 seeds | Final best-lineage accuracy; per-seed contrasts against the unselected arms | Never merge + selection 0.804, declinable + selection 0.793 (declinable − never merge −0.011±0.003, below in 3/3); unselected 0.796 / 0.792; selection − no selection +0.007±0.030 |
|
||
| Second base lineage | SmolLM2-1.7B-Instruct; Fisher–Muller 5 seeds, headroom (hard) 3 seeds; fixed tests | Same contrasts as the Qwen rows | Soup − best specialist +0.049±0.022, TIES − best +0.097±0.020 (5/5 each; worst family +0.19 / +0.20). Routing − soup +0.162±0.036 (3/3); soup − best specialist −0.029±0.017 (below in 3/3) |
|
||
| Declinable-merge acceptance | Latin-square and decorrelated curricula; 9 merge decisions per generation; 3 seeds each | Fraction of proposed merges declined vs partner complementarity, with generation controlled (partial Spearman, seed-clustered bootstrap CI) | Latin square: 0.44 → 1.00 (raw ρ with complementarity −0.57, n = 18). Decorrelated curriculum (complementarity 0.00, 0.67, 0.70, 0.58, 0.33, 0.00): 0.44 → 0.89. Pooled (n = 36): partial ρ with complementarity −0.07, CI (−0.21, +0.09); partial ρ with generation +0.31 |
|
||
|
||
## SI Methods: experimental procedures
|
||
|
||
Every experiment in this paper is one YAML config, one runner invocation, and one artifact triple
|
||
(`results.parquet` + the resolved config + a manifest carrying the master seed, git commit, library
|
||
versions, and a content hash). The configs named below are the authority on any parameter; this
|
||
section gives the scientific reasoning behind the choices. `REPRODUCING.md` maps each manuscript
|
||
panel to the config and seed that produced it.
|
||
|
||
### M1. Design principles
|
||
|
||
Four rules govern every choice that follows.
|
||
|
||
*Test each claim at the cheapest tier that can falsify it.* A closed form beats a simulation, a
|
||
simulation beats a trained network, and a small network beats a language model, whenever the cheaper
|
||
instrument can still return the answer "no". A costlier tier is entered only where it adds a
|
||
discriminating test rather than a replication — which is why several cells of the programme (Fig. 1A)
|
||
are deliberately empty.
|
||
|
||
*Match the precision of the claim to the precision of the instrument.* The inheritance model is exact,
|
||
so it carries the paper's quantitative statements. Trained systems add optimisation noise and
|
||
inductive bias, so at those tiers I claim signs and orderings, never magnitudes.
|
||
|
||
*Make reality able to refuse.* Every tier has an oracle that is independent of the model being
|
||
measured: a fixed true distribution in the inheritance model, a lossless identity code or a frozen
|
||
classifier in the neural tier, an exact-match verifier over procedurally generated tasks in the
|
||
language-model tier.
|
||
|
||
*Declare the falsifier before running.* Each experiment states the outcome that would refute the
|
||
claim it tests (per-experiment READMEs; SI Table S1). Two pre-registered predictions failed, and are
|
||
reported as failures in the main text.
|
||
|
||
### M2. Replication: what a replicate is, and how many
|
||
|
||
A replicate means something different at each tier, and conflating the three would misstate what the
|
||
error bars cover.
|
||
|
||
In the inheritance model a replicate is an independent lineage: a fresh random stream driving the same
|
||
resolved config, with sub-seeds derived from the master seed by `SeedSequence.spawn`. Because drift
|
||
*is* the object of study, the spread across replicates is signal rather than nuisance, and replicate
|
||
counts are set so that the confidence interval on the summary statistic is small relative to the
|
||
effect being reported.
|
||
|
||
In the trained-network tier a replicate is an independent lineage including fresh weight
|
||
initialisation and data ordering, so it carries optimisation noise on top of drift.
|
||
|
||
In the language-model tier a replicate is an independent *training* seed evaluated on *fixed* test
|
||
sets. Holding the evaluation data constant while varying the training seed isolates training
|
||
stochasticity, which is the quantity in doubt; it also means the seed-to-seed spread I report is not
|
||
inflated by resampling the benchmark.
|
||
|
||
Replicate counts, and why each is what it is:
|
||
|
||
| Experiment | Replicates | Reasoning |
|
||
|---|---|---|
|
||
| `fig2_grounding_sweep`, `figS5_aimed_grounding`, `figS12_quality_diversity`, `figS3_rebaselining` | 100 lineages | Long horizons (400–600 generations) with drift-dominated variance; 100 lineages put the CI on stationary diversity well inside the effect being resolved |
|
||
| `figS8_multiparent_union` | 200 | Outcomes are per-item binary retentions, the highest-variance quantity in the paper |
|
||
| `figS9_specialist_superparent` | 40 | The vertical claim; the headline separation, so the most replicated of the genotype experiments |
|
||
| `figS10_rugged_landscapes`, `figS11_directed_recombination` | 24 | Landscape sweeps where each point aggregates 200 offspring internally |
|
||
| `fig4_society_ablation` | 12 | Four-arm ablation over 80 generations; arms separate by margins far exceeding the CI |
|
||
| `fig5_speciation_bdm` | 15 | Each point already averages 500 offspring |
|
||
| `figS13_mating_breadth` | 20 | Breadth × ruggedness grid, 60 generations per cell |
|
||
| `figS2_kernel_sharpen`, `figS2_kernel_smooth` | 24 | Two-parameter kernel fits against neural reference endpoints |
|
||
| `bridge` | 60 | The harness gate: must detect *any* departure from the inheritance model, so the most replicated neural run |
|
||
| `figS6_grounding_rnn` | 18 | Nine-point grounding sweep with per-generation network retraining |
|
||
| `collapse`, `figS1_architectures` | 5 | Sign-level demonstrations across architectures; each lineage retrains a network 22–25 times |
|
||
| recombination | 8 | Operator contrast in trained weights |
|
||
| `fig2_mnist_collapse` | 4 | 15 generations × a conv-VAE retrained from scratch each generation; the contrast (30 modes vs 1) is categorical |
|
||
| speciation_real, _cliff | 3 | Barrier decomposition; the quantity is a near-deterministic function of the training condition (residual 0.001 vs 0.497) |
|
||
| speciation_real_emergent | 4 | A null: replicates are spent on longer divergence horizons rather than more repeats |
|
||
| llm_merge_seeds | 5 training seeds | The Fisher–Muller signature, the most-replicated language-model claim |
|
||
| llm_moe_hard_seeds, llm_directed_hard_seeds, llm_epistasis(+compat), llm_speciation_add | 3 training seeds | Per-seed orderings reported individually rather than averaged |
|
||
| 7B runs (llm_merge_hpc, llm_moe_hard_hpc, llm_directed_hard_hpc) | 3 training seeds | Seeds 2–3 added 2026-09-11 (`hpc/llm_7b_seeds.pbs`, ~33 min per seed on one L40S); per-seed contrasts in `figures/stats_llm_7b_seeds.py` |
|
||
| llm_curriculum_v5_{early,late}(_obl) | 3 training seeds each | Conflict-arrival curricula; per-seed contrasts and the pooled partial-correlation test |
|
||
| llm_curriculum_v5_cull | 3 training seeds | Differential reproduction; per-seed contrasts against the unselected arms |
|
||
| llm_merge_seeds_smol, llm_moe_hard_seeds_smol | 5 and 3 training seeds | Second base lineage; per-seed orderings as for the Qwen runs |
|
||
| llm_speciation | 3 training seeds | Conflict cliff and duration null checked seed by seed (`figures/stats_llm_speciation_seeds.py`); seeds 2–3 added 2026-09-12 |
|
||
| llm_curriculum_v5, llm_curriculum_v5_veto, llm_curriculum_v5_stop3, llm_curriculum_v5_decor | 3 training seeds | The six-generation population; arm separations (≈0.5) far exceed seed spread (≈0.02), and the declinable-vs-never contrast is reported per seed because its mean is near zero |
|
||
|
||
The asymmetry is deliberate: replicates are cheap exactly where the quantitative claims live, and the
|
||
expensive tiers are asked only for the sign of an effect the cheap tier has already quantified. Where
|
||
a single run is all there is, the manuscript says so.
|
||
|
||
### M3. The inheritance-model tier
|
||
|
||
Knowledge is a distribution over `K` discrete items; reality is a fixed Zipf-tailed distribution
|
||
`p*`; one generation resamples `n` draws from the parent, optionally mixes in `m` verified draws from
|
||
`p*`, and refits. Implementation: NumPy/SciPy, no GPU, bitwise reproducible.
|
||
|
||
*Parameter choices.* `K = 500`–`1000` with `zipf_s = 1.1` and half the items designated tail: large
|
||
enough that the rare tail contains hundreds of items (so tail statistics are not dominated by a
|
||
handful of them) and small enough to sweep densely. `n = 100`–`200` sets drift strength; it is the
|
||
population size in the Wright–Fisher correspondence and the distillation sample size in the AI
|
||
reading. Horizons of 400–600 generations were chosen so that ungrounded lineages reach fixation and
|
||
grounded ones reach stationarity within the run, which the trajectories confirm.
|
||
|
||
*Sweeps.* The grounding sweep (`fig2_grounding_sweep`) sweeps `g ∈ {0, 0.005, 0.01, 0.02, 0.05, 0.1, 0.2, 0.4}`;
|
||
the aimed-grounding experiment (`figS5_aimed_grounding`) contrasts uniform against region-matched
|
||
grounding allocation; the multi-parent union experiment (`figS8_multiparent_union`) crosses parent count
|
||
`K_T ∈ {1,2,3,5}` with parent correlation `ρ ∈ {0, 0.25, 0.5, 0.75, 1}` and `g ∈ {0, 0.02, 0.05}`; the
|
||
selection experiment (`figS12_quality_diversity`) crosses selection mode (none / greedy /
|
||
quality-diversity) with novelty weight; the re-baselining experiment (`figS3_rebaselining`) compares
|
||
four re-minting arms.
|
||
|
||
*The correlated-parent construction (`figS8_multiparent_union`).* Parent correlation is constructed directly rather than
|
||
obtained by tuning drift, so that `ρ` is not confounded with `n`, `m`, tail size, or generation
|
||
count. For each tail item a shared switch `z ~ Bern(ρ)`, a shared retention `s ~ Bern(q)`, and
|
||
per-parent `u⁽ᵏ⁾ ~ Bern(q)` give parent `k` retention `s` if `z` else `u⁽ᵏ⁾`. This yields exact
|
||
marginal retention `q` and exact pairwise correlation `ρ`, and is exchangeable, so `ρ` is a single
|
||
scalar knob.
|
||
|
||
*Multi-locus experiments* (`figS9_specialist_superparent`, `figS10_rugged_landscapes`,
|
||
`figS11_directed_recombination`, `fig4_society_ablation`, `figS13_mating_breadth`). Genotypes are `L = 12` biallelic loci (4096 genotypes —
|
||
effectively open-ended relative to the population sizes used), with fitness either additive or a
|
||
Kauffman NK landscape whose interaction count `K` tunes ruggedness from 0 to 10. The landscape and directed-recombination experiments breed
|
||
from `n_parents = 6` local optima into populations of 200 offspring; the directed one additionally
|
||
screens offspring and iterates (5 rounds, keeping 8). The society ablation runs a population of `N = 60` agents for 80
|
||
generations at ruggedness `K = 8`, with mutation `μ = 0.03`, 120 offspring per generation, and
|
||
selection weighting true fitness against consensus conformity at `g = 0.85`. The mating-breadth experiment sweeps mate-pool
|
||
breadth on a ring of `N = 48` against ruggedness.
|
||
|
||
*Speciation (`fig5_speciation_bdm`).* `L = 20` loci, incompatibility density `ρ ∈ {0.1, 0.25, 0.5}`,
|
||
parental divergence swept 0–20 substitutions, 500 offspring per cell at recombination rate 0.5.
|
||
|
||
*Validation.* Three closed forms are asserted as standing tests to within 0.5%: neutral
|
||
heterozygosity decay `E[H_t] = H_0(1 − 1/n)^t`, the exact immigration–drift equilibrium, and the
|
||
multi-parent union formula. These run in CI alongside the correctness tests. If they fail, the
|
||
science is wrong rather than merely the code.
|
||
|
||
### M4. The trained-network tier
|
||
|
||
*Why a synthetic universe.* Measuring collapse requires knowing the true distribution exactly. Each
|
||
mode is rendered as a token sequence carrying an identity segment (base-2 digits encoding the mode
|
||
index losslessly, so the oracle reads the mode back with zero error) followed by style tokens drawn
|
||
uniformly at random. The style segment gives genuine within-mode entropy, so a generative model must
|
||
learn a distribution rather than memorise `K` fixed strings, while the identity segment keeps the
|
||
measurement noise-free. Mode truth comes from the same `make_true_distribution` used by the
|
||
inheritance model, so "mode", "region", and "tail" denote the same objects at both tiers.
|
||
|
||
*The bridge gate.* Before any trained model is interpreted, a histogram generator is run through the
|
||
identical harness; it must reproduce the inheritance model exactly. This separates harness bugs from
|
||
model behaviour, and is why the bridge run carries 60 replicates.
|
||
|
||
*Architectures and training.* The recurrent generator is an embedding (26) → GRU (128 hidden; 192 in
|
||
the architecture-generality run) → linear readout, trained each generation from scratch with Adam,
|
||
learning rate 2×10⁻³, batch size 256, 25 epochs, and evaluated by sampling 12,000–15,000 sequences.
|
||
Feedforward and variational autoencoder generators share the harness. Retraining from scratch each
|
||
generation (rather than fine-tuning) makes the generational step a clean refit, matching the
|
||
inheritance model's operator.
|
||
|
||
*MNIST tier.* Dataset: MNIST via torchvision (60,000 training images). Modes are digit class ×
|
||
stroke-thickness bin (10 × 3 = 30 modes) with a Zipf frequency profile, so roughly eighteen modes are
|
||
rare. The generator is a convolutional variational autoencoder (latent 32, β = 1), retrained from
|
||
scratch each generation with Adam, learning rate 10⁻³, batch 256, 30 epochs, on 6,000 images drawn
|
||
from the previous generation's own samples, for 15 generations, at `g ∈ {0, 0.1}`. The oracle is a
|
||
frozen two-convolution classifier trained once (5 epochs) combined with a deterministic thickness
|
||
measure; it reaches 98.5% mode accuracy and its 30 × 30 confusion matrix is recorded in the manifest
|
||
as the measurement floor. Build gates: the oracle's accuracy, and generation-0 recovery of all 30
|
||
modes.
|
||
|
||
*Speciation in trained weights.* Two multilayer perceptrons (784–512–512–10, ReLU, no batch
|
||
normalisation — batch statistics would break the permutation correspondence the analysis depends on)
|
||
are forked from a shared base trained for 500 steps, then trained apart for 100–3,200 further steps
|
||
(up to 6.4× the shared base) under SGD at learning rate 0.05, batch 128. Merges are weight averages;
|
||
the readout is the linear-mode-connectivity error barrier before and after alignment. Alignment
|
||
composes deterministic Re-Basin permutation matching with exact per-unit scale canonicalisation —
|
||
the unit symmetry group of this architecture — and is gated by a control that must recover a
|
||
permuted-and-rescaled copy exactly. Since the search space is that group rather than all
|
||
possible alignments, the removable share is a lower bound and the residual an upper bound.
|
||
|
||
### M5. The language-model tier
|
||
|
||
*Base models.* Qwen2.5-Instruct at 0.5B and 7B, open weights under a permissive licence, with the
|
||
revision pinned. Using two sizes from one family makes scale the only variable that changes between
|
||
the small and large runs; the 0.5B model carries five-seed protocols on the easy families; the 7B
|
||
runs are replicated over three training seeds.
|
||
|
||
*Task families, and why they are procedural.* Three deliberately disjoint families — list
|
||
operations, string transformations, and small-integer arithmetic — are generated procedurally from a
|
||
seed. Procedural generation buys four things that a standard benchmark cannot: an exact-match
|
||
verifier that plays the role of reality (an answer is right or it is not, with no judge model in the
|
||
loop); freedom from train/test contamination, since every evaluation item is generated fresh from a
|
||
disjoint seed offset; control over family disjointness, which is the precondition for specialists to
|
||
be genuinely decorrelated parents; and a difficulty knob. A `hard` variant (multi-step list
|
||
operations, Caesar ciphers and letter transforms, multi-step and larger arithmetic) exists because
|
||
the easy families saturate a 7B base at ceiling, and saturation removes the headroom in which
|
||
recombination operators can differ — a control that proved necessary, since two null results at 7B
|
||
turned out to be saturation artefacts rather than scale effects.
|
||
|
||
*Data splits.* Training, validation, routing-calibration, and test items are drawn from
|
||
non-overlapping seed offsets by construction (test from 1000 + family index, routing from 2000 +,
|
||
validation from 3000 +, training from the run seed). Test sets are fixed across seeds in the
|
||
multi-seed protocols. Selection of merge weights uses validation only; the winners are then reported
|
||
on the untouched test split.
|
||
|
||
*Specialisation.* Each parent is a LoRA adapter (rank 16, α = 32) on the frozen base, applied to all
|
||
attention and MLP projection matrices, trained with a manual supervised fine-tuning loop: answer-only
|
||
cross-entropy (prompt tokens masked out of the loss), AdamW at 2×10⁻⁴, batch size 8, 3 epochs,
|
||
bfloat16, 400–800 training items per family. Low-rank adaptation is the right instrument here for a
|
||
structural reason rather than a computational one: it confines each parent's specialisation to an
|
||
additive low-rank delta over an identical frozen base, which is what makes weight-space recombination
|
||
between parents well defined.
|
||
|
||
*Recombination operators.* Fusion by uniform weight averaging (soup) and by sign-reconciled,
|
||
magnitude-pruned task arithmetic (TIES); union by keeping specialists intact and selecting per input
|
||
(oracle routing, and a training-free nearest-centroid router over the base model's own prompt
|
||
embeddings) or per module (winner-take-all by delta norm); and directed recombination, which breeds a
|
||
population of Dirichlet-weighted merges, scores each on validation, and keeps the fittest.
|
||
|
||
*Evaluation.* Greedy decoding, exact match after canonicalisation. Alongside overall accuracy I
|
||
report worst-family accuracy, because the Fisher–Muller claim is about competence across all
|
||
families rather than an average that a single strong specialty can carry.
|
||
|
||
*The controlled predictive test.* Thirty-nine parent pairs (13 conditions × 3 seeds) span three axes
|
||
that are decorrelated by construction: conflict (contradictory conventions on shared prompts, with
|
||
private training budgets held fixed), compatible overlap (the same shared prompts under the same
|
||
convention — overlap and volume without conflict), and duration (weight divergence with no conflict,
|
||
1 to 12 epochs). Six predictors are computed before any merge: confidence-weighted functional
|
||
conflict, raw disagreement, gradient alignment at the shared base, LoRA-delta cosine and L2 distance,
|
||
and a cross-task performance baseline. Probes are drawn blind to where the conflict lives. The
|
||
outcome is the merge penalty against oracle parent potential, pre-registered, and also reported
|
||
against best-parent and mean-parent references because the predictor ordering is sensitive to that
|
||
choice.
|
||
|
||
*The six-generation population.* Base model Qwen2.5-1.5B (base weights, not the instruction-tuned
|
||
variant; 0.006 accuracy on the families untrained). Six public datasets with per-family verifiers:
|
||
natural-language inference (MNLI; label), science questions (ARC; letter), commonsense completion
|
||
(HellaSwag; letter), reading-comprehension spans (SQuAD; normalised span with aliases), yes/no
|
||
questions (BoolQ), and pronoun resolution (WinoGrande; 1/2). Each family's pool is split into disjoint
|
||
training, validation, and test items before any sampling, so validation and test never share an item.
|
||
Two further curricula, conflict-early and conflict-late, are given as explicit orders: the two families
|
||
whose answer conventions conflict (BoolQ yes/no, WinoGrande 1/2) occupy generations 1–2 or 5–6 of
|
||
every lineage and the four compatible families fill the remaining generations in rotated orders, so
|
||
adapter age and skill count rise one family per generation in both and only the arrival of conflict
|
||
differs (`configs/llm/curriculum_v5_{early,late}.yaml`; obligate arms in the `_obl` configs; three
|
||
training seeds each; `hpc/llm_curriculum_timing.pbs`). The conflict-timing readout is the partial
|
||
Spearman correlation of the per-generation decline rate with an indicator of conflict presence,
|
||
controlling for generation, with a seed-clustered percentile bootstrap (`figures/stats_llm_curriculum.py`).
|
||
Differential reproduction (`cull: true`) applies truncation selection after each generation's
|
||
measurement: the lineage with the lowest all-families accuracy is re-founded from the one with the
|
||
highest (adapter, taught families, example budget and ancestry archive are copied; the slot keeps its
|
||
curriculum order; ties leave the population unchanged), recorded as `culled` and `cull_source` rows
|
||
(`configs/llm/curriculum_v5_cull.yaml`; three training seeds; `hpc/llm_cull.pbs`).
|
||
The second base lineage is `HuggingFaceTB/SmolLM2-1.7B-Instruct` (Apache-2.0; Llama architecture),
|
||
run through the unchanged `merge_seeds` and `moe_hard_seeds` protocols with its own adapter cache
|
||
(`configs/llm/{merge_seeds,moe_hard_seeds}_smol.yaml`; `hpc/llm_smol.pbs`; `figures/stats_llm_smol.py`).
|
||
Three lineages take the six families in a cyclic Latin square (each lineage's order is the previous
|
||
lineage's shifted by two), which fixes partner complementarity — the fraction of the partner's families
|
||
a lineage has not yet seen — at 1.0, 1.0, 0.8, 0.67, 0.33, 0.0 across the six generations. Each
|
||
generation a lineage draws 300 new items from its scheduled family and 150 replay items split evenly
|
||
across families already seen; the child adapter (rank 16) is initialised from the parent's and trained
|
||
for 3 epochs at learning rate 10⁻⁴ (founders from the base at 2×10⁻⁴). Recombination averages two
|
||
adapters at each weight in {0.5/0.5, 0.3/0.7, 0.7/0.3}; the winner is chosen on 20 validation items
|
||
per family seen and then trains on the generation's new family. In the declinable arm the unchanged
|
||
parent is a fourth candidate scored identically. The contemporary partner is the next lineage in the
|
||
square; the ancestor partner is the lineage's own adapter three generations earlier; the self-replay
|
||
variant draws its replay from the parent's own answers rather than from the datasets. Reporting uses
|
||
60 test items per family. Because two of the families are binary, an accuracy threshold at 0.5 is
|
||
chance, so the text reports mean accuracy over the families a lineage has been taught and the
|
||
trajectory of its first-learned family rather than a count of families above a threshold.
|
||
|
||
*The composed society.* A population of `N` LoRA agents on a shared frozen base evolves for `G`
|
||
non-overlapping generations. Each generation every agent answers a fixed validation pool (verifier
|
||
scored) and a fresh conformity pool (whose modal answer defines the population consensus); selection
|
||
scores agents by `g·fitness + (1−g)·conformity`; parents are chosen with or without a
|
||
quality-diversity term over behavioural distance; offspring are bred by screened recombination; and
|
||
each child is a fresh adapter distilled from its source model's own answers, which makes the
|
||
inheritance channel literally self-consuming. The verifier enters the loop only where `g > 0`, but is
|
||
used for reporting in every arm. The four-arm ablation removes grounded evaluation, recombination,
|
||
or diversity preservation in turn.
|
||
|
||
### M6. Negative controls
|
||
|
||
The design leans on controls that can remove a result rather than support one, and one of them did.
|
||
|
||
The compatible-overlap axis was added to the predictive test specifically to expose overlap-and-volume
|
||
artefacts, and it did: the initial two-axis grid's best predictor (LoRA-delta cosine, ρ = +0.60)
|
||
collapsed to ρ = +0.03 once compatible overlap was present, identifying it as an artefact rather than
|
||
a signal. The duration axis supplies divergence without conflict. The emergent-speciation condition
|
||
supplies divergence with no conflicting signal anywhere, and returns a null. The budget-controlled
|
||
speciation design (`conflict_mode: add`) removes the confound between conflict fraction and private
|
||
training budget. The histogram bridge is a harness control. In the inheritance model, `m = 0` arms and
|
||
`ρ = 1` (fully correlated parents) are the null conditions against which the corresponding effects
|
||
are read. In the six-generation population, three arms are controls — a single model taught the
|
||
curriculum alone (no population), the never-merge population (no recombination), and merging with
|
||
one's own ancestor (shared conventions, partial complementarity) — and the self-replay variant of the
|
||
obligate arm controls the replay channel; SI Text S3 records the three direct tests that refuted
|
||
alternative mechanisms for the obligate arm's collapse.
|
||
|
||
### M7. Statistical procedures
|
||
|
||
Error bars on replicate means are normal-approximation 95% confidence intervals unless stated
|
||
otherwise. For the predictive test, where rows share task-data seeds across conditions and are
|
||
therefore not independent, inference uses a condition-clustered bootstrap (13 clusters, 4,000
|
||
resamples); predictors are compared by paired contrasts on the same resamples; generalisation is
|
||
assessed by leave-one-condition-out prediction; and per-seed and leave-one-seed-out sensitivity are
|
||
reported alongside, since three seeds cannot settle seed generalisation on their own. Outcome-
|
||
reference sensitivity is reported rather than resolved. Where a difference is not significant at the
|
||
sample size available, the manuscript says so rather than reporting the point estimate alone.
|
||
|
||
## SI Statistics
|
||
|
||
Output of `figures/stats_llm_epistasis.py` (clustered CIs, paired predictor contrasts, LOCO held-out
|
||
prediction, outcome-reference sensitivity, within/between-axis decomposition) — reproduced verbatim at
|
||
submission. Chronology of the predictive test (prospective / adaptive / post-hoc) as disclosed in
|
||
`results/llm_epistasis/README.md`.
|
||
|
||
## SI Figures
|
||
|
||
Sixteen figures are cited from the main text by number. Each is the per-experiment figure
|
||
regenerated from the committed results artifact (`figures/plot_*.py`), reproduced here without
|
||
re-plotting, so panel titles still carry the experiment's working name. Five of them are
|
||
inheritance-model results with no real-model counterpart in this paper, reported here because each
|
||
reproduces an established result: blending versus union retention (Fig. S8), the Fisher–Muller
|
||
super-parent (Fig. S9), outbreeding depression on rugged landscapes (Fig. S10), directed
|
||
recombination (Fig. S11), and the mate-pool breadth optimum (Fig. S13).
|
||
|
||
*(FIG:s1)*
|
||
|
||
*(FIG:s2)*
|
||
|
||
*(FIG:s3)*
|
||
|
||
*(FIG:s4)*
|
||
|
||
*(FIG:s5)*
|
||
|
||
*(FIG:s6)*
|
||
|
||
*(FIG:s7)*
|
||
|
||
*(FIG:s8)*
|
||
|
||
*(FIG:s9)*
|
||
|
||
*(FIG:s10)*
|
||
|
||
*(FIG:s11)*
|
||
|
||
*(FIG:s12)*
|
||
|
||
*(FIG:s13)*
|
||
|
||
*(FIG:s14)*
|
||
|
||
*(FIG:s15)*
|
||
|
||
*(FIG:s16)*
|