SI Methods: a full experimental-procedures appendix
Replaces the three-paragraph methods sketch with a scientific account of how
the study was run (M1-M7):
- M1 design principles: cheapest falsifying tier; match claim precision to
instrument precision; every tier gets an oracle independent of the model
being measured; falsifiers declared before running.
- M2 replication: what a replicate *is* differs by tier (independent lineage /
lineage incl. fresh init and data order / training seed with test sets held
fixed), and a table giving every experiment's replicate count with the
reasoning - why 200 for E4 (per-item binary outcomes), 60 for the bridge
gate (must detect any departure), 3-5 where the contrast is categorical,
and 1 for the 7B runs, labelled as single runs.
- M3-M5 per-tier procedures: parameter choices and their justification, the
correlated-parent construction, why the neural sandbox is synthetic (a
lossless identity code plus style entropy gives an exact oracle while still
forcing the model to learn a distribution), MNIST modes and the frozen-CNN
oracle with its confusion matrix as measurement floor, why no-BatchNorm MLPs
for the alignment analysis, and for the LLM tier: why Qwen 0.5B/7B (one
family so scale is the only variable), why procedural tasks rather than a
benchmark (exact verifier, contamination-free, controlled disjointness, a
difficulty knob), why LoRA (confines each parent to an additive low-rank
delta over an identical base, which is what makes weight-space
recombination well defined), the training algorithm, and the split scheme.
- M6 negative controls, including the one that removed a result: the
compatible-overlap axis collapsed the delta-cosine predictor from rho=+0.60
to +0.03.
- M7 statistical procedures.
Also: SI voice converted to first person and terminology synced to the
"biological model" rename; removed a process ghost from the preamble
("Skeleton assembled at Phase 4"); build.py now takes a document argument and
no longer eats documents that lack a title block, so the SI compiles via a new
si.tex wrapper (10 pp). `make paper` builds both PDFs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
parent
c435cfba6e
commit
a88289964a
6 changed files with 493 additions and 40 deletions
294
paper/pnas/si.md
294
paper/pnas/si.md
|
|
@ -1,9 +1,10 @@
|
|||
# SI Appendix — The evolution of sex for artificial intelligence
|
||||
|
||||
*Skeleton assembled at Phase 4; finalised at submission. Every numbered experiment has a committed
|
||||
config (`configs/`), artifact triple (`results/<name>/results.parquet` + resolved config + manifest
|
||||
with content hashes and git commit), a README with its legend and falsifier status, and a figure that
|
||||
regenerates from the parquet alone. `reproduce.sh` re-runs everything from the master seeds.*
|
||||
*Every experiment has a committed config (`configs/`), an artifact triple
|
||||
(`results/<name>/results.parquet` + the resolved config + a manifest carrying content hashes, master
|
||||
seed, and git commit), a README with its legend and falsifier status, and a figure that regenerates
|
||||
from the parquet alone. `reproduce.sh` re-runs the whole study from the master seeds; `REPRODUCING.md`
|
||||
maps each manuscript panel to the config and seed behind it.*
|
||||
|
||||
## SI Text S1–S2: formal statements
|
||||
|
||||
|
|
@ -23,7 +24,7 @@ interpolation path. For every function-preserving `T`, the endpoint functions, h
|
|||
losses and this chord, are identical for `(A, T(B))` and `(A, B)`. The **interpolation path itself is
|
||||
generally NOT invariant** — losses along `(1−α)·A + α·T(B)` change with `T`, which is precisely why
|
||||
alignment can lower a barrier. *(Immediate from the definition of function-preserving.)* Scope
|
||||
caveat: our aligner provably recovers a permuted-and-rescaled copy exactly — an important special
|
||||
caveat: the aligner provably recovers a permuted-and-rescaled copy exactly — an important special
|
||||
case — but this does not establish global optimality of the alignment over the symmetry group for
|
||||
independently trained networks; the decomposition's "removable" share is therefore a lower bound, and
|
||||
the "residual" an upper bound, on their true values.
|
||||
|
|
@ -53,7 +54,7 @@ group can remove for *compatible* models, the sharper the meaning of the residua
|
|||
*incompatible* ones — and Proposition 2 caps what any of them could ever achieve on the conflict set.
|
||||
|
||||
**Terminology note for the paper.** "Residual (after alignment)" = the estimated functional
|
||||
incompatibility; for ReLU MLPs we align modulo the full unit symmetry group, so the estimate is not
|
||||
incompatibility; for ReLU MLPs I align modulo the full unit symmetry group, so the estimate is not
|
||||
confounded by missed symmetries of that architecture class.
|
||||
|
||||
## S2. Emergent vs imposed incompatibility (E13b framing)
|
||||
|
|
@ -82,19 +83,19 @@ long-horizon over-specialisation erodes mergeability at LLM scale (cf. arXiv:260
|
|||
|
||||
| Claim | Status | Key assumptions | Evidence | Known limits |
|
||||
|---|---|---|---|---|
|
||||
| Collapse = Wright–Fisher drift (minimal model) | Exact (diagnosis conceded to prior work) | Knowledge = categorical distribution; refit = resample | Closed forms reproduced to <0.5% | Real learners add a signed, architecture-specific estimator bias (measured) |
|
||||
| Collapse = Wright–Fisher drift (biological model) | Closed form (diagnosis conceded to prior work) | Knowledge = categorical distribution; refit = resample | Closed forms reproduced to <0.5% | Real learners add a signed, architecture-specific estimator bias (measured) |
|
||||
| Grounding = immigration; critical real-data fraction ≪ 1 | Exact + empirical sign | Fresh samples from a fixed, non-drifting truth | Exact `H_eq`; `g*≈0.048`; sign holds in RNN/MLP/VAE and on MNIST | Deepest tail unrescuable at feasible budgets (`m ∼ 1/p`); sharp threshold softens in trained nets |
|
||||
| "Merge, don't average" conservation | Exact **for the output-mean operator** | Rare-item regime; an oracle/verifier identifies the strongest source | E4 closed form + simulation; neural reproduction | Weight-averaging and routing are empirical cousins, not instances; budgets differ; bridge = the headroom rule |
|
||||
| Offspring exceed every parent (Fisher–Muller) | Interpretation + empirical | Complementary (decorrelated) parents; verifiable fitness | E8 analytic; 7B LoRA merge beats every specialist on every family | LLM tier: 3 lexically-distinct families; multi-seed replication in progress |
|
||||
| Outbreeding depression on rugged landscapes; operator design rule | Exact-model result; hypothesis at LLM scale | NK epistasis stands in for skill entanglement | E9–E10; directed selection rescues | Not yet mapped onto a real task-entanglement measure |
|
||||
| Optimal mate-pool breadth shrinks with ruggedness | Exact-model result; hypothesis for merging populations | Ring population, local selection | E14 | Phenomenon known to island-model evolutionary computation; our contribution is the mapping and the diversity/mean decomposition |
|
||||
| Offspring exceed every parent (Fisher–Muller) | Interpretation + empirical | Complementary (decorrelated) parents; verifiable fitness | E8 (biological model); 7B LoRA merge beats every specialist on every family | LLM tier: 3 lexically-distinct families; replicated over five training seeds at 0.5B |
|
||||
| Outbreeding depression on rugged landscapes; operator design rule | Biological-model result; hypothesis at LLM scale | NK epistasis stands in for skill entanglement | E9–E10; directed selection rescues | Not yet mapped onto a real task-entanglement measure |
|
||||
| Optimal mate-pool breadth shrinks with ruggedness | Biological-model result; hypothesis for merging populations | Ring population, local selection | E14 | Phenomenon known to island-model evolutionary computation; the contribution here is the mapping and the diversity/mean decomposition |
|
||||
| Merge failure decomposes into coordinate artefact + functional residual | Empirical (MLP tier; LLM tier in progress) | Alignment enumerates the architecture's unit symmetries | Full-symmetry residual ≈ 0 (compatible) vs ≈ naive (conflict); cliff in hybrid fitness | Scoped to aligned linear interpolation; conflict floor is information-theoretic, not genetic |
|
||||
| Epistasis (not divergence) sets the cliff; snowball onset | Exact-model result; **hypothesis** at the neural tier | BDM incompatibility structure | E12 | Snowball count ≠ performance cliff without the effect-size link; neural test outstanding |
|
||||
| Epistasis (not divergence) sets the cliff; snowball onset | Biological-model result; **hypothesis** at the neural tier | BDM incompatibility structure | E12 | Snowball count ≠ performance cliff without the effect-size link; neural test outstanding |
|
||||
| Pre-merge functional disagreement predicts merge penalty | Empirical, within a controlled grid (0.5B, 13 conditions × 3 seeds) | Constructed conflict/overlap/duration axes; oracle-potential outcome (pre-registered; ordering sensitive to reference) | Clustered CIs exclude 0; held-out LOCO ρ≈0.4; selected geometry baselines ≈ 0 | Head-to-head predictor differences not individually significant; only selected baselines; generalisation to real task pairs open |
|
||||
| Confidence weighting improves rank prediction over raw disagreement | **Not supported** (pre-registered internal prediction) | — | Paired Δ\|ρ\| ≈ −0.02, CI [−0.13, +0.06] | Weighting does double the conflict-vs-compat level contrast |
|
||||
| The predictor improves budget-matched operator choice | **Open** | — | Soup-vs-route gap readout noise-dominated at 0.5B | The practical payoff; untested |
|
||||
| Emergent speciation without conflict | **Not observed** (pre-registered) | Shared ancestry, compatible tasks, tested divergences | E13b: residual 0.000; merge rescues specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
|
||||
| Grounding + sex + diversity jointly necessary | Exact-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |
|
||||
| Grounding + sex + diversity jointly necessary | Biological-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |
|
||||
|
||||
## SI Table S2: headline quantitative results
|
||||
|
||||
|
|
@ -103,7 +104,7 @@ per-experiment tables and falsifier status in the per-experiment documentation).
|
|||
|
||||
| Result | Setting / n | Outcome definition | Headline |
|
||||
|---|---|---|---|
|
||||
| Closed-form validation | Analytic tier; standing tests | Simulated vs closed-form H-decay, immigration equilibrium, multi-teacher union | Agreement < 0.5% |
|
||||
| Closed-form validation | Biological model; standing tests | Simulated vs closed-form H-decay, immigration equilibrium, multi-teacher union | Agreement < 0.5% |
|
||||
| Grounding retention | Minimal model; 18+ replicates per point | Fraction of equilibrium diversity retained at grounding g (operational threshold) | g ≈ 0.05 retained ≥95% (tested setting); smooth in g |
|
||||
| MNIST collapse & rescue | Conv-VAE, 4 replicates; frozen oracle (98.5% mode acc.) | Mode support / forward-KL over generations | Dry: 30→1 modes; 10% grounding: 30/30 held |
|
||||
| Fisher–Muller in LLMs | 5 seeds (0.5B), fixed tests; single 7B run | Merged vs best-specialist accuracy (overall; worst family) | Ties 0.647±0.027 vs 0.592±0.009; 7B 0.87 vs 0.77 |
|
||||
|
|
@ -112,29 +113,256 @@ per-experiment tables and falsifier status in the per-experiment documentation).
|
|||
| Emergent isolation | MLPs 4 reps to 6.4× base training; LLM 1→12 epochs | Residual barrier; merged vs parent accuracy | 0.000 everywhere; merge rescues parents (≈0.955 vs ≈0.50) |
|
||||
| Predictive test | 13 conditions × 3 seeds (0.5B) | Merge penalty vs oracle parent potential (pre-registered; ±: clustered 95% CI) | Functional ρ +0.45/+0.46, CI excl. 0; LOCO ρ ≈ 0.4; geometry n.s.; paired differences n.s. |
|
||||
|
||||
## SI Methods (per tier — full details in the per-experiment READMEs and configs)
|
||||
## SI Methods: experimental procedures
|
||||
|
||||
**Analytic tier (E1–E14).** Wright–Fisher simulator over K-item distributions; closed-form validation
|
||||
suite (`tests/test_scientific_validation.py`, <0.5% tolerances); learning kernel; multi-locus
|
||||
genotypes, NK landscapes, n-parent crossover (E7–E11); BDM speciation model (E12); mating structure
|
||||
(E14). Bitwise reproducible from master seeds.
|
||||
Every experiment in this paper is one YAML config, one runner invocation, and one artifact triple
|
||||
(`results.parquet` + the resolved config + a manifest carrying the master seed, git commit, library
|
||||
versions, and a content hash). The configs named below are the authority on any parameter; this
|
||||
section gives the scientific reasoning behind the choices. `REPRODUCING.md` maps each manuscript
|
||||
panel to the config and seed that produced it.
|
||||
|
||||
**Neural tier.** Histogram bridge (exact reduction to the analytic tier — the harness gate);
|
||||
RNN/MLP/VAE collapse+grounding on a synthetic mode universe with an exact oracle; conv-VAE on MNIST
|
||||
with a frozen CNN oracle (98.5% mode accuracy; 30x30 confusion matrix recorded as the measurement
|
||||
floor); E13 speciation: no-BatchNorm MLPs, weight-average merges, LMC error barriers before/after
|
||||
alignment under the complete unit symmetry group (deterministic Re-Basin permutation matching composed
|
||||
with exact scale canonicalisation; sanity gate recovers a permuted-and-rescaled copy exactly);
|
||||
pre-registered emergent conditions (disjoint classes; shifted-view conventions).
|
||||
### M1. Design principles
|
||||
|
||||
**Language-model tier.** LoRA rank-16 specialists on procedural task families with an exact-match
|
||||
verifier; Qwen2.5-Instruct 0.5B/7B; operators: soup/TIES adapter arithmetic, per-input routing,
|
||||
Dirichlet offspring populations screened on held-out validation; multi-seed protocol (fixed tests,
|
||||
varied training seed); speciation knobs (conflicting conventions on ambiguous prompts; duration);
|
||||
the controlled predictive test (six pre-merge predictors; three decorrelated axes; robust statistics
|
||||
via `figures/stats_llm_epistasis.py`: condition-clustered bootstrap, paired contrasts,
|
||||
leave-one-condition-out prediction, three outcome references). Statistical (per-seed)
|
||||
reproducibility documented for GPU tiers.
|
||||
Four rules govern every choice that follows.
|
||||
|
||||
*Test each claim at the cheapest tier that can falsify it.* A closed form beats a simulation, a
|
||||
simulation beats a trained network, and a small network beats a language model, whenever the cheaper
|
||||
instrument can still return the answer "no". A costlier tier is entered only where it adds a
|
||||
discriminating test rather than a replication — which is why several cells of the programme (Fig. 1A)
|
||||
are deliberately empty.
|
||||
|
||||
*Match the precision of the claim to the precision of the instrument.* The biological model is exact,
|
||||
so it carries the paper's quantitative statements. Trained systems add optimisation noise and
|
||||
inductive bias, so at those tiers I claim signs and orderings, never magnitudes.
|
||||
|
||||
*Make reality able to refuse.* Every tier has an oracle that is independent of the model being
|
||||
measured: a fixed true distribution in the biological model, a lossless identity code or a frozen
|
||||
classifier in the neural tier, an exact-match verifier over procedurally generated tasks in the
|
||||
language-model tier.
|
||||
|
||||
*Declare the falsifier before running.* Each experiment states the outcome that would refute the
|
||||
claim it tests (per-experiment READMEs; SI Table S1). Two pre-registered predictions failed, and are
|
||||
reported as failures in the main text.
|
||||
|
||||
### M2. Replication: what a replicate is, and how many
|
||||
|
||||
A replicate means something different at each tier, and conflating the three would misstate what the
|
||||
error bars cover.
|
||||
|
||||
In the biological model a replicate is an independent lineage: a fresh random stream driving the same
|
||||
resolved config, with sub-seeds derived from the master seed by `SeedSequence.spawn`. Because drift
|
||||
*is* the object of study, the spread across replicates is signal rather than nuisance, and replicate
|
||||
counts are set so that the confidence interval on the summary statistic is small relative to the
|
||||
effect being reported.
|
||||
|
||||
In the trained-network tier a replicate is an independent lineage including fresh weight
|
||||
initialisation and data ordering, so it carries optimisation noise on top of drift.
|
||||
|
||||
In the language-model tier a replicate is an independent *training* seed evaluated on *fixed* test
|
||||
sets. Holding the evaluation data constant while varying the training seed isolates training
|
||||
stochasticity, which is the quantity in doubt; it also means the seed-to-seed spread I report is not
|
||||
inflated by resampling the benchmark.
|
||||
|
||||
Replicate counts, and why each is what it is:
|
||||
|
||||
| Experiment | Replicates | Reasoning |
|
||||
|---|---|---|
|
||||
| E1, E2, E3, E5, E6 | 100 lineages | Long horizons (400–600 generations) with drift-dominated variance; 100 lineages put the CI on stationary diversity well inside the effect being resolved |
|
||||
| E4 | 200 | Outcomes are per-item binary retentions, the highest-variance quantity in the paper |
|
||||
| E7 | 20 | Trajectory contrast (sexual vs asexual adaptation speed), large and monotone |
|
||||
| E8 | 40 | The vertical claim; the headline separation, so the most replicated of the genotype experiments |
|
||||
| E9, E10 | 24 | Landscape sweeps where each point aggregates 200 offspring internally |
|
||||
| E11 | 12 | Four-arm ablation over 80 generations; arms separate by margins far exceeding the CI |
|
||||
| E12, E12_nk | 15 | Each point already averages 500 (E12) or 200 (E12_nk) offspring |
|
||||
| E14 | 20 | Breadth × ruggedness grid, 60 generations per cell |
|
||||
| kernel_sharpen, kernel_smooth | 24 | Two-parameter kernel fits against neural reference endpoints |
|
||||
| bridge | 60 | The harness gate: must detect *any* departure from the biological model, so the most replicated neural run |
|
||||
| grounding | 18 | Nine-point grounding sweep with per-generation network retraining |
|
||||
| collapse, architectures | 5 | Sign-level demonstrations across architectures; each lineage retrains a network 22–25 times |
|
||||
| recombination | 8 | Operator contrast in trained weights |
|
||||
| mnist_collapse | 4 | 15 generations × a conv-VAE retrained from scratch each generation; the contrast (30 modes vs 1) is categorical |
|
||||
| speciation_real, _cliff | 3 | Barrier decomposition; the quantity is a near-deterministic function of the training condition (residual 0.001 vs 0.497) |
|
||||
| speciation_real_emergent | 4 | A null: replicates are spent on longer divergence horizons rather than more repeats |
|
||||
| llm_merge_seeds | 5 training seeds | The Fisher–Muller signature, the most-replicated language-model claim |
|
||||
| llm_moe_hard_seeds, llm_directed_hard_seeds, llm_epistasis(+compat), llm_speciation_add | 3 training seeds | Per-seed orderings reported individually rather than averaged |
|
||||
| 7B runs, llm_speciation | 1 | Single-run confirmations at a scale where each run costs GPU-hours; reported as sign-level and labelled as single runs |
|
||||
|
||||
The asymmetry is deliberate: replicates are cheap exactly where the quantitative claims live, and the
|
||||
expensive tiers are asked only for the sign of an effect the cheap tier has already quantified. Where
|
||||
a single run is all there is, the manuscript says so.
|
||||
|
||||
### M3. The biological-model tier
|
||||
|
||||
Knowledge is a distribution over `K` discrete items; reality is a fixed Zipf-tailed distribution
|
||||
`p*`; one generation resamples `n` draws from the parent, optionally mixes in `m` verified draws from
|
||||
`p*`, and refits. Implementation: NumPy/SciPy, no GPU, bitwise reproducible.
|
||||
|
||||
*Parameter choices.* `K = 500`–`1000` with `zipf_s = 1.1` and half the items designated tail: large
|
||||
enough that the rare tail contains hundreds of items (so tail statistics are not dominated by a
|
||||
handful of them) and small enough to sweep densely. `n = 100`–`200` sets drift strength; it is the
|
||||
population size in the Wright–Fisher correspondence and the distillation sample size in the AI
|
||||
reading. Horizons of 400–600 generations were chosen so that ungrounded lineages reach fixation and
|
||||
grounded ones reach stationarity within the run, which the trajectories confirm.
|
||||
|
||||
*Sweeps.* E2 sweeps grounding `g ∈ {0, 0.005, 0.01, 0.02, 0.05, 0.1, 0.2, 0.4}`; E3 contrasts uniform
|
||||
against region-matched grounding allocation; E4 crosses parent count `K_T ∈ {1,2,3,5}` with teacher
|
||||
correlation `ρ ∈ {0, 0.25, 0.5, 0.75, 1}` and `g ∈ {0, 0.02, 0.05}`; E5 crosses selection mode
|
||||
(none / greedy / quality-diversity) with novelty weight; E6 compares four re-minting arms.
|
||||
|
||||
*The correlated-parent construction (E4).* Teacher correlation is constructed directly rather than
|
||||
obtained by tuning drift, so that `ρ` is not confounded with `n`, `m`, tail size, or generation
|
||||
count. For each tail item a shared switch `z ~ Bern(ρ)`, a shared retention `s ~ Bern(q)`, and
|
||||
per-teacher `u⁽ᵏ⁾ ~ Bern(q)` give teacher `k` retention `s` if `z` else `u⁽ᵏ⁾`. This yields exact
|
||||
marginal retention `q` and exact pairwise correlation `ρ`, and is exchangeable, so `ρ` is a single
|
||||
scalar knob.
|
||||
|
||||
*Multi-locus experiments (E7–E11, E14).* Genotypes are `L = 12` biallelic loci (4096 genotypes —
|
||||
effectively open-ended relative to the population sizes used), with fitness either additive or a
|
||||
Kauffman NK landscape whose interaction count `K` tunes ruggedness from 0 to 10. E9 and E10 breed
|
||||
from `n_parents = 6` local optima into populations of 200 offspring; E10 additionally screens
|
||||
offspring and iterates (5 rounds, keeping 8). E11 runs a population of `N = 60` agents for 80
|
||||
generations at ruggedness `K = 8`, with mutation `μ = 0.03`, 120 offspring per generation, and
|
||||
selection weighting true fitness against consensus conformity at `g = 0.85`. E14 sweeps mate-pool
|
||||
breadth on a ring of `N = 48` against ruggedness.
|
||||
|
||||
*Speciation (E12).* `L = 20` loci, incompatibility density `ρ ∈ {0.1, 0.25, 0.5}`, parental
|
||||
divergence swept 0–20 substitutions, 500 offspring per cell at recombination rate 0.5. E12_nk repeats
|
||||
the question on NK landscapes (`L = 16`, `K` 0–10, 40 parent pairs, 200 offspring).
|
||||
|
||||
*Validation.* Three closed forms are asserted as standing tests to within 0.5%: neutral
|
||||
heterozygosity decay `E[H_t] = H_0(1 − 1/n)^t`, the exact immigration–drift equilibrium, and the
|
||||
multi-parent union formula. These run in CI alongside the correctness tests. If they fail, the
|
||||
science is wrong rather than merely the code.
|
||||
|
||||
### M4. The trained-network tier
|
||||
|
||||
*Why a synthetic universe.* Measuring collapse requires knowing the true distribution exactly. Each
|
||||
mode is rendered as a token sequence carrying an identity segment (base-2 digits encoding the mode
|
||||
index losslessly, so the oracle reads the mode back with zero error) followed by style tokens drawn
|
||||
uniformly at random. The style segment gives genuine within-mode entropy, so a generative model must
|
||||
learn a distribution rather than memorise `K` fixed strings, while the identity segment keeps the
|
||||
measurement noise-free. Mode truth comes from the same `make_true_distribution` used by the
|
||||
biological model, so "mode", "region", and "tail" denote the same objects at both tiers.
|
||||
|
||||
*The bridge gate.* Before any trained model is interpreted, a histogram generator is run through the
|
||||
identical harness; it must reproduce the biological model exactly. This separates harness bugs from
|
||||
model behaviour, and is why the bridge run carries 60 replicates.
|
||||
|
||||
*Architectures and training.* The recurrent generator is an embedding (24) → GRU (128 hidden; 192 in
|
||||
the architecture-generality run) → linear readout, trained each generation from scratch with Adam,
|
||||
learning rate 2×10⁻³, batch size 256, 25 epochs, and evaluated by sampling 12,000–15,000 sequences.
|
||||
Feedforward and variational autoencoder generators share the harness. Retraining from scratch each
|
||||
generation (rather than fine-tuning) makes the generational step a clean refit, matching the
|
||||
biological model's operator.
|
||||
|
||||
*MNIST tier.* Dataset: MNIST via torchvision (60,000 training images). Modes are digit class ×
|
||||
stroke-thickness bin (10 × 3 = 30 modes) with a Zipf frequency profile, so roughly eighteen modes are
|
||||
rare. The generator is a convolutional variational autoencoder (latent 32, β = 1), retrained from
|
||||
scratch each generation with Adam, learning rate 10⁻³, batch 256, 30 epochs, on 6,000 images drawn
|
||||
from the previous generation's own samples, for 15 generations, at `g ∈ {0, 0.1}`. The oracle is a
|
||||
frozen two-convolution classifier trained once (5 epochs) combined with a deterministic thickness
|
||||
measure; it reaches 98.5% mode accuracy and its 30 × 30 confusion matrix is recorded in the manifest
|
||||
as the measurement floor. Build gates: the oracle's accuracy, and generation-0 recovery of all 30
|
||||
modes.
|
||||
|
||||
*Speciation in trained weights.* Two multilayer perceptrons (784–512–512–10, ReLU, no batch
|
||||
normalisation — batch statistics would break the permutation correspondence the analysis depends on)
|
||||
are forked from a shared base trained for 500 steps, then trained apart for 100–3,200 further steps
|
||||
(up to 6.4× the shared base) under SGD at learning rate 0.05, batch 128. Merges are weight averages;
|
||||
the readout is the linear-mode-connectivity error barrier before and after alignment. Alignment
|
||||
composes deterministic Re-Basin permutation matching with exact per-unit scale canonicalisation —
|
||||
the unit symmetry group of this architecture — and is gated by a control that must recover a
|
||||
permuted-and-rescaled copy exactly. Since the search space is that group rather than all
|
||||
possible alignments, the removable share is a lower bound and the residual an upper bound.
|
||||
|
||||
### M5. The language-model tier
|
||||
|
||||
*Base models.* Qwen2.5-Instruct at 0.5B and 7B, open weights under a permissive licence, with the
|
||||
revision pinned. Using two sizes from one family makes scale the only variable that changes between
|
||||
the small and large runs; the 0.5B model carries the multi-seed protocols and the 7B model the
|
||||
single-run confirmations.
|
||||
|
||||
*Task families, and why they are procedural.* Three deliberately disjoint families — list
|
||||
operations, string transformations, and small-integer arithmetic — are generated procedurally from a
|
||||
seed. Procedural generation buys four things that a standard benchmark cannot: an exact-match
|
||||
verifier that plays the role of reality (an answer is right or it is not, with no judge model in the
|
||||
loop); freedom from train/test contamination, since every evaluation item is generated fresh from a
|
||||
disjoint seed offset; control over family disjointness, which is the precondition for specialists to
|
||||
be genuinely decorrelated parents; and a difficulty knob. A `hard` variant (multi-step list
|
||||
operations, Caesar ciphers and letter transforms, multi-step and larger arithmetic) exists because
|
||||
the easy families saturate a 7B base at ceiling, and saturation removes the headroom in which
|
||||
recombination operators can differ — a control that proved necessary, since two null results at 7B
|
||||
turned out to be saturation artefacts rather than scale effects.
|
||||
|
||||
*Data splits.* Training, validation, routing-calibration, and test items are drawn from
|
||||
non-overlapping seed offsets by construction (test from 1000 + family index, routing from 2000 +,
|
||||
validation from 3000 +, training from the run seed). Test sets are fixed across seeds in the
|
||||
multi-seed protocols. Selection of merge weights uses validation only; the winners are then reported
|
||||
on the untouched test split.
|
||||
|
||||
*Specialisation.* Each parent is a LoRA adapter (rank 16, α = 32) on the frozen base, applied to all
|
||||
attention and MLP projection matrices, trained with a manual supervised fine-tuning loop: answer-only
|
||||
cross-entropy (prompt tokens masked out of the loss), AdamW at 2×10⁻⁴, batch size 8, 3 epochs,
|
||||
bfloat16, 400–800 training items per family. Low-rank adaptation is the right instrument here for a
|
||||
structural reason rather than a computational one: it confines each parent's specialisation to an
|
||||
additive low-rank delta over an identical frozen base, which is what makes weight-space recombination
|
||||
between parents well defined.
|
||||
|
||||
*Recombination operators.* Fusion by uniform weight averaging (soup) and by sign-reconciled,
|
||||
magnitude-pruned task arithmetic (TIES); union by keeping specialists intact and selecting per input
|
||||
(oracle routing, and a training-free nearest-centroid router over the base model's own prompt
|
||||
embeddings) or per module (winner-take-all by delta norm); and directed recombination, which breeds a
|
||||
population of Dirichlet-weighted merges, scores each on validation, and keeps the fittest.
|
||||
|
||||
*Evaluation.* Greedy decoding, exact match after canonicalisation. Alongside overall accuracy I
|
||||
report worst-family accuracy, because the Fisher–Muller claim is about competence across all
|
||||
families rather than an average that a single strong specialty can carry.
|
||||
|
||||
*The controlled predictive test.* Thirty-nine parent pairs (13 conditions × 3 seeds) span three axes
|
||||
that are decorrelated by construction: conflict (contradictory conventions on shared prompts, with
|
||||
private training budgets held fixed), compatible overlap (the same shared prompts under the same
|
||||
convention — overlap and volume without conflict), and duration (weight divergence with no conflict,
|
||||
1 to 12 epochs). Six predictors are computed before any merge: confidence-weighted functional
|
||||
conflict, raw disagreement, gradient alignment at the shared base, LoRA-delta cosine and L2 distance,
|
||||
and a cross-task performance baseline. Probes are drawn blind to where the conflict lives. The
|
||||
outcome is the merge penalty against oracle parent potential, pre-registered, and also reported
|
||||
against best-parent and mean-parent references because the predictor ordering is sensitive to that
|
||||
choice.
|
||||
|
||||
*The composed society.* A population of `N` LoRA agents on a shared frozen base evolves for `G`
|
||||
non-overlapping generations. Each generation every agent answers a fixed validation pool (verifier
|
||||
scored) and a fresh conformity pool (whose modal answer defines the population consensus); selection
|
||||
scores agents by `g·fitness + (1−g)·conformity`; parents are chosen with or without a
|
||||
quality-diversity term over behavioural distance; offspring are bred by screened recombination; and
|
||||
each child is a fresh adapter distilled from its source model's own answers, which makes the
|
||||
inheritance channel literally self-consuming. The verifier enters the loop only where `g > 0`, but is
|
||||
used for reporting in every arm. The four-arm ablation removes grounded evaluation, recombination,
|
||||
or diversity preservation in turn.
|
||||
|
||||
### M6. Negative controls
|
||||
|
||||
The design leans on controls that can remove a result rather than support one, and one of them did.
|
||||
|
||||
The compatible-overlap axis was added to the predictive test specifically to expose overlap-and-volume
|
||||
artefacts, and it did: the initial two-axis grid's best predictor (LoRA-delta cosine, ρ = +0.60)
|
||||
collapsed to ρ = +0.03 once compatible overlap was present, identifying it as an artefact rather than
|
||||
a signal. The duration axis supplies divergence without conflict. The emergent-speciation condition
|
||||
supplies divergence with no conflicting signal anywhere, and returns a null. The budget-controlled
|
||||
speciation design (`conflict_mode: add`) removes the confound between conflict fraction and private
|
||||
training budget. The histogram bridge is a harness control. In the biological model, `m = 0` arms and
|
||||
`ρ = 1` (fully correlated parents) are the null conditions against which the corresponding effects
|
||||
are read.
|
||||
|
||||
### M7. Statistical procedures
|
||||
|
||||
Error bars on replicate means are normal-approximation 95% confidence intervals unless stated
|
||||
otherwise. For the predictive test, where rows share task-data seeds across conditions and are
|
||||
therefore not independent, inference uses a condition-clustered bootstrap (13 clusters, 4,000
|
||||
resamples); predictors are compared by paired contrasts on the same resamples; generalisation is
|
||||
assessed by leave-one-condition-out prediction; and per-seed and leave-one-seed-out sensitivity are
|
||||
reported alongside, since three seeds cannot settle seed generalisation on their own. Outcome-
|
||||
reference sensitivity is reported rather than resolved. Where a difference is not significant at the
|
||||
sample size available, the manuscript says so rather than reporting the point estimate alone.
|
||||
|
||||
## SI Statistics
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue