paper: claim-narrowing revision from the external review

The review's core instruments adopted: the interpretation/explanation/
prediction ladder is now explicit in §1 (with the decisive pre-merge
epistasis-prediction test stated as the open bar, not claimed); identity
claims scoped (WF exact only in the minimal model, with the
learning-kernel deviation cited against ourselves; Muller's ratchet
scoped to the irreversible arm — recombination reassembles only what
survives); "nobody has / none imports / theory outrun" removed;
merge-don't-average given explicit operator boundaries (output-mean vs
weight-average vs routing vs max-with-oracle; budgets; oracle; capacity
handoff to speciation); a "what these experiments do and do not
establish" scope block added to the speciation section (conflict floor
is information-theoretic, not genetic; epistasis-cliff + snowball =
hypotheses at the neural tier; emergent DMIs = flagship hypothesis,
bounded by our null); "control theory" -> "framework" (subtitle
included); §3/§11 overstatements fixed (frozen core != frozen behaviour;
Baldwin echo, not identity; operational vs archival irreversibility);
claims-at-a-glance table (status/assumptions/evidence/limits) added to
§13. Reviewer's framing sentence adopted as the stated core
contribution. Accessible version calibrated to match. md2tex gains pipe-
table support; PDF rebuilds clean (22 pp). Lessons recorded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
Giorgio Gilestro 2026-09-06 15:14:28 +01:00
parent d6a5c5cacd
commit 58e6c74609
6 changed files with 238 additions and 86 deletions

View file

@ -1,6 +1,6 @@
# The Evolution of Sex for Artificial Intelligence
### A population-genetic control theory for societies of agents that reproduce, recombine, and stay open-ended
### A population-genetic framework for societies of agents that reproduce, recombine, and stay open-ended
*A perspective, written from a geneticist's chair. Companion to a set of minimal, reproducible working
models and a first language-model prototype (both built).*
@ -28,13 +28,15 @@ you skipped a definition:
- **Genetic drift** *(population genetics)* — the random loss of rare variants that happens in any
finite population simply because not everyone leaves offspring. It is the neutral, no-selection
baseline of evolution.
- **WrightFisher process** *(population genetics)* — the standard mathematical model of drift. We
will claim, and show, that generational model-training *is* this process, not merely like it.
- **WrightFisher process** *(population genetics)* — the standard mathematical model of drift. Our
minimal model of knowledge transmission *is* this process exactly; a real trained network is this
process plus a measurable, architecture-specific bias we quantify.
- **Recombination / sexual reproduction** *(biology)* — making an offspring by combining pieces from
more than one parent, rather than copying a single parent (which is *asexual* reproduction).
- **Muller's ratchet** *(population genetics)* — the way an asexual lineage, one that never
recombines, accumulates damage it can never undo. It is, we will argue, the same thing as model
collapse.
recombines, accumulates damage it can never undo. We will argue it is the right lens for the
*irreversible* part of model collapse — the capabilities that, once lost from every parent, no
merging can rebuild.
- **Catastrophic forgetting** *(machine learning / neuroscience)* — a neural network overwriting what
it knew when it learns something new.
@ -53,27 +55,35 @@ AI is turning from single frozen models to **populations of agents** that persis
increasingly *recombined* into new models — a shift visible in multi-agent societies, population-based
self-improvement, and the explosion of **model merging**. The field is doing this with the vocabulary
of evolution — "crossover," "mutation," "mate choice," "offspring that beat their parents" — but as
loose metaphor draped over search algorithms. This paper argues that the right theory is already
written, in the branch of biology that studies exactly this: the **evolution of sex**. Ninety years of
population genetics say precisely when reproducing a population by *recombination* beats copying, when
it backfires, and how to do it better — and, read as a control theory, it tells an engineer how to keep
a society of models learning across generations instead of decaying.
loose metaphor draped over search algorithms. This paper argues that a rich, quantitative body of
applicable theory already exists in the branch of biology that studies exactly this: the **evolution
of sex**. Ninety years of population genetics analyse when reproducing a population by *recombination*
beats copying, when it backfires, and how to do it better — and, read as an engineering framework, it
supplies overlooked variables and testable design rules for keeping a society of models learning
across generations instead of decaying. The underlying shift of perspective is the contribution we
most want to land: **treat multigenerational model populations as systems whose inheritance,
diversity, and compatibility must be managed — not merely as collections of models to optimise.**
We take one diagnosis as settled and cite it as such: training each generation on the last is
**genetic drift**, and the resulting **model collapse** is the loss of rare variants a finite
population always suffers (the WrightFisher process; formalised for language models by Shumailov et
al., 2024, and Riis, 2026). We claim none of that. Our contribution is the other half — the **cure**,
and its assembly into a theory with predictions. Single-teacher copying is **asexual** reproduction,
and asexual lineages decay by **Muller's ratchet**, which *is* model collapse; the remedy nature found
is **sex**. A society of models should reproduce sexually — each new model **recombined from several
complementary parents** (which the field already does, as *model merging*), selection **anchored to a
reality that can say no** (not to the consensus of other models), and diversity actively **preserved**.
With those three ingredients a lineage does not merely avoid collapse; it **climbs** — producing models
fitter than any ancestor (the **FisherMuller effect**) while each specialty is re-earned and exceeded.
al., 2024, and Riis, 2026). We claim none of that. Our contribution is on the remedy side. Single-
teacher copying is **asexual** reproduction, and the irreversible arm of its decay corresponds to
**Muller's ratchet** (a correspondence we state with its scope, not as identity); the remedy biology
found for the ratchet is **sex**. A society of models should reproduce sexually — each new model
**recombined from several complementary parents** (which the field already does, as *model merging*),
selection **anchored to a reality that can say no** (not to the consensus of other models), and
diversity actively **preserved**. In our models — from closed-form to trained networks to a
language-model prototype — those three ingredients together let a lineage not merely avoid collapse
but **climb**, producing models fitter than any ancestor (the **FisherMuller effect**) while each
specialty is re-earned and exceeded; whether the full recipe holds at frontier scale is the open
question the framework is built to test.
From the geneticist's apparatus we extract falsifiable, load-bearing claims the merging literature has
not: (i) **"merge, don't average"** — recombination preserves the union of what parents kept, while
averaging (a "model soup") is *blending inheritance* that mathematically cancels the benefit; (ii)
From the geneticist's apparatus we extract falsifiable, load-bearing claims (each stated with its
operator and scope in the text): (i) **"merge, don't average"** — a conservation result: refitting a
child to the *mean of its parents' output distributions* conserves rare-capability mass at the
single-parent level, so adding parents cannot help, while union-preserving operators realise the gain
— exact in the minimal model, with its weight-space image verified as the headroom rule below; (ii)
**offspring can exceed every parent** (FisherMuller), the real argument for sex in model societies;
(iii) on **rugged, epistatic** task landscapes, blind recombination causes **outbreeding depression**,
yielding a design rule — *merge freely when skills are additive, sparingly and with selection when
@ -143,18 +153,24 @@ mixture-of-experts routing): all established. Third, that merge success can be *
machine-learning-native predictors exist, from interpretable pairwise metrics (gradient alignment —
Zhou et al., 2026) to capacity/rate-distortion accounts of "merging collapse" (2026); what they lack,
and we supply, is the *mechanism* — when and why the failure is a coordinate artefact versus genuine
functional incompatibility, and what moves the cliff. What is genuinely unoccupied — and what a
geneticist is placed to supply — is a **theory** rather than a search heuristic. The nearest precursor
is a theory-of-computation tradition reading sex as an algorithm for *mixability* (Livnat &
Papadimitriou, 2016), pre-dating model merging and never applied to it. Every one of the works above
uses evolution as *metaphor over an optimiser*; none imports the predictive apparatus of the evolution
of sex. Nobody has stated the **"merge, don't average" conservation law**, derived **offspring-exceed-parents
as FisherMuller**, predicted **outbreeding depression on rugged task landscapes**, framed **grounding
as migrationdrift balance** with a critical fraction, or connected **reproductive isolation** to when
two models can be merged at all. An evolutionary algorithm that *finds* a super-parent is evidence for
the theory, not a substitute for it — the way CMA-ES existing does not make fitness-landscape theory
redundant. This paper supplies the theory the tinkering has outrun, and states what it predicts and
where it would fail.
functional incompatibility, and what moves the cliff. What a geneticist is placed to supply is a
**framework** rather than a search heuristic. The nearest precursor is a theory-of-computation
tradition reading sex as an algorithm for *mixability* (Livnat & Papadimitriou, 2016), pre-dating
model merging; the works above use evolution chiefly as vocabulary over an optimiser, and — to our
knowledge — the quantitative apparatus of the evolution of sex (FisherMuller, outbreeding depression,
migrationdrift balance, reproductive isolation) has not previously been carried over as more than
metaphor. We are also candid about what *kind* of contribution each of our claims is, because three
different things are easily conflated: **interpretation** (an existing result is usefully understood
in these terms — e.g., merged offspring beating their parents as FisherMuller), **explanation** (the
transferred mechanism accounts for observations existing accounts leave open — e.g., which merge
failures are coordinate artefacts and which are functional), and **prediction** (the framework
forecasts an unmeasured outcome and improves a design decision — e.g., an epistasis measure taken
*before* merging that beats geometry-based predictors of merge success). This paper is strongest on
the first, makes concrete progress on the second, and states the third as its open, decisive test —
proposed here with pre-registered falsifiers, not claimed as done. The organising shift we argue for
is prior to any single mechanism: **treat multigenerational model populations as systems whose
inheritance, diversity, and compatibility must be managed — not merely as collections of models to
optimise.**
## 2. Why today's models cannot do this
@ -177,9 +193,11 @@ The individual model needs two properties.
core frozen and only *readable*, and carves each new skill into freshly-added capacity beside it. In
machine learning this is called *parameter isolation* (progressive networks — Rusu et al., 2016;
prune-and-freeze — Mallya & Lazebnik, 2018; and, most practically, **LoRA** and other small trainable
"patches" bolted onto a frozen model — Hu et al., 2021). If the core is never altered, forgetting it
is not merely unlikely but structurally impossible. This is what lets a model accumulate a coherent
working life of expertise — the kind of stable knowledge worth passing on.
"patches" bolted onto a frozen model — Hu et al., 2021). If the core is never altered, its *parameters*
cannot be forgotten — though a precise reader should note the system's *behaviour* can still shift
while adapters are active, so the guarantee is of a recoverable core, not of unchanging conduct. This
is what lets a model accumulate a coherent working life of expertise — the kind of stable knowledge
worth passing on.
The brain offers a partial blueprint. *Complementary Learning Systems* theory (McClelland,
McNaughton & O'Reilly, 1995) — itself a response to the forgetting problem — describes two subsystems:
@ -221,20 +239,33 @@ distribution (the rare cases) first, and drifts toward its own most common outpu
idiosyncratic* — **is** tail-deletion by design. The operation that would power a cultural ratchet and
the operation that drives model collapse are the same act.
**The population-genetics statement (the same thing).** Represent a model's knowledge as a
distribution over discrete "items" — capabilities, facts, modes of behaviour. One generation is:
*draw a finite sample from the parent, and refit the child to it.* That finite-sampling step is
**mathematically identical** to **genetic drift** — the random loss of rare variants in a finite
population — described by the century-old **WrightFisher** model (Wright, 1931; Fisher, 1930). This is
not an analogy we find pretty; it is the same equations, and we use them as an exact check on our
simulations (the first of the minimal models below). Rare items go extinct first, roughly ten times
faster than common ones, precisely as drift predicts.
**The population-genetics statement (the same thing, for the minimal model).** Represent a model's
knowledge as a distribution over discrete "items" — capabilities, facts, modes of behaviour. One
generation is: *draw a finite sample from the parent, and refit the child to it.* In this **minimal
inheritance model** the finite-sampling step is **exactly** genetic drift — the random loss of rare
variants in a finite population — described by the century-old **WrightFisher** model (Wright, 1931;
Fisher, 1930): the same equations, which we use as closed-form checks on our simulations. Rare items
go extinct first, roughly ten times faster than common ones, precisely as drift predicts. **The
boundary of the identity matters, and we measured it:** real neural training adds approximation,
optimisation noise, and inductive bias on top of sampling, and when we fit trained networks against
the exact drift null they deviate in *opposite, architecture-specific directions* — a smoothing
recurrent model resists collapse (it keeps spurious variants alive), a sharpening image generator
accelerates it (our learning-kernel result, below). So the honest statement is: the minimal
inheritance model is exactly WrightFisher; a real learner is WrightFisher *plus a signed,
measurable estimator-bias operator* — and the drift signs (rare-first loss, the grounding response)
survive that operator in every architecture we tested.
And single-teacher copying is **asexual reproduction** — cloning one parent. Nature already knows what
happens to an asexual lineage that never recombines: it accumulates damage it can never repair, a
one-way decline geneticists call **Muller's ratchet** (Muller, 1964). *Muller's ratchet is model
collapse.* Naming it that way is not decoration; it tells us where the cure is, because biology solved
this problem.
one-way decline geneticists call **Muller's ratchet** (Muller, 1964). We use the ratchet as the
*organising correspondence* for model collapse, with its scope stated: strictly, the ratchet is the
stochastic loss of the least-degraded class under recurring deleterious change in an asexual
population, so it maps onto the *irreversible* component of capability loss (once every copy of a rare
capability is gone from all parents and sources, no recombination can rebuild it) rather than onto
every form of degradation. That is exactly why the correspondence is useful rather than decorative: it
says the cure must act *before* fixation-by-loss — keep complementary variants alive somewhere in the
population — because recombination can only reassemble what still survives. Biology solved this
problem, and its solution is the subject of this paper.
Two ingredients turn the collapse operation into a climb. Both are things nature does.
@ -282,6 +313,23 @@ beats the average in exact proportion to how far the average is from the best at
that sharpens rather than weakens the claim, and that a practitioner needs before spending compute on
the fancier operator.
**The operator boundaries (stated, because "merge, don't average" is not one claim but a family).**
Four different operators travel under these words, and the conservation result belongs to exactly one
of them. What is *derived* is this: when a pupil's knowledge is refit to the **mean of the parents'
output distributions**, the expected mass on any rare item is conserved at the single-parent level —
in the rare-item regime the 1/K dilution of averaging exactly cancels the union gain of having K
parents — so adding parents cannot help; whereas an operator that keeps, per item, its **strongest
source** realises the union. That statement is exact in the minimal model, and it presupposes an
oracle (or verifier) able to say which source is strongest. The two operators the LLM prototype
tests — **weight averaging** (a nonlinear network's weight-mean does not compute the mean of its
parents' outputs) and **routing among intact specialists** (which keeps K models' storage and an input
classifier, a different parameter and inference budget from one fixed-size child) — are *empirical
cousins* of the two sides of that law, not instances of it. The headroom rule above is precisely the
empirical bridge: it says when the weight-average behaves like the diluting mean (hard tasks, weak
base) and when a capable base absorbs the dilution (easy tasks). And all of it operates within a
capacity boundary: when parental capabilities genuinely cannot coexist in the child's capacity, no
operator preserves the union — that regime is the subject of the speciation section below.
Three results keep this honest, and all are results, not hand-waving.
*Sex can backfire.* When the parents' skills are not cleanly separable but **entangled** — when the
@ -413,6 +461,26 @@ longer merge worse under averaging) — is exactly the next tier's question, and
prediction crisp: it should depend on whether extended training induces *conflicting conventions on
shared circuitry*, not on divergence time itself.
**What these experiments do and do not establish.** Stated at exactly the strength of the evidence:
they establish that *some merge failures reflect incompatible functional requirements rather than a
mismatch of coordinates* — a residual that survives the full unit-symmetry group of the architecture
tested, rises with functional conflict, and is absent under compatible specialisation. Three
qualifiers. First, the impossibility at the heart of the conflict condition — one deterministic model
cannot satisfy two contradictory answer conventions — is information-theoretic and needs no population
genetics; what the genetic frame adds is *structure around it*: which divergences generate such
conflicts, the prediction that epistasis rather than distance sets the cliff's position, and the
snowball's super-linear onset — the latter two verified so far only in the analytic model, and
therefore carried as **hypotheses at the neural tier, not results**. Second, our alignment removes the
symmetries we enumerate for this architecture class; richer transformation families for other
architectures could reapportion removable vs residual, though not below the conflict floor. Third,
"unmergeable" here means by aligned linear interpolation of weights — a barrier to that operator does
not preclude every conceivable recombination method (routing, for one, sidesteps it by not blending).
Emergent DobzhanskyMuller incompatibilities in real weights remain the flagship *hypothesis* of this
programme: our tested regimes found none, which bounds where they can live — longer horizons, shifted
data distributions, capacity pressure — and the decisive experiment (predicting merge success *before*
merging from an operational epistasis measure, against geometry- and gradient-based predictors) is
posed in the closing section.
One question remains, and the rest of the paper is largely about it: recombination combines what the
parents kept — but *who decides what each parent keeps, and which offspring are worth keeping?*
@ -571,25 +639,29 @@ across enough generations, **re-mint the base**: distil the accumulated soft inh
generations to acquire in patches. The soft budget resets; the next epoch begins from a richer floor.
What was hard-won and *learned* becomes cheap and *innate*.
This has a precise name, and it is not Lamarck's. Knowledge that is acquired and re-learned every
generation, and — once reliably present for long enough — becomes part of the innate endowment so that
it need no longer be re-learned, is the **Baldwin effect** (Baldwin, 1896; and its clean computational
demonstration, Hinton & Nowlan, 1987). It is the valve between the two substrates: the soft, learned
patches, and the hard base weights every model is born with.
The pattern **echoes the Baldwin effect** (Baldwin, 1896; its clean computational demonstration is
Hinton & Nowlan, 1987): knowledge acquired and re-learned every generation eventually becoming part of
the innate endowment. We use the echo advisedly — Baldwin's mechanism is *selection* favouring
genotypes that learn the trait ever more easily, whereas re-minting is direct distillation, a
deliberate engineering shortcut through the same soft-to-innate valve. The valve is the point: two
substrates, the soft learned patches and the hard base weights every model is born with, with a
controlled passage between them.
Three honest riders, because re-minting is the most consequential step in the scheme:
- **Cost.** This is the one step that re-pays part of the pre-training bill, breaking §10's cheapness
*locally*. It is bearable only because it is *rare*, amortised over many cheap generations, and is
continued training from the lineage's own rich outputs rather than a de-novo run.
- **Irreversibility.** Until now, one thing was always recoverable — the original pristine base, whose
lost tails could be restored just by reloading the file. Bake the current lineage into new immutable
weights and that escape hatch closes: if the lineage had been quietly collapsing, re-minting *fixes
the collapse in place* and discards the one uncollapsed reference that could have diagnosed it. In our
minimal models this is exactly what happens, and a cheap safeguard prevents it: **re-mint only while
the lineage is demonstrably diverse and healthy**, never as a rescue for a line already drifting. It
is the sharpest instance of the human seat of §9 — choosing what no future generation will think to
question.
- **Irreversibility (of the lineage, not the archive).** A digital system can, of course, keep every
old base on disk — nothing forces deletion, and archives should be kept. The irreversibility is
*operational*: once the lineage's production base, training mixtures, and selection all run downstream
of the re-minted weights, a quiet collapse baked into them propagates to every descendant, and the
archived ancestor helps only if some process still compares against it — which nothing in the loop
does by default. In our minimal models a collapsed-then-re-minted lineage locks in its loss exactly
this way, and a cheap safeguard prevents it: **re-mint only while the lineage is demonstrably diverse
and healthy** (and keep an audit that diffs against the archived ancestor), never as a rescue for a
line already drifting. It is the sharpest instance of the human seat of §9 — choosing what no future
generation will think to question.
- **Speciation.** A re-minting is a founder event. Different laboratories, re-basing on different
criteria, will mint divergent bases; the lineage branches. This is not a defect but *adaptive
radiation*, and it is exactly what open weights make possible. The society grows not as one heavy
@ -678,6 +750,26 @@ shape.)
differently; only the whole system climbs. This is the closest thing we have to a test of the actual
thesis, rather than of the borrowed scaffolding around it.
### The claims at a glance: status, assumptions, evidence, limits
Because a perspective of this breadth risks blurring what is proved, what is measured, and what is
proposed, here is the ledger of the load-bearing claims — each labelled **exact** (closed-form in the
minimal model), **empirical** (measured in trained systems), or **hypothesis** (stated with a
falsifier, not yet established):
| Claim | Status | Key assumptions | Evidence | Known limits |
|---|---|---|---|---|
| Collapse = WrightFisher drift (minimal model) | Exact (diagnosis conceded to prior work) | Knowledge = categorical distribution; refit = resample | Closed forms reproduced to <0.5% | Real learners add a signed, architecture-specific estimator bias (measured) |
| Grounding = immigration; critical real-data fraction ≪ 1 | Exact + empirical sign | Fresh samples from a fixed, non-drifting truth | Exact `H_eq`; `g*≈0.048`; sign holds in RNN/MLP/VAE and on MNIST | Deepest tail unrescuable at feasible budgets (`m 1/p`); sharp threshold softens in trained nets |
| "Merge, don't average" conservation | Exact **for the output-mean operator** | Rare-item regime; an oracle/verifier identifies the strongest source | E4 closed form + simulation; neural reproduction | Weight-averaging and routing are empirical cousins, not instances; budgets differ; bridge = the headroom rule |
| Offspring exceed every parent (FisherMuller) | Interpretation + empirical | Complementary (decorrelated) parents; verifiable fitness | E8 analytic; 7B LoRA merge beats every specialist on every family | LLM tier: 3 lexically-distinct families; multi-seed replication in progress |
| Outbreeding depression on rugged landscapes; operator design rule | Exact-model result; hypothesis at LLM scale | NK epistasis stands in for skill entanglement | E9E10; directed selection rescues | Not yet mapped onto a real task-entanglement measure |
| Optimal mate-pool breadth shrinks with ruggedness | Exact-model result; hypothesis for merging populations | Ring population, local selection | E14 | Phenomenon known to island-model evolutionary computation; our contribution is the mapping and the diversity/mean decomposition |
| Merge failure decomposes into coordinate artefact + functional residual | Empirical (MLP tier; LLM tier in progress) | Alignment enumerates the architecture's unit symmetries | Full-symmetry residual ≈ 0 (compatible) vs ≈ naive (conflict); cliff in hybrid fitness | Scoped to aligned linear interpolation; conflict floor is information-theoretic, not genetic |
| Epistasis (not divergence) sets the cliff; snowball onset | Exact-model result; **hypothesis** at the neural tier | BDM incompatibility structure | E12 | The decisive pre-merge prediction test is proposed, not run |
| Emergent speciation without conflict | **Not observed** (pre-registered) | Shared ancestry, compatible tasks, tested divergences | E13b: residual 0.000; merge rescues specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
| Grounding + sex + diversity jointly necessary | Exact-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |
**What is borrowed, and what is ours.** We are deliberate about the ledger, because the surrounding
literature is crowded and a reader deserves to know exactly where the line falls. **Conceded as prior
art:** (a) *model collapse is genetic drift* — derived independently and cleanly (Riis, 2026; the
@ -692,8 +784,8 @@ capacity/rate-distortion accounts of merging collapse (Cao et al., 2026), and st
analyses of multi-task degradation; and (e) that verifier-screened synthetic data can avert collapse
(Yi et al., 2025) — the statistical cousin of our grounding operator. We claim none of these.
**Ours** is the theory those results have outrun: a **population-genetics of sex** applied to model
societies, which is *generative* where the incumbents are empirical. Concretely — the **"merge, don't
**Ours** is the framework those results invite: a **population-genetics of sex** applied to model
societies, generative where the incumbents are empirical. Concretely — the **"merge, don't
average" conservation law** (recombination preserves the union; blending inheritance cancels it),
derived not observed; **FisherMuller** named and used to explain *why* offspring exceed parents;
**outbreeding depression on rugged/epistatic landscapes**, which turns "when does merging help vs hurt"