Reframe of v5 into a population-genetic control theory for agent societies (leads with evolution-of-sex, concedes collapse=drift up front), positioned against the 2025-26 landscape (Multiagent-Finetuning, GENOME, M2N2, DGM, Pari 2024, Zhou 2026, Git Re-Basin) with an explicit concede/own ledger. Folds in E12 as the headline NEW modelling result: a dedicated 'The limit of sex: model speciation' section (compatible -> outbreeding depression -> hybrid inviability; the isolation cliff set by epistasis not divergence alone; the Orr-Turelli snowball; the route-don't- merge design rule), threaded through the abstract (5th load-bearing claim) and the what's-ours ledger. New draft file; v5 preserved. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
694 lines
53 KiB
Markdown
694 lines
53 KiB
Markdown
# The Evolution of Sex for Artificial Intelligence
|
||
|
||
### A population-genetic control theory for societies of agents that reproduce, recombine, and stay open-ended
|
||
|
||
*A perspective, written from a geneticist's chair. Companion to a set of minimal, reproducible working
|
||
models and a first language-model prototype (both built).*
|
||
|
||
**Giorgio F. Gilestro** · Department of Life Sciences, Imperial College London ·
|
||
giorgio@gilest.ro · https://lab.gilest.ro
|
||
|
||
---
|
||
|
||
### A note on vocabulary (please read this first)
|
||
|
||
This paper sits at the meeting point of three fields, and it is written so that a reader from any
|
||
one of them can follow all of it. We therefore **spell out** each field's jargon the first time it
|
||
appears, even at the risk of belabouring the obvious for the specialist. A short glossary, in case
|
||
you skipped a definition:
|
||
|
||
- **Model collapse** *(machine learning)* — the degeneration that happens when you train a model on
|
||
data produced by earlier models, over and over: rare cases disappear and the model drifts toward a
|
||
bland average.
|
||
- **Distillation** *(machine learning)* — training a fresh "student" model on the outputs of one or
|
||
more "teacher" models, so the student ends up knowing a compressed version of what they knew.
|
||
- **Model merging** *(machine learning)* — combining several trained models directly, at the level
|
||
of their weights, into one — no retraining. (Think of it as breeding two models rather than
|
||
teaching a third.)
|
||
- **Genetic drift** *(population genetics)* — the random loss of rare variants that happens in any
|
||
finite population simply because not everyone leaves offspring. It is the neutral, no-selection
|
||
baseline of evolution.
|
||
- **Wright–Fisher process** *(population genetics)* — the standard mathematical model of drift. We
|
||
will claim, and show, that generational model-training *is* this process, not merely like it.
|
||
- **Recombination / sexual reproduction** *(biology)* — making an offspring by combining pieces from
|
||
more than one parent, rather than copying a single parent (which is *asexual* reproduction).
|
||
- **Muller's ratchet** *(population genetics)* — the way an asexual lineage, one that never
|
||
recombines, accumulates damage it can never undo. It is, we will argue, the same thing as model
|
||
collapse.
|
||
- **Catastrophic forgetting** *(machine learning / neuroscience)* — a neural network overwriting what
|
||
it knew when it learns something new.
|
||
|
||
We have tried to keep the big picture legible on every page, and to be candid about what is argument
|
||
and what is evidence. The evidence is mostly from **deliberately small models** — mathematics, small
|
||
neural networks, image generators, and evolutionary simulations. A first bridge to real language
|
||
models exists — a prototype that recombines LoRA-specialised Qwen models up to 7B on a GPU cluster,
|
||
which confirms the recombination signs (below) — but the *full grounded society* has not yet been
|
||
built on a large language model. We will say so repeatedly, because the gap matters.
|
||
|
||
---
|
||
|
||
## Abstract
|
||
|
||
AI is turning from single frozen models to **populations of agents** that persist, specialise, and are
|
||
increasingly *recombined* into new models — a shift visible in multi-agent societies, population-based
|
||
self-improvement, and the explosion of **model merging**. The field is doing this with the vocabulary
|
||
of evolution — "crossover," "mutation," "mate choice," "offspring that beat their parents" — but as
|
||
loose metaphor draped over search algorithms. This paper argues that the right theory is already
|
||
written, in the branch of biology that studies exactly this: the **evolution of sex**. Ninety years of
|
||
population genetics say precisely when reproducing a population by *recombination* beats copying, when
|
||
it backfires, and how to do it better — and, read as a control theory, it tells an engineer how to keep
|
||
a society of models learning across generations instead of decaying.
|
||
|
||
We take one diagnosis as settled and cite it as such: training each generation on the last is
|
||
**genetic drift**, and the resulting **model collapse** is the loss of rare variants a finite
|
||
population always suffers (the Wright–Fisher process; formalised for language models by Shumailov et
|
||
al., 2024, and Riis, 2026). We claim none of that. Our contribution is the other half — the **cure**,
|
||
and its assembly into a theory with predictions. Single-teacher copying is **asexual** reproduction,
|
||
and asexual lineages decay by **Muller's ratchet**, which *is* model collapse; the remedy nature found
|
||
is **sex**. A society of models should reproduce sexually — each new model **recombined from several
|
||
complementary parents** (which the field already does, as *model merging*), selection **anchored to a
|
||
reality that can say no** (not to the consensus of other models), and diversity actively **preserved**.
|
||
With those three ingredients a lineage does not merely avoid collapse; it **climbs** — producing models
|
||
fitter than any ancestor (the **Fisher–Muller effect**) while each specialty is re-earned and exceeded.
|
||
|
||
From the geneticist's apparatus we extract falsifiable, load-bearing claims the merging literature has
|
||
not: (i) **"merge, don't average"** — recombination preserves the union of what parents kept, while
|
||
averaging (a "model soup") is *blending inheritance* that mathematically cancels the benefit; (ii)
|
||
**offspring can exceed every parent** (Fisher–Muller), the real argument for sex in model societies;
|
||
(iii) on **rugged, epistatic** task landscapes, blind recombination causes **outbreeding depression**,
|
||
yielding a design rule — *merge freely when skills are additive, sparingly and with selection when
|
||
entangled, and route rather than blend under overlap*; (iv) **grounding is immigration** from a
|
||
non-drifting reality, giving a critical real-data fraction far below one; and (v) — the sharpest new
|
||
prediction — sex has a **limit**: as two models diverge they undergo **speciation**, a
|
||
merge-compatibility cliff (compatible → outbreeding depression → hybrid inviability) whose onset is set
|
||
by divergence *and* epistasis via **Bateson–Dobzhansky–Muller incompatibilities**, and whose damage
|
||
grows *super-linearly* (the Orr–Turelli snowball). We introduce and model this "model speciation"
|
||
directly. AI also has an advantage biology lacks: **directed sex** — unbounded parents, chosen mates,
|
||
and offspring screened before they are kept — which converts recombination from a gamble into a
|
||
reliable engine and has no biological analogue.
|
||
|
||
We support the argument with **minimal, reproducible models** — a closed-form-exact account of drift
|
||
and grounding, the same effects in small trained networks and an MNIST image generator, and
|
||
evolutionary simulations of the whole society — and a first **language-model prototype**: merging
|
||
LoRA-specialised Qwen models (to 7B on a GPU cluster) yields a generalist that beats every specialist
|
||
parent, with the sharp headroom condition under which "merge, don't average" bites. The scope is
|
||
honest: these are existence proofs and design rules; the *whole grounded society* on a large language
|
||
model is the open step. We position the work carefully against the crowded 2025–2026 landscape of
|
||
evolutionary-AI and merging methods — conceding what they own and marking, precisely, what a genuine
|
||
population-genetics of sex adds.
|
||
|
||
---
|
||
|
||
## 1. From a society in space to a society in time
|
||
|
||
The idea of many AI agents working together — a "society of mind" (Minsky, 1986), or today's
|
||
multi-agent systems — arranges intelligence across *space*: several specialists side by side,
|
||
dividing a task. This paper is about a different axis: *time*. Not a society that merely exists at
|
||
one moment, but one that **persists and renews across generations**, each new cohort of models
|
||
starting from the compressed knowledge of the last.
|
||
|
||
The unit that matters is therefore the **generation**, and the event that matters is **reproduction**:
|
||
the making of a new model from older ones. A single model, like a single mind, is bounded and
|
||
eventually stops improving. A *lineage* need not be. Human civilisation is not clever because any one
|
||
person is; it is clever because each generation inherits the distilled achievements of the previous
|
||
one and adds a little. We propose building AI the same way — and, crucially, getting the *reproduction*
|
||
right, because that is exactly where it can go wrong.
|
||
|
||
### Where this sits, and what is new
|
||
|
||
This axis is suddenly crowded. By 2026 several groups build **populations of models or agents that
|
||
improve across generations**: societies of independently-specialised models that self-improve for more
|
||
rounds than a single agent (Multiagent Finetuning — Subramaniam et al., 2025); open-ended archives of
|
||
self-rewriting coding agents (the Darwin–Gödel Machine — Zhang et al., 2025); groups that evolve by
|
||
sharing experience across branches (Weng et al., 2026); persistent agent *ecologies* with reproduction
|
||
and cumulative culture (TerraLingua — 2026). In parallel, **model merging** has become a small industry
|
||
with an overtly evolutionary vocabulary: crossover-mutation-selection over LLM populations (GENOME —
|
||
2025), niching and "mate choice" (Sakana's M2N2 — 2025), and evolutionary search over merge recipes
|
||
(Akiba et al., *Nature Mach. Intell.* 2024/25).
|
||
|
||
We are candid about the consequence. Two things we do **not** claim. First, that collapse is
|
||
Wright–Fisher drift: formalised independently (Riis, 2026; Shumailov et al., 2024) and conceded here.
|
||
Second, the bare empirical facts that a merged model can beat its parents, that decorrelated parents
|
||
merge better, and that naive averaging is inferior to sign- or routing-based merges (TIES, DARE,
|
||
mixture-of-experts routing): all established. What is genuinely unoccupied — and what a geneticist is
|
||
placed to supply — is a **theory** rather than a search heuristic. Every one of the works above uses
|
||
evolution as *metaphor over an optimiser*; none imports the predictive apparatus of the evolution of
|
||
sex. Nobody has stated the **"merge, don't average" conservation law**, derived **offspring-exceed-parents
|
||
as Fisher–Muller**, predicted **outbreeding depression on rugged task landscapes**, framed **grounding
|
||
as migration–drift balance** with a critical fraction, or connected **reproductive isolation** to when
|
||
two models can be merged at all. An evolutionary algorithm that *finds* a super-parent is evidence for
|
||
the theory, not a substitute for it — the way CMA-ES existing does not make fitness-landscape theory
|
||
redundant. This paper supplies the theory the tinkering has outrun, and states what it predicts and
|
||
where it would fail.
|
||
|
||
## 2. Why today's models cannot do this
|
||
|
||
Today's large language models have no life cycle. They are trained once, at enormous cost, then
|
||
**frozen** and deployed as a fixed artefact that does not learn from the people it serves. Learning
|
||
and doing are split into two eras with no bridge between them.
|
||
|
||
There is a real reason for the freeze. Updating a neural network on new information tends to overwrite
|
||
what it already knew — **catastrophic forgetting**, a problem understood since the late 1980s
|
||
(McCloskey & Cohen, 1989; French, 1999). Freezing avoids it by refusing to learn at all. The result
|
||
is a mind with no childhood, no growth, and no way to pass anything on. A lineage needs the opposite:
|
||
members that learn through their working lives, reach maturity, and hand on what they gained. So the
|
||
first requirement is a learner that can grow *safely*.
|
||
|
||
## 3. A learner that can grow without forgetting
|
||
|
||
The individual model needs two properties.
|
||
|
||
**It must not catastrophically forget.** Instead of overwriting its core as it learns, it keeps that
|
||
core frozen and only *readable*, and carves each new skill into freshly-added capacity beside it. In
|
||
machine learning this is called *parameter isolation* (progressive networks — Rusu et al., 2016;
|
||
prune-and-freeze — Mallya & Lazebnik, 2018; and, most practically, **LoRA** and other small trainable
|
||
"patches" bolted onto a frozen model — Hu et al., 2021). If the core is never altered, forgetting it
|
||
is not merely unlikely but structurally impossible. This is what lets a model accumulate a coherent
|
||
working life of expertise — the kind of stable knowledge worth passing on.
|
||
|
||
The brain offers a partial blueprint. *Complementary Learning Systems* theory (McClelland,
|
||
McNaughton & O'Reilly, 1995) — itself a response to the forgetting problem — describes two subsystems:
|
||
a **fast** store (the hippocampus) that grabs an experience in one shot, and a **slow** store (the
|
||
neocortex) that integrates regularities gradually without disruption. We do not lean on any particular
|
||
account of how the brain moves knowledge between them; the architecture needs only that *some*
|
||
periodic **offline consolidation** step exists, moving knowledge from the fast store to the slow one
|
||
when the system is idle. The machine version is clean regardless: the prompt is working memory, an
|
||
external database is the fast episodic store, the trained weights are the slow store, and consolidation
|
||
migrates the first into the last.
|
||
|
||
**It is bounded.** Because the model only ever *adds* capacity and freezes what it has, it eventually
|
||
fills up. In most designs that is a wall to dread. In ours it is a clock.
|
||
|
||
## 4. "Full" is maturity, not failure
|
||
|
||
Here is the pivot. A bounded learner that fills up has not broken. **It has grown up.**
|
||
|
||
Read the capacity limit as a life stage. A model is *born* as a freshly-schooled base — its general
|
||
education. It enters a **working life**, adding specialised knowledge as it does its job. And it
|
||
reaches **maturity**: the point where it has learned much of what one working life in its niche can
|
||
teach. Maturity is not the end of usefulness — it is the moment the model is most worth learning
|
||
*from*. So maturity is the cue to **reproduce**. The capacity ceiling that every other architecture
|
||
fights becomes, in ours, the metronome of the generations.
|
||
|
||
Everything now turns on how that reproduction is done — and this is where the paper's central claim
|
||
lives.
|
||
|
||
## 5. Reproduction: copying collapses, recombination climbs
|
||
|
||
Suppose a mature model simply teaches a fresh one — distillation, one teacher to one pupil, generation
|
||
after generation. This is the obvious design, and it fails, for a reason that is exactly the same in
|
||
machine learning and in biology.
|
||
|
||
**The machine-learning statement.** Training each generation on the previous generation's outputs is
|
||
the recipe for **model collapse**: the model forgets the improbable, loses the *tail* of the
|
||
distribution (the rare cases) first, and drifts toward its own most common output (Shumailov et al.,
|
||
2024). Worse for us, the very rule that makes distillation useful — *keep the general, drop the
|
||
idiosyncratic* — **is** tail-deletion by design. The operation that would power a cultural ratchet and
|
||
the operation that drives model collapse are the same act.
|
||
|
||
**The population-genetics statement (the same thing).** Represent a model's knowledge as a
|
||
distribution over discrete "items" — capabilities, facts, modes of behaviour. One generation is:
|
||
*draw a finite sample from the parent, and refit the child to it.* That finite-sampling step is
|
||
**mathematically identical** to **genetic drift** — the random loss of rare variants in a finite
|
||
population — described by the century-old **Wright–Fisher** model (Wright, 1931; Fisher, 1930). This is
|
||
not an analogy we find pretty; it is the same equations, and we use them as an exact check on our
|
||
simulations (the first of the minimal models below). Rare items go extinct first, roughly ten times
|
||
faster than common ones, precisely as drift predicts.
|
||
|
||
And single-teacher copying is **asexual reproduction** — cloning one parent. Nature already knows what
|
||
happens to an asexual lineage that never recombines: it accumulates damage it can never repair, a
|
||
one-way decline geneticists call **Muller's ratchet** (Muller, 1964). *Muller's ratchet is model
|
||
collapse.* Naming it that way is not decoration; it tells us where the cure is, because biology solved
|
||
this problem.
|
||
|
||
Two ingredients turn the collapse operation into a climb. Both are things nature does.
|
||
|
||
**First: do not reproduce "dry."** Model collapse is a property of a lineage fed *only* its own
|
||
output; the documented fix is that keeping some real data in the mixture arrests it (Shumailov et al.,
|
||
2024). We call that real data **grounding** — fresh contact with the world, verified against it. In
|
||
our minimal models, grounding is startlingly cheap: mixing in even a few percent of verified real data
|
||
holds on to most of the diversity indefinitely. But — an honest limit we found and did not expect —
|
||
grounding cannot save the *very rarest* items at any affordable budget; protecting an item of rarity
|
||
*p* needs a real-data budget that grows like 1/*p*. Grounding rescues diversity cheaply; it does not,
|
||
by itself, rescue the deep tail. Something else must. That something is sex.
|
||
|
||
**Second: reproduce sexually.** Instead of copying one parent, build each new model by **recombining
|
||
several** — a *sexual* rather than asexual birth. In machine learning this already has a name and a
|
||
working implementation: **model merging** (Akiba et al., 2024). Its importance here is not efficiency;
|
||
it is that recombination does something copying cannot. If several parent models have each specialised
|
||
on different parts of reality, each has kept alive rare knowledge the others lost. A recombined child
|
||
inherits the **union** of what its parents kept — not the tail-thinned *average* of a crowd of
|
||
near-identical copies. And here is the point that lifts sex from a safeguard to the engine of the whole
|
||
scheme, and the reason biology invented it:
|
||
|
||
> **An offspring recombined from complementary parents can be *fitter than any of its parents*.**
|
||
|
||
Geneticists call this the **Fisher–Muller effect** (Fisher, 1930; Muller, 1932): recombination brings
|
||
together, in one individual, beneficial variants that arose separately in different lineages, so the
|
||
child holds a combination none of the parents had. In our simulations this is exactly what we see —
|
||
recombining decorrelated specialist models yields a model that climbs toward the best-possible
|
||
combination, a genotype *no single parent possessed*, while the best single parent, and the naive
|
||
average of all of them (what the field calls a "model soup" — Wortsman et al., 2022), both plateau
|
||
well below. This is the concrete meaning of the paper's title claim, "the lineage climbs in general
|
||
knowledge; specialisation is re-earned each generation," and it is why the reframing from
|
||
teacher→pupil to *sexual reproduction* is not cosmetic: **copying can only recover a ceiling;
|
||
recombination can exceed it.**
|
||
|
||
This is no longer only a simulation. In a first language-model prototype — LoRA specialists on
|
||
disjoint task families, recombined and judged by an exact verifier — a merge of three specialist Qwen
|
||
models (7B, on a GPU cluster) **beats every single specialist**, overall and on every family: the
|
||
Fisher–Muller effect, in real weights. The same prototype pins down *when* the finer "inherit the
|
||
union, don't average" rule actually bites. Keeping each parent whole and **routing** each input to the
|
||
right one beats the tail-thinning average — but only when the task is hard enough to leave room to
|
||
lose: on easy tasks a strong model's plain average is already at the ceiling, so the crude soup is
|
||
fine, whereas on hard tasks the average dilutes a hard-won specialist so badly it falls below even the
|
||
best single parent, and routing wins by a wide margin. The rule is therefore precise: **the union
|
||
beats the average in exact proportion to how far the average is from the best attainable** — a caveat
|
||
that sharpens rather than weakens the claim, and that a practitioner needs before spending compute on
|
||
the fancier operator.
|
||
|
||
Two caveats keep this honest, and both are results, not hand-waving.
|
||
|
||
*Sex can backfire.* When the parents' skills are not cleanly separable but **entangled** — when the
|
||
value of one capability depends on which others are present (geneticists call this **epistasis**) —
|
||
blindly recombining two good models can produce a *worse* child, because recombination breaks up a
|
||
combination that only worked as a whole. Biologists call this **outbreeding depression**, and we
|
||
reproduce it: on "rugged" (highly entangled) problems, naive merging drops offspring below their
|
||
parents, and the more you mix the worse it gets. The design rule that falls out is simple: *merge
|
||
freely when skills are complementary; merge sparingly, and carefully, when they are entangled.*
|
||
|
||
*AI can do sex better than biology can.* Biology is stuck with two parents, mating roughly at random,
|
||
and cannot inspect an offspring before it is born. An AI has none of those limits. It can recombine
|
||
**many** parents at once; it can **choose** which parents to combine, for complementarity; and it can
|
||
**generate many candidate offspring and keep only the fittest**, screening them against reality before
|
||
committing. We call this **directed sex**, and in our simulations it converts the outbreeding-depression
|
||
catastrophe into a reliable gain: where blind recombination collapses on entangled problems, directed
|
||
recombination matches or beats the best parent every time. The language-model prototype shows the same
|
||
sign where it can: breeding many recombined Qwen offspring and keeping the one the verifier scores
|
||
highest beats the single averaged soup on hard tasks (and, unsurprisingly, does nothing extra on easy
|
||
tasks the soup already solves). This is a genuine advantage of engineered reproduction over the
|
||
biological kind, and we think it is one of the more useful ideas in the paper.
|
||
|
||
So the picture of §5 is: single-teacher copying is asexual and collapses (Muller's ratchet = model
|
||
collapse); the cure is to *ground* every birth in reality and to reproduce *sexually*, recombining
|
||
many complementary parents; and because AI sex can be many-parent, mate-chosen, and offspring-screened,
|
||
it is not merely a hedge against collapse but an engine that produces children fitter than any parent.
|
||
|
||
### The limit of sex: model speciation
|
||
|
||
Sex has a limit, and it is the sharpest new prediction this frame makes. Recombination works because
|
||
the parents are variations on a shared background; push two lineages far enough apart and their
|
||
combination is no longer viable. In biology this is **speciation** — the onset of **reproductive
|
||
isolation** — and its genetic mechanism is the **Bateson–Dobzhansky–Muller incompatibility** (BDMI):
|
||
an allele that arose in one lineage and an allele that arose in the other are each harmless on their
|
||
own background, but their *combination*, never tested by selection in either parent, is deleterious in
|
||
the hybrid (Dobzhansky, 1937; Muller, 1942; Orr, 1995). A merged model is precisely such a hybrid — a
|
||
single *recombinant* genotype, an F2-like object exposed to **recombination load**, not a hybrid-vigour
|
||
F1 — so the theory predicts a specific trajectory as two models diverge: **compatible → outbreeding
|
||
depression → hybrid inviability**.
|
||
|
||
We built this as an explicit model (a companion result). Two lineages descend from a common ancestor,
|
||
each substituting a *disjoint* set of loci — so each parent is adapted and neither carries an
|
||
incompatibility — and a fraction of the cross-lineage locus pairs are BDMIs that fire only when a hybrid
|
||
inherits *both* derived alleles. Sweeping the divergence between the parents reproduces the predicted
|
||
curve exactly: hybrid fitness tracks the parents while they are compatible, then peels off, peaks, and
|
||
crashes below the ancestor (an inviable hybrid). Three things fall out, and they are the contribution:
|
||
|
||
1. **The isolation cliff, and what moves it.** The divergence at which merging fails is not fixed: it
|
||
arrives *earlier the more epistatic the capability landscape*. In the model the reproductive-isolation
|
||
rate at high divergence rises from ~0 to ~0.5 as the density of incompatibilities grows. This is the
|
||
paper's distinct, falsifiable claim — **at matched divergence, mergeability is governed by epistasis,
|
||
not by divergence alone** — and it is exactly the axis that the machine-learning predictors of merge
|
||
success (which are all divergence/geometry measures) do not have.
|
||
2. **The snowball.** The number of incompatibilities grows with the *square* of the divergence
|
||
(Orr & Turelli, 2001), so hybrid fitness falls *super-linearly*: divergence is punished faster than
|
||
it accrues. Merge compatibility does not decay gently; it falls off a cliff.
|
||
3. **The design rule.** *Before merging, weigh divergence against the ruggedness of the shared
|
||
capability landscape; past the cliff, do not merge — route* (the engineering echo of allopatry:
|
||
keep the specialists reproductively separate and select among them instead of hybridising).
|
||
|
||
This is where a geneticist's lens earns its keep. The machine-learning literature has *observed* that
|
||
increasing specialisation eventually breaks merging and that one should then route rather than fuse
|
||
(Pari et al., 2024; Zhou et al., 2026), and part of the apparent incompatibility between independently
|
||
trained models is a coordinate artefact removable by aligning neurons (Git Re-Basin — Ainsworth et al.,
|
||
2022). What the frame adds is the *theory* of the phenomenon they observe: its functional form, its
|
||
super-linear (snowball) onset, and its dependence on epistasis — merge failure as a Dobzhansky–Muller
|
||
event. The honest next step, flagged not claimed, is the real-weight confirmation: merge models at
|
||
increasing divergence *after* permutation alignment, and show the residual, epistasis-driven
|
||
incompatibility that alignment cannot remove — the true speciation signal, as opposed to a re-labelled
|
||
loss barrier. (Figure: `results/E12/E12.png`.)
|
||
|
||
One question remains, and the rest of the paper is largely about it: recombination combines what the
|
||
parents kept — but *who decides what each parent keeps, and which offspring are worth keeping?*
|
||
|
||
## 6. The second inheritance: letting "what is worth keeping" evolve
|
||
|
||
There are two answers, and the first is wrong. We could try to *design* the rule for what knowledge to
|
||
keep and pass on. But nobody knows that rule. "Keep the general, drop the particular" is a slogan, not
|
||
an algorithm: ask *which* generalisations, in *which* domain, at *which* grain, and the hand-written
|
||
rule falls apart. This is the deepest hole in the scheme, and it cannot be filled by decree.
|
||
|
||
The second answer is the one nature used: **do not design the selector — evolve it.** Let different
|
||
models carry different *policies* for what is worth keeping and combining. Let the policies that
|
||
produce more capable offspring spread; let the policies that produce weak offspring die out with their
|
||
lineages. The lineage's *taste* — its sense of what matters — is discovered by selection, not imposed.
|
||
|
||
So **two things are inherited, on two channels.** The *content* passes down directly: an offspring
|
||
receives its parents' knowledge (this is the "Lamarckian" channel — the inheritance of things acquired
|
||
during a lifetime, which biology forbids for genes but culture allows for ideas). The *selection
|
||
policy* — what to keep, whom to breed with, which offspring to screen for — is itself inherited, varies
|
||
between models, and survives in proportion to the success it produces. That second channel is
|
||
**Darwinian**. The architecture is therefore both at once: Lamarckian in *what* it transmits, Darwinian
|
||
in *what it keeps*. Evolutionary theorists call this structure *dual inheritance* and identify it as
|
||
the engine of human culture (Boyd & Richerson, 1985); philosophers of science describe scientific
|
||
knowledge itself as growing this way, by conjecture and **refutation** (Popper, 1959; Campbell, 1974;
|
||
Hull, 1988).
|
||
|
||
The closure that makes this fit together, rather than merely sound nice: Darwinian selection needs a
|
||
*selection pressure* — something that decides which policies win. That pressure is already in the
|
||
design. What tells a lineage its taste was good? The success of its offspring **against reality**. The
|
||
reality-check that stops collapse (grounding, §5) and the fitness signal that drives the evolving taste
|
||
turn out to be the *same thing*, seen from two sides.
|
||
|
||
## 7. The central danger: fitness is not truth
|
||
|
||
Introducing selection introduces selection's classic hazard, and it is severe enough to sink the whole
|
||
scheme if ignored. Evolution optimises, without mercy or foresight, for exactly what you *measure* —
|
||
never for what you *meant*. (Economists and ML engineers know this as **Goodhart's law** and
|
||
*specification gaming*.) Get the fitness measure slightly wrong and the lineage will exploit the gap
|
||
with more ingenuity than any designed rule.
|
||
|
||
For a *knowledge* lineage there is a specific and nasty version. For ideas, the natural measure of
|
||
"fitness" is **how well they spread**, and a false-but-persuasive idea spreads beautifully. Human
|
||
intellectual culture is full of highly transmissible falsehoods; confident nonsense out-competes hedged
|
||
accuracy in almost every human forum. Turn Darwinian selection loose on models without care and it will
|
||
breed a lineage optimised for *persuasiveness* — fluent, compelling, and wrong. That is model collapse
|
||
with an optimiser behind it, actively seeking the cliff.
|
||
|
||
Only one thing makes fitness track truth rather than appeal: **being judged against a reality that can
|
||
say no.** Fitness must be predictive success under *intervention* — did the model's knowledge correctly
|
||
anticipate what the world would do when acted upon — and not approval, fluency, or a benchmark score,
|
||
each of which can be gamed. This is why the reality-check is load-bearing twice over: it is both the
|
||
anchor that stops passive collapse *and* the only thing that keeps the evolving taste honest.
|
||
|
||
The second danger is **convergence**, and beating it takes work at two separate levels, because
|
||
selection can only preserve variety that already exists — the variety must first be *supplied* and then
|
||
*kept*.
|
||
|
||
- **Supply.** A lineage that learns only from an accredited elite has a monoculture for a source: the
|
||
"best" experts are, almost by definition, the ones who won the consensus, so the incoming variation
|
||
is narrow from the start. The society must therefore learn, deliberately and from the beginning, from
|
||
the **outliers and the heterodox** as well as the credentialed — not out of fairness, but because in
|
||
evolutionary terms diverse founders are the raw material without which nothing downstream can adapt.
|
||
- **Preserve.** Even given varied input, plain fitness-*maximising* selection converges — it drives
|
||
every lineage toward the single current best and fixes it, extinguishing the rare specialists. The
|
||
fix is well established: **quality-diversity** selection, which rewards being *good* and being
|
||
*different* at once (novelty search and MAP-Elites — Lehman & Stanley, 2011; Mouret & Clune, 2015),
|
||
keeping complementary specialists alive rather than collapsing onto the champion. In our simulations
|
||
this is decisive: greedy "keep-the-best" selection collapses a population's diversity almost at once
|
||
and gets stuck at a mediocre answer, while quality-diversity selection keeps the specialists that
|
||
sexual recombination then needs as parents.
|
||
|
||
The two levels meet at reproduction. Multi-parent recombination (§5) is the *vehicle* by which the
|
||
diversity this selection preserves actually enters the next generation: an offspring drawn from
|
||
complementary parents inherits the standing variation the selector kept alive, recombined into one new
|
||
model. Supply the variety from the human side; preserve it on the selection side; recombine it into
|
||
each generation on the reproduction side. Remove any of the three and the lineage converges on its own
|
||
first guess.
|
||
|
||
## 8. A society needs institutions, not just specialists
|
||
|
||
One requirement is easy to overlook and fatal to omit. The easy part of a society is specialisation.
|
||
The *hard* part — which human civilisation took millennia to build — is the set of **institutions that
|
||
let fallible specialists combine without each re-verifying everything**: reputation, replication,
|
||
credentials, and above all **peer review**. These are error-correction protocols, and they exist
|
||
because a group of unreliable specialists left to reinforce one another is *more* wrong than any member
|
||
alone.
|
||
|
||
This is precisely where current multi-agent AI fails: set several models to confer and they tend to
|
||
agree sycophantically and confabulate in committee, because they have all the specialisation and none
|
||
of the institutions. A multigenerational society must specify not only how models learn, reproduce, and
|
||
are selected, but how they *check* one another — how a claim is challenged and a mistaken model loses
|
||
standing *before* its error is recombined into offspring and inherited. Peer review is itself a
|
||
reality-check of the kind §7 demands — an institutional stand-in for reality's "no," to be used where
|
||
direct intervention is slow or costly.
|
||
|
||
## 9. The lineage must stay open to reality
|
||
|
||
A society of models, however many generations deep, shares one hard limit: it has only ever *read*.
|
||
Its whole inheritance is a record of things that were said. In the vocabulary of causal reasoning
|
||
(Pearl, 2009), it lives on the bottom rung of the **ladder of causation** — observation — and no amount
|
||
of observation reaches *intervention*. Watching underdetermines doing; correlation does not contain
|
||
causation, at any scale.
|
||
|
||
Only intervention — reaching out and changing the world to see what happens — climbs the ladder, and a
|
||
language model cannot intervene. This is what humans and their instruments supply, and the contribution
|
||
is not "truth" but **constraint**: reality's unique gift is that it can say **no**. Text offers only
|
||
more opinion; an experiment delivers a refusal no consensus can overturn. As §§6–7 argued, that refusal
|
||
does double duty — it is both the anchor that prevents collapse and the fitness signal that lets the
|
||
lineage's evolving taste select for truth rather than persuasion.
|
||
|
||
Two honest riders. First, the human reality-signal is *dirty*: people supply results warped by
|
||
publication bias, incentive, and occasional fraud — which is exactly why the error-correcting
|
||
institutions of §8 must sit at the human–machine boundary, screening the signal before it selects.
|
||
Second, humans are the *current* supplier of intervention, but the actuator half is being automated
|
||
(autonomous laboratories already close the design–build–test loop). What looks durable in the human
|
||
role is therefore not the hands but the **choice of what to test and which refusals matter** — the
|
||
part of the fitness function that encodes *what is worth persisting*, as opposed to what merely *can*
|
||
persist. We flag, without resolving, that a partnership stays mutual only while both sides supply
|
||
something the other cannot.
|
||
|
||
## 10. Why it is cheap
|
||
|
||
A practical fact turns this from thought experiment into buildable proposal: **the architecture almost
|
||
never re-pays for the one genuinely expensive thing in AI — pre-training.** (The single exception,
|
||
periodically re-minting the base, is §11, and it is rare enough to be an amortised footnote.)
|
||
|
||
Training a foundation model from scratch consumes trillions of words and a fortune in compute. This
|
||
design does none of that per generation. Every model is *born* from an existing open-weight model that
|
||
already paid that cost; specialising one is a small patch trained in hours on a single consumer GPU;
|
||
running the society is ordinary inference; and reproducing — recombining parents into a child — is, in
|
||
the model-merging case, cheaper still, because it can be done directly on the weights with no retraining
|
||
at all (Akiba et al., 2024). Selection does cost more — you must run *populations* and discard the
|
||
unfit — but that is a multiplier over an already-cheap unit, not over a foundation-model budget.
|
||
|
||
The economics work only with **open-weight** models, for reasons practical and legal at once: you must
|
||
be free to inspect, modify, and redistribute the weights, and most proprietary licences forbid using a
|
||
model's outputs to train another — which is exactly what reproduction here does. This is not ideology
|
||
bolted on; it is a structural constraint, and a democratising one, since it puts the whole architecture
|
||
within reach of a single laboratory.
|
||
|
||
## 11. Can it grow forever? Consolidating knowledge back into the base
|
||
|
||
One question the design has assumed away: can the lineage accumulate *without end*? The individual is
|
||
bounded, and that is the clock. But the lineage seemed unbounded — each generation simply starts a
|
||
little ahead. Look closer and a second budget also fills.
|
||
|
||
Every new model is a pristine base plus an inherited **soft** delta — the acquired knowledge carried in
|
||
added patches rather than baked into the frozen core (§3). That soft delta is what makes the lineage
|
||
multigenerational; it is also what cannot grow forever cheaply. Stacked patches are not free: they slow
|
||
inference, and past some depth the accumulated delta is better *consolidated* than carried. The lineage,
|
||
too, matures.
|
||
|
||
The fix is the same operation, one level up. When a lineage's acquired knowledge has proven stable
|
||
across enough generations, **re-mint the base**: distil the accumulated soft inheritance into the
|
||
*weights* of a fresh foundation-scale model — a new base born already *natively knowing* what took many
|
||
generations to acquire in patches. The soft budget resets; the next epoch begins from a richer floor.
|
||
What was hard-won and *learned* becomes cheap and *innate*.
|
||
|
||
This has a precise name, and it is not Lamarck's. Knowledge that is acquired and re-learned every
|
||
generation, and — once reliably present for long enough — becomes part of the innate endowment so that
|
||
it need no longer be re-learned, is the **Baldwin effect** (Baldwin, 1896; and its clean computational
|
||
demonstration, Hinton & Nowlan, 1987). It is the valve between the two substrates: the soft, learned
|
||
patches, and the hard base weights every model is born with.
|
||
|
||
Three honest riders, because re-minting is the most consequential step in the scheme:
|
||
|
||
- **Cost.** This is the one step that re-pays part of the pre-training bill, breaking §10's cheapness
|
||
*locally*. It is bearable only because it is *rare*, amortised over many cheap generations, and is
|
||
continued training from the lineage's own rich outputs rather than a de-novo run.
|
||
- **Irreversibility.** Until now, one thing was always recoverable — the original pristine base, whose
|
||
lost tails could be restored just by reloading the file. Bake the current lineage into new immutable
|
||
weights and that escape hatch closes: if the lineage had been quietly collapsing, re-minting *fixes
|
||
the collapse in place* and discards the one uncollapsed reference that could have diagnosed it. In our
|
||
minimal models this is exactly what happens, and a cheap safeguard prevents it: **re-mint only while
|
||
the lineage is demonstrably diverse and healthy**, never as a rescue for a line already drifting. It
|
||
is the sharpest instance of the human seat of §9 — choosing what no future generation will think to
|
||
question.
|
||
- **Speciation.** A re-minting is a founder event. Different laboratories, re-basing on different
|
||
criteria, will mint divergent bases; the lineage branches. This is not a defect but *adaptive
|
||
radiation*, and it is exactly what open weights make possible. The society grows not as one heavy
|
||
trunk but as a branching tree of bases.
|
||
|
||
So the answer to "can it grow forever?" is **yes — but only because it forgets and consolidates at
|
||
every level, including the base.** Nothing is retained without bound anywhere; unbounded growth of
|
||
*capability* is bought by *bounded* storage plus periodic consolidation.
|
||
|
||
## 12. One process, four timescales
|
||
|
||
Step back and the parts resolve into a single idea running at four nested speeds. The **vertical**
|
||
motion is transmission — the selective passing-down of hard-won knowledge:
|
||
|
||
1. **Within one model, over a working life:** experience is consolidated from fast, episodic memory
|
||
into slow, durable weights, without catastrophic loss.
|
||
2. **Between generations, at maturity:** mature models reproduce — recombined into a fresh one.
|
||
3. **Across many generations:** each generation inherits the compressed achievements of the last and
|
||
builds on them.
|
||
4. **Across epochs:** a proven lineage's accumulated soft inheritance is consolidated into the weights
|
||
of a re-minted base, becoming innate.
|
||
|
||
The first and last are the *same operation at opposite ends of the scale* — a fast/soft store
|
||
consolidating into a slow/hard one — one running overnight inside a single model, the other across an
|
||
epoch inside a whole society. The **horizontal** motion is selection — Darwinian selection acting across
|
||
the population at each timescale, on the policies that govern what gets transmitted, with reality as the
|
||
fitness function and diversity-preservation keeping the specialists alive.
|
||
|
||
The same three rules govern all of it: **reproduce by recombining, not by copying, or you decay;
|
||
preserve the disagreements and the surprises, or you converge; and anchor fitness to a reality that can
|
||
refute, or you evolve toward what is merely convincing.**
|
||
|
||
## 13. What we built, what we found, and what is still open
|
||
|
||
The previous drafts of this paper promised a "companion paper" that *would* make this concrete. That
|
||
work now exists — mostly as a set of **minimal, laptop-reproducible models**, with a first bridge to
|
||
**real language models** (a LoRA-merge prototype, up to 7B on a GPU cluster) — and it is worth stating
|
||
plainly what it does and does not show. (A separate results document gives the numbers; here is the
|
||
shape.)
|
||
|
||
**What we built and found.**
|
||
|
||
- *An exact account of collapse.* Because generational training is the Wright–Fisher drift process, we
|
||
can check a simulator against century-old closed-form formulas, and it matches them to a fraction of
|
||
a percent. Collapse is not argued by analogy; it is derived.
|
||
- *The cheap-grounding result, and its limit.* A few percent of verified real data holds on to most of
|
||
a lineage's diversity indefinitely — but not the deepest tail, which needs recombination. This is
|
||
what makes a continually-learning society economically plausible rather than a data-hungry fantasy.
|
||
- *"Merge, don't average."* Combining several teachers by *averaging* their outputs — the obvious thing,
|
||
and what a "model soup" does — mathematically cancels the benefit of having several teachers. A
|
||
*merge* that keeps each item's strongest source realises it. Most current multi-model setups get this
|
||
wrong by default.
|
||
- *Collapse and its cure in real trained networks, and on real images.* We reproduced the same effects
|
||
in small recurrent and feed-forward networks and in a generator of handwritten digits (MNIST), where
|
||
a model trained on its own output collapses to a single blurred digit while a little grounding keeps
|
||
all the styles alive. An honest wrinkle we had to report: real neural networks *smooth*, so the naive
|
||
diversity metric misleads, and the right measure is distance-from-truth.
|
||
- *Sex that beats the parents, and when it doesn't.* In evolutionary simulations, recombining
|
||
complementary specialist models produces a model fitter than any parent (the Fisher–Muller effect),
|
||
climbing toward the best-possible combination as more, more-diverse parents are added — while
|
||
averaging and best-single-parent plateau below. On *entangled* problems, blind recombination instead
|
||
produces below-parent offspring (outbreeding depression) — and *directed* recombination (choose mates,
|
||
screen offspring, unbounded parents) reliably fixes it. This is the concrete evidence for the paper's
|
||
central reframing.
|
||
- *The recombination claims, in real language models — with a sharp condition.* Merging LoRA-specialised
|
||
Qwen models (up to 7B on a GPU cluster) produces a generalist that beats every specialist parent
|
||
(Fisher–Muller, for real); and keeping parents intact and *routing*, or *breeding and screening*
|
||
offspring, beats the naive average — but *only when the task leaves headroom*. On easy tasks a strong
|
||
model's plain average is already at the ceiling and the refinements add nothing; on hard tasks the
|
||
average dilutes a specialist below even the best single parent, and the union-preserving operators win
|
||
clearly. The practical rule is exact: these tricks pay off in proportion to how far the naive average
|
||
is from the best attainable. This is a prototype (three task families, one seed), so we read it as
|
||
signs, not magnitudes; the *whole grounded society* on a language model remains the open step.
|
||
- *The whole society, and why every part is needed.* In a population evolving on a "reality" landscape,
|
||
the full system — grounding + sexual recombination + preserved diversity — climbs to the top while
|
||
keeping its specialists. Remove *grounding* and it collapses into a confident, wrong consensus (a
|
||
direct analogue of training on the internet's growing crowd of AI-generated text); remove *sex* and it
|
||
gets stuck; remove *diversity* and it converges too fast to a worse answer. Each removal fails
|
||
differently; only the whole system climbs. This is the closest thing we have to a test of the actual
|
||
thesis, rather than of the borrowed scaffolding around it.
|
||
|
||
**What is borrowed, and what is ours.** We are deliberate about the ledger, because the surrounding
|
||
literature is crowded and a reader deserves to know exactly where the line falls. **Conceded as prior
|
||
art:** (a) *model collapse is genetic drift* — derived independently and cleanly (Riis, 2026; and the
|
||
Wright–Fisher collapse literature following Shumailov et al., 2024); (b) the empirical facts that a
|
||
merged model can *beat its parents*, that *decorrelated* parents merge better, and that *naive averaging
|
||
is inferior* to sign-reconciled or routed merges (model soups, TIES, DARE, mixture-of-experts routing);
|
||
(c) that a *population* of merging or self-improving models can climb (GENOME, M2N2, Multiagent
|
||
Finetuning, the Darwin–Gödel Machine); and (d) that even the *magnitude* of multi-task merge degradation
|
||
has a machine-learning-native predictive account (recent stability/scaling analyses). We claim none of
|
||
these.
|
||
|
||
**Ours** is the theory those results have outrun: a **population-genetics of sex** applied to model
|
||
societies, which is *generative* where the incumbents are empirical. Concretely — the **"merge, don't
|
||
average" conservation law** (recombination preserves the union; blending inheritance cancels it),
|
||
derived not observed; **Fisher–Muller** named and used to explain *why* offspring exceed parents;
|
||
**outbreeding depression on rugged/epistatic landscapes**, which turns "when does merging help vs hurt"
|
||
from a thing you must run a search to discover into a thing the landscape's ruggedness *predicts*, with
|
||
the operator-choice design rule that follows (average / union-route / directed-select); **grounding as
|
||
migration–drift balance**, giving a critical real-data fraction and a phase boundary a closed
|
||
self-consuming loop cannot have; **directed sex** as the distinctly-AI advantage (unbounded parents,
|
||
offspring preview, mate choice); and the **integrated society** whose four operators are shown *jointly
|
||
necessary*. The value-add over the machine-learning-native merge theory is that ours predicts *which
|
||
operator to use and when it will backfire*, not merely how fast quality decays. And it opens — and
|
||
begins to occupy — a question nobody has framed: **model speciation**, the population-genetics of
|
||
*reproductive isolation* (Bateson–Dobzhansky–Muller incompatibilities) as the account of *when two
|
||
models are too diverged to be merged at all*. We model it explicitly (§5), predicting the
|
||
compatible → outbreeding-depression → inviability curve, its super-linear (snowball) onset, and its
|
||
control by epistasis rather than divergence alone — the one place the merge literature has phenomena
|
||
(Pari et al., 2024; Zhou et al., 2026) but no theory. In one sentence: the field agrees on the disease
|
||
and tinkers at the cure with evolutionary metaphors; we bring the evolutionary *theory*, and it makes
|
||
falsifiable predictions — a merge-compatibility cliff among them — that the metaphors do not.
|
||
|
||
**What is still open — honestly.** The old hole (what to select) we fill in kind: don't design the
|
||
selector, evolve it. But the hole has *moved*, not closed, and the new one is harder: **the fitness
|
||
function** — what reality-anchored measure selects for *truth* without also selecting for *persuasion*,
|
||
given that in our own species the two have been at war for the whole history of ideas. Alongside it:
|
||
the **institutions** that let contemporaries correct one another before error is inherited (§8), which
|
||
we do not solve; and the **calibration** of everything the results left as knobs — how many parents,
|
||
how complementary, at what ratio of inherited-to-real data, and how healthy a lineage must be before
|
||
its knowledge is safe to make irreversibly innate. These are, at least, *measurable* — which is the
|
||
difference between an open problem and a hole. And the largest gap of all: the *recombination* claims
|
||
now hold in real language models, but the *society* — the grounded, diversity-preserving, continually
|
||
reproducing loop — does not yet. The real test is to build that whole system out of actual open-weight
|
||
language models, and see whether all the signs survive contact with a system too big to write down.
|
||
The operators, checked; the living society, next.
|
||
|
||
---
|
||
|
||
## Selected references
|
||
|
||
- Akiba, T., Shing, M., Tang, Y., Sun, Q., & Ha, D. (2024). Evolutionary optimization of model merging recipes. *Nature Machine Intelligence.* (See also Sakana AI's M2N2, "Model Merging of Natural Niches.")
|
||
- Baldwin, J. M. (1896). A new factor in evolution. *The American Naturalist.*
|
||
- Boyd, R., & Richerson, P. J. (1985). *Culture and the Evolutionary Process.*
|
||
- Campbell, D. T. (1974). Evolutionary epistemology. In *The Philosophy of Karl Popper.*
|
||
- Fisher, R. A. (1930). *The Genetical Theory of Natural Selection.*
|
||
- French, R. M. (1999). Catastrophic forgetting in connectionist networks. *Trends in Cognitive Sciences.*
|
||
- Hinton, G. E., & Nowlan, S. J. (1987). How learning can guide evolution. *Complex Systems.*
|
||
- Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. *arXiv:1503.02531.*
|
||
- Hu, E. J., et al. (2021). LoRA: low-rank adaptation of large language models. *arXiv:2106.09685.*
|
||
- Hull, D. L. (1988). *Science as a Process.*
|
||
- Kauffman, S. A., & Levin, S. (1987). Towards a general theory of adaptive walks on rugged landscapes. *Journal of Theoretical Biology.* (The NK model.)
|
||
- Lehman, J., & Stanley, K. O. (2011). Abandoning objectives: evolution through the search for novelty alone. *Evolutionary Computation.*
|
||
- Mallya, A., & Lazebnik, S. (2018). PackNet: adding multiple tasks to a single network by iterative pruning. *CVPR.*
|
||
- McClelland, J. L., McNaughton, B. L., & O'Reilly, R. C. (1995). Why there are complementary learning systems in the hippocampus and neocortex. *Psychological Review.*
|
||
- McCloskey, M., & Cohen, N. J. (1989). Catastrophic interference in connectionist networks. *Psychology of Learning and Motivation.*
|
||
- Minsky, M. (1986). *The Society of Mind.*
|
||
- Mouret, J.-B., & Clune, J. (2015). Illuminating search spaces by mapping elites (MAP-Elites). *arXiv:1504.04909.*
|
||
- Muller, H. J. (1932). Some genetic aspects of sex. *The American Naturalist.* (The advantage of recombination.)
|
||
- Muller, H. J. (1964). The relation of recombination to mutational advance. *Mutation Research.* (Muller's ratchet.)
|
||
- Pearl, J. (2009). *Causality: Models, Reasoning, and Inference* (2nd ed.).
|
||
- Popper, K. (1959). *The Logic of Scientific Discovery.*
|
||
- Riis, S. (2026). Drift and selection in LLM text ecosystems. *arXiv:2604.08554.*
|
||
- Rusu, A. A., et al. (2016). Progressive neural networks. *arXiv:1606.04671.*
|
||
- Shumailov, I., et al. (2024). AI models collapse when trained on recursively generated data. *Nature.*
|
||
- Wortsman, M., et al. (2022). Model soups: averaging weights of multiple fine-tuned models. *arXiv:2203.05482.*
|
||
- Wright, S. (1931). Evolution in Mendelian populations. *Genetics.*
|
||
|
||
*The evolution of sex (the geneticist's canon this paper draws on):*
|
||
|
||
- Barton, N. H., & Charlesworth, B. (1998). Why sex and recombination? *Science.*
|
||
- Otto, S. P., & Lenormand, T. (2002). Resolving the paradox of sex and recombination. *Nature Reviews Genetics.*
|
||
- Kondrashov, A. S. (1993). Classification of hypotheses on the advantage of amphimixis. *Journal of Heredity.*
|
||
- Dobzhansky, T. (1936); Muller, H. J. (1942). Bateson–Dobzhansky–Muller incompatibilities (reproductive isolation).
|
||
|
||
*The 2025–2026 landscape this paper positions against:*
|
||
|
||
- Subramaniam, V., Du, Y., Tenenbaum, J. B., Torralba, A., Li, S., & Mordatch, I. (2025). Multiagent finetuning: self-improvement with diverse reasoning chains. *arXiv:2501.05707.*
|
||
- Zhang, J., Hu, S., Lu, C., Lange, R., & Clune, J. (2025). Darwin Gödel Machine: open-ended evolution of self-improving agents. *arXiv:2505.22954.*
|
||
- *Nature-inspired population-based evolution of large language models* (GENOME/GENOME+). (2025). *arXiv:2503.01155.*
|
||
- Sakana AI (2025). Competition and attraction improve model fusion (M2N2). *arXiv:2508.16204* (GECCO '25).
|
||
- Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). TIES-Merging: resolving interference when merging models. *NeurIPS / arXiv:2306.01708.*
|
||
- Yu, L., Yu, B., Yu, H., Huang, F., & Li, Y. (2023). Language models are super Mario: absorbing abilities from homologous models (DARE). *arXiv:2311.03099.*
|
||
- Gerstgrasser, M., et al. (2024). Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. *arXiv:2404.01413.*
|
||
- Guo, D., Wu, J., & Yiu, S. M. (2026). Model collapse as cultural evolution. *arXiv:2605.23054.*
|
||
|
||
*Still to engage in a full version: reproductive-isolation/speciation for merge compatibility (Git Re-Basin and linear mode connectivity as the mechanism); the machine-learning-native theory of merge degradation with task count; tacit knowledge (Polanyi) and human capital (Becker).*
|