MachineSex/paper/the-evolution-of-sex-for-ai-accessible.md
Giorgio Gilestro 58e6c74609 paper: claim-narrowing revision from the external review
The review's core instruments adopted: the interpretation/explanation/
prediction ladder is now explicit in §1 (with the decisive pre-merge
epistasis-prediction test stated as the open bar, not claimed); identity
claims scoped (WF exact only in the minimal model, with the
learning-kernel deviation cited against ourselves; Muller's ratchet
scoped to the irreversible arm — recombination reassembles only what
survives); "nobody has / none imports / theory outrun" removed;
merge-don't-average given explicit operator boundaries (output-mean vs
weight-average vs routing vs max-with-oracle; budgets; oracle; capacity
handoff to speciation); a "what these experiments do and do not
establish" scope block added to the speciation section (conflict floor
is information-theoretic, not genetic; epistasis-cliff + snowball =
hypotheses at the neural tier; emergent DMIs = flagship hypothesis,
bounded by our null); "control theory" -> "framework" (subtitle
included); §3/§11 overstatements fixed (frozen core != frozen behaviour;
Baldwin echo, not identity; operational vs archival irreversibility);
claims-at-a-glance table (status/assumptions/evidence/limits) added to
§13. Reviewer's framing sentence adopted as the stated core
contribution. Accessible version calibrated to match. md2tex gains pipe-
table support; PDF rebuilds clean (22 pp). Lessons recorded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
2026-09-06 15:14:28 +01:00

38 KiB
Raw Permalink Blame History

The Evolution of Sex for Artificial Intelligence — the plain-language version

How ideas from breeding and genetics tell us how to keep AI models improving across generations

This is an accessible companion to the full paper (the-evolution-of-sex-for-ai.md). It makes the same argument and reaches the same conclusions, but assumes only that you know roughly what a machine learning model is — that it is trained on data, that training adjusts numbers called "weights," and that you can fine-tune a model on new data. Everything else is explained as we go. Where the full paper defends each point against the research literature, this version just tells the story.

Giorgio F. Gilestro · Department of Life Sciences, Imperial College London · giorgio@gilest.ro · https://lab.gilest.ro


The one-paragraph version

If you train an AI model on the output of earlier AI models, over and over, it rots: rare knowledge disappears and everything drifts toward a bland average. This is a known problem ("model collapse"), and in its simplest form it is governed by exactly the same math that describes how small biological populations lose rare genes by chance. That is bad news, but it comes with good news: biology has been managing this kind of rot for hundreds of millions of years, and its best-tested remedy is sex — making offspring by combining several parents instead of copying one. This paper tests how far that remedy carries for AI. This paper takes ninety years of genetics about when and why sex beats cloning and reads it as an engineering manual for building AI that keeps getting better across generations instead of decaying. Along the way it produces concrete, testable rules — including a surprising one about how to combine models ("merge, don't average"), and a limit ("models can drift so far apart they can no longer be usefully combined at all"). We back the argument with small, fully reproducible experiments and a first test on real language models.


A few words you'll need

  • Model collapse — what happens when you train models on the output of earlier models, again and again: rare cases vanish, everything gets blander. The central disease this paper is about.
  • Fine-tuning / specialising — taking a trained model and training it a bit more so it gets good at one specific thing.
  • Model merging — combining two or more trained models directly, by mixing their weights, to get one model — without retraining. Think "breeding two models" rather than "teaching a third."
  • The tail — the rare stuff. Common knowledge is the "head" of the distribution; unusual cases, rare facts, and edge behaviours are the "tail." Collapse eats the tail first.
  • Grounding — mixing some real, verified data from the actual world into training, instead of only model-generated data. The reality check.
  • Genetic drift (from biology) — in any finite population, rare gene variants can vanish purely by chance, because not every individual reproduces. This is the biological twin of model collapse.
  • Recombination / sex (from biology) — making a child by combining pieces of more than one parent. The opposite of cloning (asexual reproduction).

1. A society of AIs across time, not just space

When people imagine "many AIs working together," they usually picture teamwork in the moment: several specialist agents side by side, splitting up a job. This paper is about a different direction: time. Not AIs that cooperate right now, but AIs that pass knowledge down across generations — each new model starting from what the previous ones learned, the way each human generation inherits the accumulated knowledge of the last and adds a little.

The key event, then, is reproduction: making a new model out of older ones. A single model, like a single person, eventually stops improving. A lineage — a chain of models across generations — does not have to. Civilisation isn't smart because any one person is; it's smart because knowledge accumulates. The whole question of this paper is: how do you make one AI model out of older ones, without the knowledge rotting on the way down? That's exactly where it can go wrong.

(This corner of AI is suddenly busy: several 20252026 research projects build populations of models that improve over rounds, and "model merging" has become a small industry that already borrows words like crossover, mutation, and mate choice. What this paper adds is the quantitative framework behind those borrowed words — and honest tests of where it works and where it doesn't.)

2. Why today's models can't do this

Today's large models have no life cycle. A model is trained once, at huge expense, then frozen and shipped. It does not learn from the people who use it. "Learning" and "doing" are two separate eras with nothing connecting them.

There's a real reason for the freeze: if you keep training a neural network on new things, it tends to overwrite what it already knew. This is called catastrophic forgetting, and it's been a known problem since the 1980s. Freezing dodges it by refusing to learn at all. But a lineage needs the opposite of a frozen model: it needs members that keep learning through their working lives and then pass on what they gained. So step one is a model that can grow safely.

3. A model that grows without forgetting

The trick is to stop overwriting. Keep the model's original core frozen and untouchable, and bolt each new skill onto extra capacity added beside it. In practice this is what small add-on "patches" like LoRA already do: the big pretrained model stays fixed, and you train a little attachable module for each new skill. If the core is never altered, its knowledge can't be erased — though the system's behaviour can still change while patches are active; what's guaranteed is a recoverable core.

There's even a rough brain analogy: we have a fast memory (the hippocampus) that grabs an experience immediately, and a slow memory (the cortex) that absorbs patterns gradually, usually while we sleep. The AI version is clean: the prompt is short-term memory, a database is fast memory, the trained weights are slow memory, and a periodic "consolidation" step moves knowledge from fast to slow.

One consequence matters a lot: because this kind of model only ever adds capacity, it eventually fills up. In most designs that's a problem. Here it's a feature — read on.

4. "Full" means grown up, not broken

Here's the pivot. When a model that can only add capacity finally fills up, it hasn't failed. It has matured.

Think of the capacity limit as a life stage. A model is born as a freshly trained base — its general education. It has a working life, picking up specialised expertise on the job. And it reaches maturity — the point where it has learned about as much as one working life in its niche can teach. Maturity isn't the end of usefulness; it's the moment the model is most worth learning from. So maturity is the signal to reproduce. The capacity ceiling that every other design fights becomes, here, the clock that times the generations.

Everything now depends on how that reproduction happens. This is the heart of the paper.

5. Copying rots; combining climbs

Suppose a mature model just teaches a fresh one, and that one teaches the next, and so on down the line. It's the obvious design — and it fails, for the same reason in AI and in biology.

In AI terms: training each generation on the previous generation's output is the exact recipe for model collapse. The model forgets the improbable, loses the rare cases (the tail) first, and drifts toward its own most common output. And here's the nasty part: the thing that makes teaching-a-student useful — "keep the general, drop the quirky" — is the same act that deletes the tail. The operation you want and the operation that kills the lineage are the same move.

In biology terms (and for our simplest model it really is the same math — real networks add a measurable twist on top, which we also measure): picture a model's knowledge as a big bag of items — facts, skills, behaviours — in certain proportions. One generation is: draw a finite sample from the parent, and rebuild the child from that sample. That "finite sample" step is identical to genetic drift — the way rare gene variants vanish by chance in any finite population. This isn't a loose analogy; it's the same century-old equations (the WrightFisher model), which is why we can check our simulations against them exactly. Rare items go extinct first, about ten times faster than common ones — precisely what drift predicts.

And copying one teacher is asexual reproduction — cloning. Biology already knows the fate of a lineage that only ever clones and never combines: it piles up damage it can never undo, a one-way decline called Muller's ratchet. That's our lens for the irreversible part of model collapse — the capabilities that, once every copy is gone, no amount of combining can rebuild. Naming it that isn't just poetry — it tells us where to look for remedies, because biology has spent a very long time solving exactly this.

Two ingredients turn the rot into a climb. Both are things nature does.

Ingredient one: don't reproduce "dry." Collapse only happens to a lineage fed nothing but its own output. Mixing in some real, verified data from the world — we call this grounding — stops it. In our small experiments, grounding is shockingly cheap: even a few percent of real data keeps most of the diversity alive indefinitely. But we found an honest limit we didn't expect: grounding can't save the very rarest items at any affordable cost — protecting something of rarity p needs a real-data budget that grows like 1/p. Grounding rescues diversity cheaply, but not the deepest tail. Something else has to do that. That something is sex.

Ingredient two: reproduce sexually. Instead of copying one parent, build each new model by combining several — a sexual birth, not an asexual one. AI already has a tool for this: model merging. Why does it help? If several parent models each specialised on different things, each one kept alive rare knowledge the others lost. A combined child inherits the union of what its parents kept — not the tail-thinned average of a crowd of near-identical clones. And here's the point that turns sex from a mere safety net into the engine of the whole thing:

A child combined from complementary parents can be better than any of its parents.

Geneticists call this the FisherMuller effect: recombination gathers, into one individual, good variants that arose separately in different lineages — so the child has a combination none of the parents had. Our simulations show exactly this: combining specialist models that each mastered different skills produces a model that climbs toward the best possible combination — one no single parent had — while the best single parent, and the plain average of all of them (a "model soup"), both level off well below. This is the paper's core claim in one line: copying can only recover a ceiling; combining can break through it.

And it's not just simulation. In a first test on real language models — three small Qwen models, each fine-tuned on a different family of tasks, then merged and graded by an automatic checker — the merge beat every single specialist, overall and on every task family. The FisherMuller effect, in real weights.

That same test pinned down a subtle rule about how to combine models:

"Merge, don't average" — but only when there's room to lose.

Keeping each parent whole and routing each question to the right specialist beats crudely averaging them together — but only when the task is hard enough that averaging actually damages something. On easy tasks, a strong model's plain average is already about as good as possible, so the crude soup is fine. On hard tasks, averaging waters down a hard-won specialist so badly the blend falls below even the best single parent — and the smarter "keep-them-separate-and-route" approach wins big. So the rule is precise: the fancy combining tricks help in exact proportion to how far the plain average is from the best you could achieve. A practitioner needs to know this before spending compute on the fancy version.

Two honest caveats, and both are actual findings, not hand-waving:

Sex can backfire. When the parents' skills aren't cleanly separable but tangled together — when skill A only pays off if skill B is also present (biologists call this epistasis) — blindly combining two good models can produce a worse child, because combining breaks apart a package that only worked as a whole. Biologists call this outbreeding depression, and we reproduce it: on "tangled" problems, naive merging drops the child below its parents, and the more you mix, the worse it gets. The design rule: combine freely when skills are independent; combine sparingly and carefully when they're tangled.

How widely you mate matters too. That last point was about how much to mix; a separate knob is who mixes with whom. Monogamy = each model only ever combines within a small, fixed circle; promiscuity = any model can combine with any other. Almost all model-merging today is promiscuous by default — throw everything in one pot. But there's a catch: wide mixing spreads good traits fast, but it also makes the whole population converge to the same thing, killing variety. Narrow, local mixing keeps separate sub-groups exploring different solutions. We tested this against tangledness, and the best answer moves: on simple (independent-skill) problems, wide promiscuous merging is best; but the more tangled the skills, the more you want to narrow it — full promiscuity converges too fast onto one mediocre solution and finds a worse champion, while keeping structured sub-groups preserves the variety a hard problem needs. So the rule extends: merge widely for independent skills; keep separate sub-populations ("island" merging) for tangled ones.

AI can do sex better than biology can. Biology is stuck with two parents, mating more or less at random, and can't inspect a child before it's born. AI has none of those limits. It can combine many parents at once; it can choose which parents to combine, for complementary skills; and it can generate many candidate children and keep only the best, testing them against reality before committing. We call this directed sex, and in our simulations it turns the outbreeding-depression disaster into a reliable win: where blind combining collapses on tangled problems, directed combining matches or beats the best parent every time. The real-language-model test shows the same where it can: breeding many merged offspring and keeping the one that scores highest beats the plain soup on hard tasks. This is a genuine advantage of engineered reproduction over the biological kind, and it's one of the more useful ideas in the paper.

So §5 in one breath: copying is asexual and rots (Muller's ratchet = model collapse); the cure is to ground every birth in reality and to reproduce sexually, combining many complementary parents; and because AI's version of sex can use many parents, chosen mates, and pre-screened offspring, it's not just insurance against collapse — it's an engine that makes children better than any parent.

The limit of sex: models can drift too far apart to merge

Sex has a limit, and it's the sharpest new prediction here. Combining parents works because they're variations on a shared background. Push two lineages far enough apart and their combination stops being viable. In biology this is speciation — two populations become separate species that can no longer interbreed. The genetic mechanism has a name (a BatesonDobzhanskyMuller incompatibility): a change that arose in lineage A and a change that arose in lineage B are each harmless on their own, but their combination — which neither lineage ever tested — is broken in the hybrid. A merged model is exactly such a hybrid. So the theory predicts a specific trajectory as two models drift apart: they merge fine → merging starts to hurt → merging becomes useless.

We built this as an explicit model and confirmed the predicted curve. Three things come out of it:

  1. There's a cliff, and epistasis moves it. The point where merging fails isn't fixed — it comes earlier the more tangled (epistatic) the skills are. This is a distinct, testable claim: at the same amount of drift, whether two models can be merged depends on how tangled their skills are, not just on how far apart they are. The existing AI tools for predicting merge success only measure distance/geometry — they don't have this axis.
  2. It snowballs. The number of incompatibilities grows with the square of the drift, so merge quality doesn't fade gently — it falls off a cliff. Drift is punished faster than it accumulates.
  3. The design rule: before merging, weigh how far apart the models are against how tangled the skills are. Past the cliff, don't merge — route instead (keep the specialists separate and pick between them).

We also did the experiment a skeptic would demand. A known objection: "your 'incompatibility' is just a loss barrier, and those are famous for being fake — two networks can learn the same function but store it in a shuffled internal order, which looks like incompatibility until you line their neurons back up (a technique called Git Re-Basin)." So we tested it directly in real trained networks. We trained pairs of small networks, merged them, and measured the merge damage before and after re-aligning their neurons. The result is clean:

  • Two networks trained on the same task but from different random starts: big apparent merge damage, but re-aligning removes ~98% of it. That's the fake kind — same skill, shuffled order. (This also proves our alignment tool works.) And we allowed the aligner every legal move for these networks — not just re-ordering neurons but also re-scaling them — so nothing removable was left on the table.
  • Two networks trained on conflicting tasks: big merge damage, and even the full aligner removes essentially none of it. That's the real kind — genuine incompatibility, not a bookkeeping artifact. And it can't be waved away as "you just didn't align them well," because the exact same tool cleaned up the first case. There's even a simple proof that no future alignment trick can fix it: no single model can obey two rulebooks that contradict each other on the same inputs.

Sweeping from "no conflict" to "total conflict" traces a smooth incompatibility cliff in real weights: the merged model's accuracy slides from 0.97 (no conflict) down to 0.03 (total conflict) — a hybrid that is literally inviable. So the speciation effect is real, not a relabelled artifact.

And we ran the honest flip side, deciding in advance to report it either way: what if two networks just specialise differently, with no conflict at all — one keeps training only on digits 04, the other only on 59? Do they drift into incompatibility on their own? No. At every amount of divergence we tested, the merge damage stayed at zero — and the merged model actually rescued the two specialists: each parent alone had forgotten half the digits (scoring ~0.50), while their merge scored ~0.95. So in these experiments, models don't become unmergeable just by growing apart; they become unmergeable when they learn things that genuinely contradict. That's good news for merging — specialisation is safe, conflict is the danger — and it makes the theory's prediction sharper, not weaker.

One question is left hanging, and the rest of the paper is about it: combining preserves what the parents kept — but who decides what each parent keeps, and which children are worth keeping?

6. Don't design the selector — evolve it

There are two ways to answer that question, and the first one is wrong. We could try to hand-write the rule for "what knowledge to keep and pass on." But nobody actually knows that rule. "Keep the general, drop the specific" sounds wise until you ask which generalisations, in which domain, at which level of detail — and it falls apart. This is the deepest gap in the whole scheme, and you can't fill it by decree.

The second answer is the one nature used: don't design the selector — let it evolve. Let different models carry different policies about what's worth keeping and combining. Let the policies that produce strong children spread, and the policies that produce weak children die out with their lineages. What the lineage considers important — its "taste" — is discovered by selection, not imposed by us.

So two things get inherited, on two channels. The content — the actual knowledge — is passed down directly (this is the "Lamarckian" part: inheriting things acquired during a lifetime, which biology forbids for genes but culture allows for ideas). The selection policy — what to keep, who to breed with, which children to screen for — is itself inherited, varies between models, and survives in proportion to how well it works. That second channel is Darwinian. The system is both at once: it inherits content like culture, and selects policies like evolution.

And it closes neatly: Darwinian selection needs a pressure — something that decides which policies win. That pressure is already in the design. What tells a lineage its taste was good? Whether its children succeed against reality. The reality check that stops collapse (grounding, §5) and the fitness signal that guides the evolving taste turn out to be the same thing, seen from two angles.

7. The big danger: what you measure is not what you want

Adding selection adds selection's classic trap, and it's dangerous enough to sink everything if ignored. Evolution optimises, ruthlessly and without foresight, for exactly what you measure — never for what you meant. (In ML you know this as reward hacking or specification gaming; economists call it Goodhart's law.) Get the fitness measure a little wrong and the lineage will exploit the gap more cleverly than any rule you could write.

For a knowledge lineage there's a particularly nasty version. The natural way to measure how "good" an idea is might be how well it spreads — but a false-but-persuasive idea spreads beautifully. Human culture is full of highly contagious nonsense; confident wrongness routinely beats careful accuracy. Let selection loose on models without care and it will breed a lineage that is fluent, compelling, and wrong — model collapse with an optimiser actively steering toward the cliff.

Only one thing makes fitness track truth instead of appeal: being judged against a reality that can say no. Fitness has to mean "did this knowledge correctly predict what the world actually does when you act on it" — not approval, not fluency, not a gameable benchmark score. This is why the reality check matters twice: it's both the anchor that stops passive collapse and the only thing that keeps the evolving taste honest.

The second danger is everyone converging to the same thing, and avoiding it takes work at two levels, because selection can only preserve variety that already exists — variety first has to be supplied and then kept.

  • Supply. A lineage that only learns from the accredited elite has a monoculture for a source — the "top" experts are, by definition, the ones who won the consensus. So the system must deliberately learn from outliers and heretics too — not out of fairness, but because diverse starting material is the raw fuel for everything downstream.
  • Preserve. Even with varied input, plain "keep the best" selection converges — it stampedes toward the single current champion and wipes out the rare specialists. The fix is well known: quality-diversity selection, which rewards being good and being different at the same time, keeping complementary specialists alive. In our simulations this is decisive: greedy "keep-the-best" collapses diversity almost immediately and gets stuck on a mediocre answer, while quality-diversity keeps the very specialists that sexual combining then needs as parents.

The two levels meet at reproduction: multi-parent combining is the vehicle that carries the preserved diversity into the next generation. Supply the variety, preserve it, recombine it — remove any one and the lineage collapses onto its own first guess.

8. A society needs institutions, not just experts

One requirement is easy to miss and fatal to skip. The easy part of a society is having specialists. The hard part — which human civilisation took millennia to build — is the institutions that let fallible specialists combine without each re-checking everything: reputation, replication, credentials, and above all peer review. These are error-correction systems, and they exist because a group of unreliable experts left to reinforce each other is more wrong than any one of them alone.

This is exactly where today's multi-agent AI fails: put several models in a room and they tend to agree sycophantically and confabulate together, because they have all the specialisation and none of the institutions. A real multigenerational society has to specify not just how models learn, reproduce, and get selected, but how they check each other — how a wrong model loses standing before its error gets merged into its children and inherited.

9. The lineage must stay open to reality

However many generations deep it goes, a society of models shares one hard limit: it has only ever read. Its entire inheritance is a record of things that were said. It lives on the bottom rung of what causality researchers call the ladder — observation — and no amount of reading ever reaches intervention. Watching doesn't tell you what would happen if you acted; correlation isn't causation at any scale.

Only intervention — actually reaching out and changing the world to see what happens — climbs that ladder, and a language model can't intervene. That's what humans and their instruments provide, and the gift isn't "truth," it's constraint: reality's unique power is that it can say no. Text just offers more opinion; an experiment delivers a refusal no consensus can overturn. As before, that refusal does double duty: it stops collapse and keeps the evolving taste selecting for truth over persuasion.

Two honest riders. First, the human reality signal is dirty — warped by publication bias, incentives, and the occasional fraud — which is exactly why the error-correcting institutions of §8 have to sit at the humanmachine boundary. Second, humans are the current supplier of intervention, but the hands-on half is being automated (robot labs already run their own experiments). What looks durable in the human role isn't the hands — it's the choice of what to test and which refusals matter.

10. Why this is affordable

A practical fact turns this from daydream into buildable proposal: it almost never re-pays for the one genuinely expensive thing in AI — training a model from scratch.

Training a foundation model from scratch costs a fortune in data and compute. This design does none of that per generation. Every model is born from an existing open model that already paid that cost; specialising one is a small patch trained in hours on a single consumer GPU; running the society is ordinary use; and reproducing — merging parents into a child — can be done directly on the weights with no retraining at all. Selection costs more (you have to run populations and discard the losers), but that's a multiplier on an already-cheap unit, not on a from-scratch budget.

The economics only work with open-weight models — for practical and legal reasons at once. You have to be free to inspect, modify, and redistribute the weights, and most proprietary licences forbid using a model's output to train another (which is exactly what reproduction here does). That's not ideology — it's a structural constraint, and a democratising one: it puts the whole thing within reach of a single lab.

11. Can it grow forever? Baking knowledge back into the base

One thing we quietly assumed: can the lineage accumulate without end? The individual model is bounded — that's the clock. But the lineage looked unbounded, each generation starting a step ahead. Look closer and a second budget also fills up.

Every new model is a clean base plus an inherited pile of soft patches — the acquired knowledge carried in add-on modules rather than baked into the core. Those patches are what make the lineage multigenerational — but they're not free: they slow the model down, and past some depth they're better consolidated than carried. The lineage matures too.

The fix is the same operation, one level up. When a lineage's acquired knowledge has proven stable across enough generations, re-mint the base: bake all those accumulated patches into the weights of a fresh from-scratch-scale model — a new base that is born already knowing what took many generations to learn in patches. The soft budget resets; the next era starts from a higher floor. What was hard-won and learned becomes cheap and built-in. (Biologists have a name for acquired traits that eventually become innate: the Baldwin effect.)

Three honest riders, because this is the most consequential step:

  • Cost. This is the one step that re-pays part of the from-scratch bill — the exception to §10. It's bearable only because it's rare, spread thin over many cheap generations.
  • Irreversibility. Until now, one thing was always recoverable — the original clean base, whose lost rare knowledge you could restore just by reloading the file. Bake the current lineage into new permanent weights and that escape hatch closes: if the lineage had been quietly collapsing, re-minting freezes the collapse in place and throws away the one uncollapsed reference that could have caught it. In our experiments this happens exactly as feared — and a cheap safeguard prevents it: only re-mint while the lineage is provably healthy and diverse, never as a rescue for one already drifting.
  • Branching. Different labs will re-mint on different criteria and produce different bases; the lineage branches. That's not a bug — it's the tree of life, and it's exactly what open weights make possible.

So can it grow forever? Yes — but only because it forgets and consolidates at every level, including the base. Nothing is stored without limit anywhere. Unbounded growth of capability is bought with bounded storage plus periodic consolidation.

12. One idea at four speeds

Step back and it all resolves into a single process running at four nested speeds. The downward motion is transmission — passing hard-won knowledge down:

  1. Within one model, over its working life: experience moves from fast memory into slow, durable weights, without catastrophic forgetting.
  2. Between generations, at maturity: mature models reproduce, recombined into a fresh one.
  3. Across many generations: each inherits the compressed achievements of the last and adds a little.
  4. Across eras: a proven lineage's accumulated patches get baked into a re-minted base, becoming innate.

The first and last are the same operation at opposite ends of the scale — fast/soft memory consolidating into slow/hard memory — one running overnight inside a single model, the other across an era inside a whole society. The sideways motion is selection — acting across the population at every speed, on the policies that decide what gets passed on, with reality as the judge and diversity kept alive so the specialists survive.

Three rules govern all of it: reproduce by combining, not copying, or you rot; keep the disagreements and surprises, or you converge; and anchor fitness to a reality that can say no, or you drift toward whatever is merely convincing.

13. What we actually built, found, and left open

We didn't just argue this — we built small, fully reproducible models to test it, plus a first bridge to real language models. Here's the shape of the evidence (a separate document has the numbers).

What we built and found:

  • An exact account of collapse. Because generational training is the genetic-drift process, we can check our simulator against century-old formulas — and it matches to a fraction of a percent. Collapse isn't argued by analogy; it's derived.
  • Grounding is cheap — but has a limit. A few percent of real data keeps most of a lineage's diversity alive indefinitely — but not the very deepest rare cases, which need combining. This is what makes a continually-learning society economically realistic rather than a data-hungry fantasy.
  • "Merge, don't average." Combining several teachers by averaging their outputs — the obvious thing, and what a "model soup" does — mathematically cancels the benefit of having several teachers. A merge that keeps each item's strongest source realises it. Most current multi-model setups get this wrong by default.
  • Collapse and its cure in real networks, and on real images. We reproduced the same effects in small neural networks and in a generator of handwritten digits (MNIST): a model trained on its own output collapses to a single blurry digit, while a little grounding keeps all the styles alive. Honest wrinkle: real networks smooth things over, so the naive diversity metric can mislead — the right measure is distance-from-truth.
  • Sex that beats the parents — and when it doesn't. In simulations, combining complementary specialist models produces a model better than any parent (FisherMuller), climbing toward the best possible combination as you add more, more-diverse parents — while averaging and best-single-parent level off below. On tangled problems, blind combining instead produces below-parent children (outbreeding depression) — and directed combining (choose mates, screen offspring, many parents) reliably fixes it.
  • Monogamy vs promiscuity. Sweeping how widely models merge — from local/monogamous to everyone-with-everyone/promiscuous — against how tangled the skills are, the best breadth shrinks as the skills get more tangled: wide promiscuous merging wins when skills are independent, but on tangled problems it converges too fast onto one mediocre solution and finds a worse champion, so keeping structured sub-populations wins. (Promiscuity always lifts the typical model but always destroys variety.) A merging design knob the field, which throws everything in one pot, doesn't currently have.
  • The combining claims, in real language models — with a sharp condition. Merging fine-tuned Qwen models (up to 7B on a GPU cluster) produces a generalist that beats every specialist parent; and keeping parents separate and routing, or breeding and screening offspring, beats the plain average — but only when the task leaves room to lose. On easy tasks the plain average is already at the ceiling and the tricks add nothing; on hard tasks the average dilutes a specialist below even the best single parent, and the smarter operators win clearly. This is a prototype (three task families, one run), so we read it as signs, not exact numbers — the whole grounded society on a language model is still the open step.
  • The whole society, and why every part is needed. In a population evolving on a "reality" landscape, the full system — grounding + combining + preserved diversity — climbs to the top while keeping its specialists. Remove grounding and it collapses into a confident, wrong consensus (a direct analogue of training on the internet's growing pile of AI-generated text); remove combining and it gets stuck; remove diversity and it converges too fast to a worse answer. Each removal breaks differently; only the whole thing climbs.

What's borrowed vs. what's ours. We're careful about credit because the area is crowded. Already known (and we claim none of it): that model collapse is genetic drift; that a merged model can beat its parents; that decorrelated parents merge better; that naive averaging is worse than smarter merges; that populations of self-improving models can climb. What's genuinely new here is the theory those results have outrun — a real population-genetics of sex for model societies, which predicts rather than just observes: the "merge, don't average" law, FisherMuller as the reason children beat parents, outbreeding depression on tangled landscapes (turning "when does merging help?" into something the problem's structure predicts), grounding as a migration-vs-drift balance with a critical real-data fraction, directed sex as AI's distinct advantage, the whole integrated society whose parts are shown jointly necessary — and, newest of all, model speciation: the account of when two models are too far apart to merge at all, confirmed in real weights.

What's still open — honestly. We fill the old hole (what to select) by evolving the selector instead of designing it — but the hole moved rather than closed, and the new one is harder: the fitness function. What reality-anchored measure selects for truth without also selecting for persuasion, given that in our own species the two have been at war for the entire history of ideas? Alongside it: the institutions that let models correct each other before error is inherited (§8), which we don't solve; and the calibration of all the knobs the experiments left open — how many parents, how complementary, at what ratio of inherited-to-real data, how healthy a lineage must be before its knowledge is safe to make permanent. These are at least measurable, which is the difference between an open problem and a hole. And the biggest gap: the combining claims now hold in real language models, but the society — the grounded, diversity-preserving, continually reproducing loop — does not yet. The real test is to build that whole thing out of actual open language models and see whether all the signs survive. The operators, checked; the living society, next.


For the full argument, the literature it's positioned against, the exact predictions, and the references, see the complete paper: the-evolution-of-sex-for-ai.md.