papers: consolidate the real-LLM recombination arc (merge/moe/directed + headroom law)

Folds the four LLM runs into both papers now that the arc is fully characterised.

results-summary.md: new section 5 "The claims tested in real LLM weights" — the
Fisher-Muller generalist (7B merge 0.87 > best specialist 0.77), union-vs-fusion
(routing beats averaging where there is headroom), directed sex (breed + select),
and the unifying HEADROOM law that resolves the earlier saturation confound.
Updates design rule 2 ("merge, don't average — where there is headroom"), adds a
plain-language point 8, refreshes the validation counts (131 tests, real-LLM tier).

the-lamarckian-society-v5.md: softens the three "not a language model yet" claims
to acknowledge the prototype; adds the real-LLM confirmation after the
Fisher-Muller and directed-sex claims (with the headroom caveat as a sharpening,
not a weakening); adds an LLM bullet to section 13; reframes the closing gap to
"the operators, checked; the living society, next."

Tone held to the accessibility/honesty bar: prototype scope flagged (signs not
magnitudes), caveats presented as sharpening the claims, nothing over-celebrated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Giorgio Gilestro 2026-07-05 19:20:19 +01:00
parent 39f6c9f4df
commit 4e86d8b602
2 changed files with 111 additions and 27 deletions

View file

@ -13,8 +13,9 @@ sexual society**: reality-checking (grounding) plus **recombination across many
parents** (sexual reproduction, not teacher→pupil copying) plus **quality-diversity** selection.
The payoff is not merely arrested collapse but a population whose **offspring exceed their
parents** and whose general capability **climbs** while specialties are re-earned — and removing
any one operator breaks it. **122 tests pass**; three of the core predictions are exact closed
forms, so the headline curves sit on analytic targets rather than eyeballing.
any one operator breaks it, and a first **real-LLM prototype** (up to 7B on HPC) confirms the
recombination claims in trained weights. **131 tests pass**; three of the core predictions are exact
closed forms, so the headline curves sit on analytic targets rather than eyeballing.
---
@ -145,7 +146,47 @@ FisherMuller effect):
a rugged landscape you need diversity to explore basins, sex to recombine them, and grounding to
select on reality.
### 5. Positioning — what is prior art, what is ours
### 5. The claims tested in real LLM weights (Layer 2 prototype)
The analytic operators above make three claims that are testable in *real language-model weights*:
"merge, don't average" (E4), the FisherMuller generalist-from-specialists (E8), and directed sex
(E10). We tested them with a small, reproducible pipeline — LoRA **specialists** fine-tuned on three
*disjoint*, procedurally-generated task families (`lists`, `strings`, `arith`) with an **exact-match
verifier** as reality's "no" — recombined by various operators and scored on held-out tasks. Runs are
at two scales: **Qwen2.5-0.5B** locally and **Qwen2.5-7B** on Imperial College's CX3 HPC (one L40S).
This is a prototype (three families, one seed), so read it as **signs, not magnitudes**; the analytic
layer carries the quantities.
- **Recombining specialists yields a generalist that exceeds every parent (E8).** At 7B, a weight-space
**merge of the three specialists reaches 0.87 overall vs 0.77 for the best single specialist**, and
beats every specialist on every family — the FisherMuller signature, in real weights. At 0.5B the
same sign is present but marginal, because a weak base *dilutes* (see below).
- **Union beats averaging — when there is headroom (E4).** Two ways to recombine: **fusion** (average
the LoRA deltas — a "model soup") versus **union** (keep each specialist intact and *route* each
input to the right one — a Mixture-of-Experts). On tasks with headroom, **routing beats fusion**
because averaging dilutes a specialist's contribution: at 0.5B routing 0.74 > soup 0.64; on *hard*
7B tasks routing 0.50 > soup 0.40, with fusion diluting the fragile string-cipher specialist so badly
(0.67→0.30) that the soup even loses to the best single specialist. This is E8's `max > mean` in real
weights.
- **Directed sex helps — when there is headroom (E10).** Breeding a population of recombinant offspring
(specialists merged at many weightings), scoring each on a held-out split with the verifier, and
keeping the fittest **beats the single a-priori soup** where the soup is suboptimal: 0.69 > 0.64 at
0.5B, and 0.49 > 0.39 on hard 7B tasks — recovering most of routing's benefit from one deployable
model. This is offspring selection (which biology cannot preview) realised in weight space.
- **The one law that unifies them: headroom, not scale.** Our first 7B runs on *easy* tasks showed the
opposite — fusion beating union, and directed selection tied with the soup. That was a **saturation
artefact**: the easy families saturated 7B at 1.00, leaving no room to lose to dilution, so the naive
soup was already optimal. On *hard* (unsaturated) 7B tasks the original ordering returns. **The
operative variable is headroom**: union-preserving merging and offspring selection pay off whenever
there is room to lose to dilution — a weak base *or* hard tasks — and only the degenerate corner of
easy tasks on a strong base makes naive averaging look sufficient. This both confirms E4/E8/E10 in
real weights and delimits exactly when the refinements matter.
Honest scope: three lexically-distinct families (so the router is trivially accurate — an ambiguous-
skill benchmark is the interesting next stress test), one seed, small LoRA. The full *grounded*
society on LLMs (collapse under dry self-training, the dynamic society) is the HPC-scale next step.
### 6. Positioning — what is prior art, what is ours
An independent 2026 paper (Riis, *Drift and selection in LLM text ecosystems*) rigorously formalises
**collapse = WrightFisher drift** (martingale of minority mass, rare-first extinction, de Bruijn
@ -166,8 +207,11 @@ engineered cure and its integration.
1. **Never inherit dry, and ground where it matters** — a little reality (`g*≈5%`) protects most
diversity, but it protects only what it overlaps, and it can't hold the deep tail.
2. **Merge, don't average** — union-preserving model-merging realises the multi-teacher benefit;
averaging cancels it.
2. **Merge, don't average — where there is headroom** — union-preserving model-merging (routing /
max-merge) realises the multi-teacher benefit; averaging dilutes it. Confirmed in real LLM weights,
with a sharp caveat: dilution only bites when the task leaves room for it — on easy tasks a strong
base's naive soup already composes to the ceiling, so averaging looks fine. The refinement matters
exactly in proportion to how far the naive soup is from optimal.
3. **Sex, with no parent limit** — recombining decorrelated specialists yields offspring that exceed
any parent; more, complementary parents climb higher.
4. **Match recombination to entanglement** — merge freely for complementary skills; sparingly for
@ -230,6 +274,13 @@ that lets a society climb instead of collapse.
consensus (the exact analogue of training on the internet's AI-generated crowd); remove **sex** and
it gets stuck; remove **diversity** and it converges too fast to a worse answer. Each failure is
distinct; only the full society wins.
8. **It shows up in real language models — with one clean caveat.** We merged LoRA-specialised
Qwen models (0.5B locally, 7B on a university GPU cluster): the recombined model beats every
specialist (FisherMuller, for real), and *routing* / *offspring-selection* beat naive averaging —
but only when the tasks are hard enough to leave room. On easy tasks a strong model's plain average
is already at ceiling, so the fancier operators don't help. The lesson is precise: **these
recombination tricks matter exactly in proportion to how far the naive average is from the best you
could do** — which is a genuinely useful thing to know before you spend compute on them.
### Why it is novel and why it matters
@ -240,9 +291,10 @@ that lets a society climb instead of collapse.
- **The sexual-reproduction frame is, we believe, genuinely new for AI** — model merging reframed as
meiotic recombination, with a rigorous account of when it helps (complementary skills), when it
hurts (entangled skills), and how to make it reliably win (directed sex, unbounded parents).
- **It is validated, not vibes.** Three core predictions are exact closed forms; the neural and
image results confirm the *signs* in real trained weights; 122 automated tests pass; the whole
study is laptop-reproducible from a seed.
- **It is validated, not vibes.** Three core predictions are exact closed forms; the neural, image,
and **real-LLM** results confirm the *signs* in real trained weights (up to 7B on HPC); 131
automated tests pass; the whole study is laptop-reproducible from a seed (the LLM tier statistically
reproducible on one GPU).
- **It gives concrete design rules** for anyone building self-improving or model-merging systems:
ground where it matters, merge-don't-average, match recombination to skill-entanglement, select
offspring, keep diversity, and gate irreversible consolidation on health.

View file

@ -35,9 +35,11 @@ you skipped a definition:
it knew when it learns something new.
We have tried to keep the big picture legible on every page, and to be candid about what is argument
and what is evidence. The evidence is from **deliberately small models** — mathematics, small neural
networks, image generators, and evolutionary simulations. None of it is a large language model yet.
We will say so repeatedly, because the gap matters.
and what is evidence. The evidence is mostly from **deliberately small models** — mathematics, small
neural networks, image generators, and evolutionary simulations. A first bridge to real language
models exists — a prototype that recombines LoRA-specialised Qwen models up to 7B on a GPU cluster,
which confirms the recombination signs (below) — but the *full grounded society* has not yet been
built on a large language model. We will say so repeatedly, because the gap matters.
---
@ -67,13 +69,16 @@ AI has one advantage biology lacks: its "sex" has **no two-parent limit**, its m
for complementarity, and its offspring can be **screened before they are kept**. We call this
*directed sex*, and it turns recombination from a gamble into a reliable engine.
We support the argument with a set of **minimal models** (not language models): a mathematically exact
account of collapse and its cure; reproductions of the same effects in small trained neural networks
and in a generator of handwritten digits; and evolutionary simulations in which a population of models
climbs a "fitness landscape" that stands in for reality. In these, every claim above either holds or
fails visibly, and removing any single ingredient breaks the system in a distinct way. The scope is
honest: these are existence proofs and design rules, and the eventual test is to build the same thing
out of real language models. We also note where our diagnosis is no longer novel — the reading of
We support the argument with a set of **minimal models**: a mathematically exact account of collapse
and its cure; reproductions of the same effects in small trained neural networks and in a generator of
handwritten digits; and evolutionary simulations in which a population of models climbs a "fitness
landscape" that stands in for reality. In these, every claim above either holds or fails visibly, and
removing any single ingredient breaks the system in a distinct way. A first **language-model
prototype** then confirms the recombination claims in real weights — merging LoRA-specialised Qwen
models (up to 7B on a GPU cluster) produces a generalist that exceeds every parent, and shows the
sharp condition under which the "merge, don't average" refinement matters (below). The scope is
honest: these are existence proofs and design rules, and the eventual test is to build the *whole
grounded society* out of real language models. We also note where our diagnosis is no longer novel — the reading of
collapse as genetic drift has since been derived independently — and locate our contribution in the
**cure** rather than the diagnosis.
@ -207,6 +212,19 @@ knowledge; specialisation is re-earned each generation," and it is why the refra
teacher→pupil to *sexual reproduction* is not cosmetic: **copying can only recover a ceiling;
recombination can exceed it.**
This is no longer only a simulation. In a first language-model prototype — LoRA specialists on
disjoint task families, recombined and judged by an exact verifier — a merge of three specialist Qwen
models (7B, on a GPU cluster) **beats every single specialist**, overall and on every family: the
FisherMuller effect, in real weights. The same prototype pins down *when* the finer "inherit the
union, don't average" rule actually bites. Keeping each parent whole and **routing** each input to the
right one beats the tail-thinning average — but only when the task is hard enough to leave room to
lose: on easy tasks a strong model's plain average is already at the ceiling, so the crude soup is
fine, whereas on hard tasks the average dilutes a hard-won specialist so badly it falls below even the
best single parent, and routing wins by a wide margin. The rule is therefore precise: **the union
beats the average in exact proportion to how far the average is from the best attainable** — a caveat
that sharpens rather than weakens the claim, and that a practitioner needs before spending compute on
the fancier operator.
Two caveats keep this honest, and both are results, not hand-waving.
*Sex can backfire.* When the parents' skills are not cleanly separable but **entangled** — when the
@ -223,8 +241,11 @@ and cannot inspect an offspring before it is born. An AI has none of those limit
**generate many candidate offspring and keep only the fittest**, screening them against reality before
committing. We call this **directed sex**, and in our simulations it converts the outbreeding-depression
catastrophe into a reliable gain: where blind recombination collapses on entangled problems, directed
recombination matches or beats the best parent every time. This is a genuine advantage of engineered
reproduction over the biological kind, and we think it is one of the more useful ideas in the paper.
recombination matches or beats the best parent every time. The language-model prototype shows the same
sign where it can: breeding many recombined Qwen offspring and keeping the one the verifier scores
highest beats the single averaged soup on hard tasks (and, unsurprisingly, does nothing extra on easy
tasks the soup already solves). This is a genuine advantage of engineered reproduction over the
biological kind, and we think it is one of the more useful ideas in the paper.
So the picture of §5 is: single-teacher copying is asexual and collapses (Muller's ratchet = model
collapse); the cure is to *ground* every birth in reality and to reproduce *sexually*, recombining
@ -443,9 +464,10 @@ refute, or you evolve toward what is merely convincing.**
## 13. What we built, what we found, and what is still open
The previous drafts of this paper promised a "companion paper" that *would* make this concrete. That
work now exists — as a set of **minimal, laptop-reproducible models**, not language models — and it is
worth stating plainly what it does and does not show. (A separate results document gives the numbers;
here is the shape.)
work now exists — mostly as a set of **minimal, laptop-reproducible models**, with a first bridge to
**real language models** (a LoRA-merge prototype, up to 7B on a GPU cluster) — and it is worth stating
plainly what it does and does not show. (A separate results document gives the numbers; here is the
shape.)
**What we built and found.**
@ -471,6 +493,15 @@ here is the shape.)
produces below-parent offspring (outbreeding depression) — and *directed* recombination (choose mates,
screen offspring, unbounded parents) reliably fixes it. This is the concrete evidence for the paper's
central reframing.
- *The recombination claims, in real language models — with a sharp condition.* Merging LoRA-specialised
Qwen models (up to 7B on a GPU cluster) produces a generalist that beats every specialist parent
(FisherMuller, for real); and keeping parents intact and *routing*, or *breeding and screening*
offspring, beats the naive average — but *only when the task leaves headroom*. On easy tasks a strong
model's plain average is already at the ceiling and the refinements add nothing; on hard tasks the
average dilutes a specialist below even the best single parent, and the union-preserving operators win
clearly. The practical rule is exact: these tricks pay off in proportion to how far the naive average
is from the best attainable. This is a prototype (three task families, one seed), so we read it as
signs, not magnitudes; the *whole grounded society* on a language model remains the open step.
- *The whole society, and why every part is needed.* In a population evolving on a "reality" landscape,
the full system — grounding + sexual recombination + preserved diversity — climbs to the top while
keeping its specialists. Remove *grounding* and it collapses into a confident, wrong consensus (a
@ -499,10 +530,11 @@ the **institutions** that let contemporaries correct one another before error is
we do not solve; and the **calibration** of everything the results left as knobs — how many parents,
how complementary, at what ratio of inherited-to-real data, and how healthy a lineage must be before
its knowledge is safe to make irreversibly innate. These are, at least, *measurable* — which is the
difference between an open problem and a hole. And the largest gap of all: every result here is from a
minimal model. The real test is to build the same grounded, sexually-reproducing, diversity-preserving
society out of actual open-weight language models, and see whether the three signs survive contact with
a system too big to write down. Theory and toys here; the real thing, next.
difference between an open problem and a hole. And the largest gap of all: the *recombination* claims
now hold in real language models, but the *society* — the grounded, diversity-preserving, continually
reproducing loop — does not yet. The real test is to build that whole system out of actual open-weight
language models, and see whether all the signs survive contact with a system too big to write down.
The operators, checked; the living society, next.
---