narrative revision: substantiated premise, continual-learning positioning, convergence framing

Per GG's directive: (1) the model-societies premise is no longer asserted
— the Introduction opens with the verified evidence base (3M-model
ecosystem with phylogenetic lineage-mapping literature, >98%-synthetic
alignment pipelines, machine-generated web share, the human-data
ceiling, mainstream merging tooling, agent economies; refs 31-44, all
identifiers verified by the literature scan). (2) The findings are
contextualised in CONTINUAL LEARNING, where they land hardest: a new
Introduction block maps the CL canon onto the operators — replay <->
grounding, with the field's measured replay fractions (1%/5%/25%)
sitting on our theorized g*~0.05; pseudo-rehearsal/generative replay as
precisely our ungrounded null; parameter isolation; CLS consolidation;
merging-for-CL vs cross-lineage recombination; tail-first forgetting <->
tail-allele extinction; CF-vs-collapse mechanism distinction kept
explicit — plus a Discussion block with five CL impact points (replay-
ratio theory testable against published sweeps; a failure theory for
generative replay; pre-merge interference prediction with a mechanism; a
consolidate-vs-modular decision rule; tail monitoring, engaging the
latent-vs-extinct objection). The scan verified the bridge is open: no
prior work carries pop-gen formalism into CL. (3) Downplaying replaced
by convergence framing: the diagnosis was reached independently and is
corroborated by parallel arrivals (Riis; Benati; Yoon; and Crutchfield &
Whalen 2012, pre-deep-learning) — cited for priority of publication, the
full arc owned as one framework. References 30 -> 65; Significance
carries the CL frame; 20-pp rebuild; 151 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
Giorgio Gilestro 2026-09-06 20:15:40 +01:00
parent 6b5591c92f
commit e1532abb4b
5 changed files with 202 additions and 36 deletions

View file

@ -12,10 +12,11 @@ on their own output, model lineages degenerate — a process already recognised
genetic drift. This paper imports the other half of population genetics: the biology of sexual
reproduction. It treats model merging as recombination, real data as immigration, and merge failure as
reproductive isolation, and tests each correspondence in simulations, small neural networks, and
language models. The framework yields design rules — when to average models, when to keep them
separate, how much real data suffices — and a controlled small-model test in which pre-merge
functional disagreement predicted merge damage, motivating further comparison with weight-space
measures.
language models. The framework recasts machine learning's oldest problem — continual learning without
forgetting — at the population scale, and yields design rules: when to average models, when to keep
them separate, how much real data suffices (a theory for the field's empirical replay fractions), and
a controlled small-model test in which pre-merge functional disagreement predicted merge damage,
motivating further comparison with weight-space measures.
## Abstract
@ -23,7 +24,10 @@ AI development increasingly resembles a population process: models are specialis
output, and recombined by weight merging, with an openly evolutionary vocabulary but little use of
evolutionary theory. Here we treat multigenerational model populations as systems whose inheritance,
diversity, and compatibility must be managed, and transfer the quantitative apparatus of the evolution
of sex. We take as settled that training on model output is genetic drift (model collapse). In a
of sex. Its entry point — training on model output is genetic drift, and model collapse is its
signature — we developed independently, and parallel work has now formalised the same diagnosis from
several directions, a convergence we read as evidence for the frame rather than as a shared discovery
to be divided. In a
minimal inheritance model that is exactly WrightFisher — and measurably WrightFisher-plus-bias in
trained networks — we derive and test the remedies: grounding as immigration, where a real-data
fraction far below one retained most equilibrium diversity in the tested settings, with a
@ -44,27 +48,71 @@ association. We state precisely what is exact, what is measured, and what remain
## Introduction
The unit of AI progress is quietly changing. Multi-agent systems arrange many models across *space*
specialists cooperating on a task. A newer axis is *time*: populations of models that persist across
generations, each new model built from older ones — specialised by fine-tuning, trained on data earlier
models generated, and, increasingly, produced by **model merging**, the direct combination of trained
weights (1, 2). The engineering literature describes this openly in evolutionary vocabulary —
"crossover," "mutation," "mate choice," populations of merging models that climb benchmarks (25) —
but as metaphor over search algorithms. The organising claim of this paper is that the vocabulary
deserves its mathematics: **multigenerational model populations are systems whose inheritance,
diversity, and compatibility must be managed — not merely collections of models to optimise — and the
branch of biology that studies exactly this problem, the population genetics of the evolution of sex,
transfers as a quantitative framework.**
Machine learning has quietly become a population-scale phenomenon. Public repositories now host
millions of models — Hugging Face alone grew past three million by 2026 — and these are not
independent creations: the overwhelming majority are fine-tunes, distillations, or merges of a small
number of foundation models, forming sprawling family trees whose lineage structure, inherited traits,
and mutation dynamics are already being mapped with explicitly phylogenetic methods (3133).
Reproduction in this population is no longer metaphorical. Weight-space **model merging** — the direct
combination of trained parents into a new model — is mainstream community practice with standard
tooling and thousands of hybrid checkpoints, including leaderboard-topping ones (1, 2, 37, 38), and
the engineering literature describes it in openly evolutionary vocabulary: "crossover," "mutation,"
"mate choice," populations of merging models that climb benchmarks (25).
One half of the transfer is settled and is not our contribution. Training each generation of a model
on the previous generation's output degrades it — *model collapse*: rare capabilities vanish first and
the lineage drifts toward its own most common behaviour (6). That this is the mathematics of **genetic
drift** in a finite population is now established from several directions (79); a closed-form
first-extinction law even places collapse onset at the WrightFisher first-extinction time (8). We cite
this literature as the diagnosis and build on it.
The generations are coupled through data as well as through weights. Successive models increasingly
learn from model output rather than from fresh human experience: frontier alignment pipelines are now
predominantly synthetic — over 98% in documented cases (43, 44) — self-generated instruction data
seeds whole lineages of descendants (5), a large and growing share of the public web is
machine-generated or machine-translated text (35, 36), and the stock of human text is projected to be
exhausted by frontier training within this decade (34). Meanwhile persistent multi-agent systems and
emerging agent economies put many interacting models into sustained contact (3942). A population
whose members inherit from one another, recombine, and retransmit under these conditions is an
evolving population in the technical sense, whatever one thinks of the metaphors. The organising claim
of this paper is that the vocabulary deserves its mathematics: **multigenerational model populations
are systems whose inheritance, diversity, and compatibility must be managed — not merely collections
of models to optimise — and the branch of biology that studies exactly this problem, the population
genetics of the evolution of sex, transfers as a quantitative framework.**
Our contribution is on the remedy side, and we are explicit about what kind of contribution each claim
is, distinguishing **interpretation** (an existing result understood in population-genetic terms),
The frame's entry point is the diagnosis. Training each generation of a model on the previous
generation's output degrades it — *model collapse*: rare capabilities vanish first and the lineage
drifts toward its own most common behaviour (6). That this is the mathematics of **genetic drift** in
a finite population is a conclusion we reached independently in building the present framework, and
one that has been derived in parallel from several other directions (79), including a closed-form
first-extinction law placing collapse onset at the WrightFisher first-extinction time (8) — and that
was anticipated, before deep learning, in an analysis of sequential inference chains as generalised
genetic drift (63). We cite these works for priority of publication on the diagnosis and read the
convergence — independent arrivals at the same population-genetic account by different routes and in
different decades — as corroboration that the frame is the natural one. What none of that parallel work develops, and what this paper is about, is the
structure the diagnosis opens: the full arc from drift through its remedies (immigration,
recombination, selection, population structure) to its limit (reproductive isolation), carried as one
framework from closed forms to trained networks to language models.
Seen from machine learning's own history, the problem this frame addresses is the field's oldest —
**continual learning** — reappearing one level up. Within a single network, sequential learning
overwrites prior knowledge (catastrophic forgetting; 45, 46), and the discipline's remedies are, one
by one, the population operators of this paper in single-model form: **rehearsal and replay** of past
data is grounding's within-lineage counterpart, and the field's empirically settled replay fractions —
on the order of 1% for instruction tuning (53), 5% for weak and 25% for strong distribution shift in
continual pretraining (52) — sit exactly where the minimal model's operational grounding threshold
lies, a correspondence for which the framework supplies the missing theory (equilibrium diversity,
and a per-capability survival law). **Pseudo-rehearsal** — replaying the network's own generated
samples, proposed as a cure in 1995 (47) and revived as generative replay (48) — is precisely this
paper's ungrounded null: immigration from a drifting source, benign for one hop and compounding into
collapse over generations, with verifier-filtering (12, 62) as what converts it back into grounding.
**Parameter isolation** (65, and frozen-base adapters, which forget far less; 54) is the engineered
decorrelation our specialists use; **complementary-learning-systems consolidation** (4951) is our
periodic adapter-into-base merge; the recent turn to **merging as a continual-learning mechanism**
(5558) applies recombination within one lineage over time, where we apply it across lineages; and
the observation that **rare examples and long-tail knowledge are forgotten first** (5961) is
tail-allele extinction observed one model at a time. One distinction is kept explicit throughout:
catastrophic forgetting is largely deterministic interference from shifted training, whereas collapse
is stochastic sampling drift — the phenomena share their victims (the rare) and their remedies, not
their mechanism. To our knowledge, no prior work carries population-genetic formalism into continual
learning itself; that bridge — replay as immigration with a survival law, merging as recombination
with a compatibility criterion, consolidation as the slow store of a two-speed memory — is where this
framework may matter most.
We are explicit about what kind of contribution each claim is, distinguishing **interpretation** (an existing result understood in population-genetic terms),
**explanation** (the transferred mechanism accounts for observations existing accounts leave open),
and **prediction** (the framework forecasts an unmeasured outcome). The paper is strongest on the
first; makes concrete progress on the second — separating merge failures that are coordinate artefacts
@ -350,8 +398,29 @@ tested weight-distance baselines were not; and *do not treat divergence or speci
evidence of incompatibility* — in every regime we tested, what broke merging was conflicting
conventions on shared circuitry, which is the thing to detect.
**What is borrowed and what is ours.** The diagnosis — collapse as drift — is prior art (69), as are
the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
**What this offers continual learning.** Read into the field where these results most directly land:
(i) a first-principles account of the **replay ratio** — the folklore constants (≈1%, 5%, 25%; 52, 53)
acquire an equilibrium theory and a sharper prediction, that the required fraction is set by the
rarest capability one refuses to lose (the `1 e^{m·p}` law), not by average loss — directly
testable against published replay sweeps; (ii) a **failure theory for generative replay**:
self-generated rehearsal is safe for short horizons and compounds into collapse across generations
unless verifier-filtered back into grounding (47, 48, 12, 62); (iii) **pre-merge interference
prediction with a mechanism**: where the current state of the art fits regressions over candidate
metrics (21), the functional-conflict measure arrives at a convergent signal from principle and comes
with an operator prescription — when conflict is high, do not average; route or breed-and-screen;
(iv) a candidate **decision rule for the consolidate-versus-stay-modular question** that currently
splits the field's practice (keep adapters separate vs merge them; 5458): union-preserving operators
where headroom exists, fusion where the base composes, consolidation as the slow-store step; and (v)
**tail monitoring as the leading indicator**: continual-learning evaluation that averages over
capabilities hides exactly the losses that drift theory says come first and, past a threshold, become
irreversible. On that last point we note the standing objection that apparent forgetting can be
skewed task-inference over latent capability rather than erasure (64); our irreversibility results
concern oracle-measured behavioural distributions, and distinguishing latent from extinct capability
at language-model scale is an open and, we think, decisive experiment for both readings.
**What is borrowed and what is ours.** The diagnosis — collapse as drift — was published first by
others and we cite it so (69), while noting the derivations are independent and convergent; prior art
in the strict sense are the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
naive averaging loses to interference-aware or routed merges (1, 19, 20), that model populations can
climb (25), and that merge success admits ML-native predictors (21, 22). Ours is the framework-level
synthesis — inheritance, diversity, and compatibility as managed quantities — together with: the
@ -442,3 +511,38 @@ publication; every figure in this paper regenerates from committed artifacts wit
28. Hu EJ, et al. (2021) LoRA: low-rank adaptation of large language models. arXiv:2106.09685.
29. Sharma E, Roy DM, Dziugaite GK (2024) The non-local model merging problem: permutation symmetries and variance collapse. arXiv:2410.12766.
30. Kozodoi N, Afolabi Z, Butler J (2026) Are we merging the right models? Impact of expert training duration on model merging for LLMs. arXiv:2607.11997.
31. Laufer B, Oderinwale H, Kleinberg J (2025) Anatomy of a machine learning ecosystem: 2 million models on Hugging Face. arXiv:2508.06811.
32. Horwitz E, Shul A, Hoshen Y (2025) Unsupervised model tree heritage recovery. *ICLR*. arXiv:2405.18432.
33. Jiang W, et al. (2024) PeaTMOSS: a dataset and initial analysis of pre-trained models in open-source software. *MSR*. arXiv:2402.00699.
34. Villalobos P, Ho A, Sevilla J, Besiroglu T, Heim L, Hobbhahn M (2024) Position: will we run out of data? Limits of LLM scaling based on human-generated data. *ICML*. arXiv:2211.04325.
35. Thompson B, et al. (2024) A shocking amount of the web is machine translated. *Findings of ACL*. arXiv:2401.05749.
36. Liang W, et al. (2024) Monitoring AI-modified content at scale. *ICML*. arXiv:2403.07183.
37. Goddard C, et al. (2024) Arcee's MergeKit: a toolkit for merging large language models. *EMNLP Industry Track*, 477485. arXiv:2403.13257.
38. Yang E, et al. (2024) Model merging in LLMs, MLLMs, and beyond: methods, theories, applications and opportunities. arXiv:2408.07666.
39. Brinkmann L, et al. (2023) Machine culture. *Nat Hum Behav* 7:18551868.
40. Park JS, O'Brien JC, Cai CJ, et al. (2023) Generative agents: interactive simulacra of human behavior. *UIST*. arXiv:2304.03442.
41. Guo T, Chen X, Wang Y, et al. (2024) Large language model based multi-agents: a survey of progress and challenges. *IJCAI*. arXiv:2402.01680.
42. Tomasev N, Franklin M, Leibo JZ, et al. (2025) Virtual agent economies. arXiv:2509.10147.
43. Adler B, et al. (2024) Nemotron-4 340B technical report. arXiv:2406.11704.
44. Abdin M, et al. (2024) Phi-4 technical report. arXiv:2412.08905.
45. McCloskey M, Cohen NJ (1989) Catastrophic interference in connectionist networks. *Psychol Learn Motiv* 24:109165.
46. French RM (1999) Catastrophic forgetting in connectionist networks. *Trends Cogn Sci* 3:128135.
47. Robins A (1995) Catastrophic forgetting, rehearsal and pseudorehearsal. *Connect Sci* 7:123146.
48. Shin H, Lee JK, Kim J, Kim J (2017) Continual learning with deep generative replay. *NeurIPS*. arXiv:1705.08690.
49. McClelland JL, McNaughton BL, O'Reilly RC (1995) Why there are complementary learning systems in the hippocampus and neocortex. *Psychol Rev* 102:419457.
50. Kumaran D, Hassabis D, McClelland JL (2016) What learning systems do intelligent agents need? *Trends Cogn Sci* 20:512534.
51. Schwarz J, et al. (2018) Progress & Compress: a scalable framework for continual learning. *ICML*.
52. Ibrahim A, et al. (2024) Simple and scalable strategies to continually pre-train large language models. *TMLR*. arXiv:2403.08763.
53. Scialom T, Chakrabarty T, Muresan S (2022) Fine-tuned language models are continual learners. *EMNLP*. arXiv:2205.12393.
54. Biderman D, et al. (2024) LoRA learns less and forgets less. *TMLR*. arXiv:2405.09673.
55. Ilharco G, et al. (2023) Editing models with task arithmetic. *ICLR*. arXiv:2212.04089.
56. Marczak D, et al. (2024) MagMax: leveraging model merging for seamless continual learning. *ECCV*. arXiv:2407.06322.
57. Alexandrov A, et al. (2024) Mitigating catastrophic forgetting in language transfer via model merging. *Findings of EMNLP*. arXiv:2407.08699.
58. Dziadzio S, et al. (2025) How to merge your multimodal models over time? *CVPR*. arXiv:2412.06712.
59. Toneva M, et al. (2019) An empirical study of example forgetting during deep neural network learning. *ICLR*. arXiv:1812.05159.
60. Kandpal N, et al. (2023) Large language models struggle to learn long-tail knowledge. *ICML*.
61. Liu X, et al. (2022) Long-tailed class incremental learning. *ECCV*. arXiv:2210.00266.
62. Feng Y, et al. (2024) Beyond model collapse: scaling up with synthesized data requires verification. arXiv:2406.07515.
63. Crutchfield JP, Whalen S (2012) Structural drift: the population dynamics of sequential learning. *PLoS Comput Biol* 8:e1002510.
64. Kotha S, Springer JM, Raghunathan A (2024) Understanding catastrophic forgetting in language models via implicit inference. *ICLR*.
65. Rusu AA, et al. (2016) Progressive neural networks. arXiv:1606.04671.