introduction rewritten: setting, diagnosis, question — nothing else

The Introduction is halved (1,360 -> 654 words, four paragraphs): the
model-population setting; the data-coupled generations + the thesis
sentence; the drift diagnosis placed in the literature; and the
motivating question (the four operator decisions with no principled
guidance + the continual-learning framing), closing on the value
anticipation without disclosing results. Evicted and rehomed: the
interpretation/explanation/prediction ladder (deleted — its content
lives in the calibrated Results and ledger); the answers-list (deleted —
results belong in Results); the correspondence walk-through (Muller's
ratchet moved to the minimal-model section with its scope clause;
immigration/Fisher-Muller/BDM citations anchored where the concepts are
developed in Results; the Livnat precursor and predictor-delta moved to
the Discussion ledger); the tiers-of-evidence and negative-results-
prominence sentences (deleted). The continual-learning operator mapping
moved into the Discussion block, retitled "Continual learning at the
population scale", deduplicated against its five offers. All 66
references wholesale-renumbered to the new first-appearance order and
the list reordered (invariant verified: in-text order = 1..66 = list).
Main text 4.7k words; 19 pp.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
Giorgio Gilestro 2026-09-07 10:20:00 +01:00
parent 02ff327b58
commit 612be58433
3 changed files with 128 additions and 176 deletions

View file

@ -82,77 +82,17 @@ structure) and of where those mechanisms reach their limits. This paper develops
for model populations: the arc from drift through its remedies to its limit, reproductive isolation,
carried as one framework from closed forms to trained networks to language models.
In machine learning's own terms, the problem this frame addresses is the field's oldest,
*continual learning*, reappearing one level up. Within a single network, sequential learning
overwrites prior knowledge (catastrophic forgetting; 26, 27), and the discipline's remedies are, one
by one, the population operators of this paper in single-model form: *rehearsal and replay* of past
data is grounding's within-lineage counterpart, and the field's empirically settled replay fractions,
on the order of 1% for instruction tuning (28) and 5% to 25% by distribution-shift strength in
continual pretraining (29), sit where the minimal model's operational grounding threshold lies, a
correspondence for which the framework supplies the missing theory (equilibrium diversity, and a
per-capability survival law). *Pseudo-rehearsal*, the replay of the network's own generated
samples, proposed as a cure in 1995 (30) and revived as generative replay (31), is this paper's
ungrounded null: immigration from a drifting source, benign for one hop and compounding into
collapse over generations; verifier-filtering (32, 33) converts it back into grounding.
*Parameter isolation* (34, and frozen-base adapters, which forget far less; 35) is the engineered
decorrelation our specialists use; *complementary-learning-systems consolidation* (3638) is our
periodic adapter-into-base merge; the recent turn to *merging as a continual-learning mechanism*
(3942) applies recombination within one lineage over time, where we apply it across lineages; and
the observation that rare examples and long-tail knowledge are forgotten first (4345) is
tail-allele extinction observed one model at a time. One distinction is kept explicit throughout:
catastrophic forgetting is largely deterministic interference from shifted training, whereas collapse
is stochastic sampling drift; the two phenomena share their victims, the rare, and their remedies, but not their mechanism. To our knowledge, no prior work carries population-genetic formalism into continual
learning itself; that bridge (replay as immigration with a survival law, merging as recombination with a
compatibility criterion, consolidation as the slow store of a two-speed memory) is where this
framework may matter most.
We are explicit about what kind of contribution each claim is, distinguishing *interpretation* (an existing result understood in population-genetic terms),
*explanation* (the transferred mechanism accounts for observations existing accounts leave open),
and *prediction* (the framework forecasts an unmeasured outcome). The paper is strongest on the
first; makes concrete progress on the second (separating merge failures that are coordinate artefacts
from those that are functional); and reports a first, bounded step on the third: a controlled
predictive test in which pre-merge functional-disagreement measures, chosen by the framework,
predicted merge damage on a constructed task grid while the tested weight-geometry baselines showed
no detectable association.
Stated as a problem: an operator of a model population today has no principled answer to four
recurring decisions. How much verified real data does retraining need before a lineage decays?
Will combining two particular models compose their abilities or damage them? Can incompatibility be
detected before paying for a failed merge? And when should specialists be kept separate rather than
consolidated? Current practice answers these with folklore constants and trial-and-error searches.
The framework prices each decision, and several of its answers are not the intuitive ones. Averaging,
the default combining operator, cancels the benefit of multiple parents to first order precisely in
the regime where that benefit matters most, the rare-capability tail. Specialisation and divergence,
widely treated as the threat to mergeability, produced no incompatibility in any regime we tested;
conflicting conventions always did. Weight distance, the field's default compatibility signal, carried
no detectable predictive signal in our controlled test, while a cheap behavioural measure did. And
where the framework's numbers can be checked against settled practice, they land on it: the replay
fractions that continual learning converged on empirically sit at the minimal model's threshold.
The correspondences we develop, summarised in Table 1: single-teacher retraining is *asexual
reproduction*, and the irreversible arm of its decay shares the defining consequence of *Muller's
ratchet* (46): once every copy of a rare capability is gone from all parents and sources, no
recombination can rebuild it, which is why remedies must act before fixation-by-loss (a consequence-
level correspondence: the minimal model lacks the ratchet's recurrent deleterious-mutation mechanism,
so irreversible loss alone does not identify that specific mechanism). Injecting verified real data is
*immigration* from a non-drifting source (32, 47, 48). Model merging is *recombination*, and its
central payoff, a merged model exceeding every parent, is the *FisherMuller effect* (49, 50).
Merging entangled skills courts *outbreeding depression*; screening many candidate merges is
engineered recombination with unusually flexible parent choice and pre-deployment screening (we use
the shorthand *directed sex*); restricting who merges with whom is *population structure*. Merging's hard limit, models too diverged in function to combine, is *reproductive
isolation*, for which the BatesonDobzhanskyMuller theory of incompatibilities (51, 52) supplies the
structure. The nearest precursor to this programme reads sex as an algorithm for mixability in the
theory of computation (53), pre-dating model merging; the model-merging literature itself has strong
empirical operators (4, 54, 55) and emerging merge-success predictors (56, 57), to which our delta is
mechanism: *when and why* failure is coordinate versus functional, and what moves the boundary.
We support the framework at three tiers of evidence, in ascending realism and descending exactness: a
*minimal analytic model* validated against closed forms to a fraction of a percent; *small trained
networks* (MLPs, recurrent networks, an MNIST image generator) where the operators are measured in
real weights; and *language models* (LoRA-specialised Qwen models, 0.5B locally and 7B on a compute
cluster) where the claims are tested as signs under seed replication. Negative results are reported with the same prominence as confirmations; they include the failure of
an internal pre-registered prediction, a null on emergent speciation that bounds the analogy, and the
sensitivity analyses on the predictive test.
The question this paper addresses is what to do with that diagnosis. An operator of a model
population faces recurring decisions for which there is no principled guidance: how much verified
real data does retraining need before a lineage decays; will combining two particular models compose
their abilities or damage them; can incompatibility be detected before paying for a failed merge; and
when should specialists be kept separate rather than consolidated? In practice these are settled by
convention and by trial-and-error search. They are also, recognisably, machine learning's oldest
problem at a new scale: *continual learning*, the struggle to acquire new abilities without losing old
ones (26, 27), transposed from a single network to a population whose members inherit from one
another. Population genetics, we will argue, prices these decisions. Table 1 summarises the
correspondences on which the argument runs; the sections that follow develop them from closed-form
theory to experiments in trained networks and language models.
## The minimal model, and where its exactness ends
@ -174,7 +114,11 @@ refit) reproduces both. Throughout, a real learner is therefore treated as Wrigh
estimator bias*, and the drift signs (rare-first loss; the grounding response)
survived that bias in every architecture we tested, including a convolutional VAE retrained on its own
generated digits, where the dry lineage collapses to a single blurred digit class while 10% grounding
holds all thirty modes (Fig. 1).
holds all thirty modes (Fig. 1). One consequence of drift deserves its genetic name. Retraining on a
single parent is *asexual reproduction*, and sustained loss under it carries the defining consequence
of *Muller's ratchet* (28): once every copy of a rare capability is gone from all parents and sources,
no recombination can rebuild it, so remedies must act before fixation-by-loss (a consequence-level
correspondence; the minimal model lacks the ratchet's recurrent-mutation driver).
**Table 1.** The dictionary. Each correspondence is stated with the level of support it currently has
(exact = closed form in the minimal model; empirical = measured in trained systems; hypothesis =
@ -197,8 +141,8 @@ known limits is SI Appendix, Table S1.
### Grounding is immigration: cheap, with a floor
In the minimal model, grounding from a fixed real source is immigration into a drifting population,
and the equilibrium diversity has a closed form our simulator matches exactly. That equilibrium is
In the minimal model, grounding from a fixed real source is *immigration* into a drifting population
(2931), and the equilibrium diversity has a closed form our simulator matches exactly. That equilibrium is
*smooth* in the grounding fraction (there is no phase transition in aggregate diversity), so the
practical number is an operational threshold, and we define it as such: under the tested population
size and Zipf source distribution, `g ≈ 0.05` retained most (≥95%) of equilibrium diversity
@ -238,7 +182,7 @@ contrasting union operator (keep each item's strongest source, then renormalise,
redistributes mass and presupposes a verifier or oracle to identify the strongest source) increases
expected retention with K in all regimes in the minimal model. The practically important
operators, *weight averaging* (a nonlinear network's weight-mean does not compute its parents'
output-mean) and *routing among intact specialists* (58) (different storage and inference budgets from a
output-mean) and *routing among intact specialists* (32) (different storage and inference budgets from a
single child), are its empirical cousins, and the measured bridge is a *headroom rule*, stated qualitatively: in language models,
union-preserving operators beat the weight-average where that average falls short of attainable
performance, and add nothing where it does not (easy-versus-hard contrasts at two scales; a
@ -246,15 +190,15 @@ quantitative form of the relationship is untested). On easy tasks a capable base
add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent
and routing wins by a wide margin (Fig. 6AB).
The generative payoff is the FisherMuller effect: recombination assembles, in one offspring,
The generative payoff is the *FisherMuller effect* (33, 34): recombination assembles, in one offspring,
complementary variants that arose in different lineages, producing a genotype fitter than any parent.
In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum, a
genotype no parent held, while the best single parent and the blended average both plateau below
(Fig. 2). In real language models the signature replicates under seed replication: merges of three
LoRA (59) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper
LoRA (35) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper
worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 6A).
Sex has risks and, for AI, an unfair advantage, both quantified on rugged (epistatic) NK landscapes (60)
Sex has risks and, for AI, an unfair advantage, both quantified on rugged (epistatic) NK landscapes (36)
(Fig. 3). When skills are entangled, blind recombination produces offspring *below* their parents
(outbreeding depression), worsening with ruggedness, and the optimal recombination rate shrinks as
entanglement grows. But an engineered population can do what biology cannot: recombine unbounded
@ -283,7 +227,7 @@ the analogue of scoring models by the crowd's approval (the fitness channel). Th
ideas, since both couple the lineage to a non-drifting external signal, but they are different operators,
and we name them separately. In the tested society (a finite agent population on a rugged NK
landscape), a four-arm ablation separates the failure modes: the full system (grounded evaluation +
directed recombination + diversity-preserving selection (61)) climbs to near the global optimum while
directed recombination + diversity-preserving selection (37)) climbs to near the global optimum while
keeping its specialists; removing grounded evaluation converges the population confidently on an unfit
consensus (self-consumption); removing recombination strands it on local optima; removing diversity
converges it prematurely to a worse answer. Each removal fails differently; the three implementations
@ -296,18 +240,18 @@ language-model scale this composed loop remains unbuilt; it is the paper's large
### The limit of sex: model speciation
Recombination presupposes compatible parents. In biology, lineages pushed far enough apart become
separate species (reproductive isolation) through BatesonDobzhanskyMuller incompatibilities:
separate species (*reproductive isolation*) through BatesonDobzhanskyMuller incompatibilities (38, 39):
changes harmless on their own background but deleterious in combination. A merged model is exactly the
exposed hybrid. We built the analytic model (Fig. 5A): hybrid fitness tracks the parents while
compatible, then peels off and crashes below the ancestor; the isolation cliff arrives earlier the
denser the incompatibilities; and the incompatibility *count* snowballs quadratically with divergence
(52). We note that a super-linear count does not by itself entail a sharp performance cliff without
(39). We note that a super-linear count does not by itself entail a sharp performance cliff without
the count-to-effect-size link, which the analytic model supplies under its assumptions and any neural
test must establish separately.
In trained networks, the claim must survive a known alternative: merge barriers between independently
trained networks are famously *coordinate artefacts*, removable by re-aligning hidden units (62);
richer symmetry groups remove more (63), with known failures beyond the shared-data regime (64). We therefore aligned under the composition of
trained networks are famously *coordinate artefacts*, removable by re-aligning hidden units (40);
richer symmetry groups remove more (41), with known failures beyond the shared-data regime (42). We therefore aligned under the composition of
permutation matching and exact per-unit rescaling (the unit symmetry group of plain ReLU MLPs, as the
search space) and decomposed the barrier (Fig. 5 C and D): two networks trained from different
initialisations on the *same* task have a barrier that this alignment removes essentially entirely
@ -330,7 +274,7 @@ catastrophically-forgetting specialists (parents ≈ 0.50, merge ≈ 0.955, a su
rescue). The same double result appears at the language-model tier (Fig. 5 E and F): conflicting conventions
produce *function-specific* hybrid breakdown (the merge scores below both parents on the conflicted
function, while a budget-controlled design shows the disjoint skills merge unharmed), and over-training
disjoint specialists 1→12 epochs (cf. the merging literature's expert-duration effect; 65) produces
disjoint specialists 1→12 epochs (cf. the merging literature's expert-duration effect; 43) produces
no isolation at all — the merge improves. Across every tier
tested, isolation had to be provoked by functional conflict; specialisation alone did not speciate
— a bound on the analogy that sharpens the design rule: what breaks merging is conflicting conventions
@ -350,7 +294,7 @@ divergence with zero conflict). Before merging, six predictors are computed: *co
functional conflict* (bilateral confident disagreement on probes drawn blind to where conflict lives —
a proposed proxy for merge-relevant interactions, motivated by the observation that raw disagreement
counts harmless complementation, one parent merely ignorant, as conflict), raw disagreement, gradient
alignment at the shared base (56), LoRA-delta cosine and distance, and a cross-task performance
alignment at the shared base (44), LoRA-delta cosine and distance, and a cross-task performance
baseline. The pre-registered outcome is the merge penalty against oracle parent potential (the
hybrid-load analogue), also reported against best- and mean-parent references because the predictor
ordering is sensitive to that choice.
@ -407,15 +351,29 @@ tested weight-distance baselines were not; and *do not treat divergence or speci
evidence of incompatibility* — in every regime we tested, what broke merging was conflicting
conventions on shared circuitry, which is the thing to detect.
**What this offers continual learning.** Read into the field where these results most directly land:
(i) a first-principles account of the *replay ratio*: the field's constants (≈1%, 5%, 25%; 28, 29)
acquire an equilibrium theory and a sharper prediction, that the required fraction is set by the
rarest capability one refuses to lose (the `1 e^{m·p}` law) rather than by average loss, which is
testable against published replay sweeps; (ii) a *failure theory for generative replay*:
**Continual learning at the population scale.** Within a single network, the discipline's remedies
for forgetting are this framework's operators writ small. Rehearsal and replay of stored data (26, 27)
is grounded inheritance within one lineage, and the replay fractions the field settled on empirically,
on the order of 1% for instruction tuning (45) and 5% to 25% by distribution-shift strength in
continual pretraining (46), sit where the minimal model's operational threshold lies.
*Pseudo-rehearsal*, the replay of a network's own generated samples, proposed as a cure in 1995 (47)
and revived as generative replay (48), is precisely the ungrounded null studied here: immigration from
a drifting source, benign for one hop, compounding over generations, with verifier-filtering (29, 49)
converting it back into grounding. Parameter isolation (50), including frozen-base adapters, which forget far
less (51), is engineered decorrelation; complementary-learning-systems consolidation (5254) is the periodic
adapter-into-base merge; the recent turn to merging as a continual-learning mechanism (5558) applies
recombination within one lineage over time, where this paper applies it across lineages; and the
observation that rare examples and long-tail knowledge are forgotten first (5961) is tail extinction
seen one model at a time. The mechanisms differ (forgetting is largely deterministic interference,
collapse is sampling drift) but the victims and the remedies coincide, and to our knowledge no prior
work carries population-genetic formalism into continual learning. Read into that field, the results
offer: (i) an equilibrium theory for the replay ratio, with the sharper prediction that the required
fraction is set by the rarest capability one refuses to lose (the `1 e^{m·p}` law) rather than by
average loss, testable against published replay sweeps; (ii) a *failure theory for generative replay*:
self-generated rehearsal is safe for short horizons and compounds into collapse across generations
unless verifier-filtered back into grounding (3033); (iii) *pre-merge interference
unless verifier-filtered back into grounding (29, 4749); (iii) *pre-merge interference
prediction with a mechanism*: where the current state of the art fits regressions over candidate
metrics (56), the functional-conflict measure arrives at a convergent signal from principle and comes
metrics (44), the functional-conflict measure arrives at a convergent signal from principle and comes
with an operator prescription — when conflict is high, do not average; route or breed-and-screen;
(iv) a candidate *decision rule for the consolidate-versus-stay-modular question* that currently
splits the field's practice (keep adapters separate vs merge them; 5458): union-preserving operators
@ -423,14 +381,16 @@ where headroom exists, fusion where the base composes, consolidation as the slow
*tail monitoring as the leading indicator*: continual-learning evaluation that averages over
capabilities hides exactly the losses that drift theory says come first and, past a threshold, become
irreversible. On that last point we note the standing objection that apparent forgetting can be
skewed task-inference over latent capability rather than erasure (66); our irreversibility results
skewed task-inference over latent capability rather than erasure (62); our irreversibility results
concern oracle-measured behavioural distributions, and distinguishing latent from extinct capability
at language-model scale is an open experiment whose outcome would be decisive for both readings.
**What is borrowed and what is ours.** The collapse-as-drift diagnosis is established prior work
(2125); so are the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
naive averaging loses to interference-aware or routed merges (4, 54, 55), that model populations can
climb (5, 810), and that merge success admits ML-native predictors (56, 57). Ours is the framework-level
naive averaging loses to interference-aware or routed merges (4, 63, 64), that model populations can
climb (5, 810), and that merge success admits ML-native predictors (44, 65), correlational where this framework
supplies mechanism; the reading of sex as an algorithm for mixability in the theory of computation
(66) anticipated the transfer before model merging existed. Ours is the framework-level
synthesis — inheritance, diversity, and compatibility as managed quantities — together with: the
conservation law for blending inheritance and its operator boundaries; the per-item grounding floor;
the society ablation with its complementary failure modes; model speciation as a named, tested question, with the
@ -463,7 +423,7 @@ our tested regimes found, freely recombinable in the absence of conflicting conv
ecosystem evolves as one interbreeding population, and the levers that matter are grounding budgets
priced per rare capability and diversity preserved deliberately. If instead long-horizon
specialisation at scale begins to produce emergent incompatibility, as the expert-training-duration
observations hint (65) and our small-scale null does not rule out, then lineages will begin to
observations hint (43) and our small-scale null does not rule out, then lineages will begin to
speciate, and the ecosystem's future is a set of diverging species connected by routing rather than by
merging. Which of the two it will be is measurable now, with the pre-merge conflict instruments this
paper tested.
@ -531,42 +491,42 @@ publication; every figure in this paper regenerates from committed artifacts wit
25. Y. Yoon, D. Hu, I. Weissburg, Y. Qin, H. Jeong, Model collapse in the self-consuming chain of diffusion finetuning: A novel perspective from quantitative trait modeling. *Int. Conf. Learn. Represent.* (2025). https://doi.org/10.48550/arXiv.2407.17493.
26. M. McCloskey, N. J. Cohen, Catastrophic interference in connectionist networks: The sequential learning problem. *Psychol. Learn. Motiv.* **24**, 109165 (1989).
27. R. M. French, Catastrophic forgetting in connectionist networks. *Trends Cogn. Sci.* **3**, 128135 (1999).
28. T. Scialom, T. Chakrabarty, S. Muresan, Fine-tuned language models are continual learners. *Proc. Conf. Empir. Methods Nat. Lang. Process.* (2022). https://doi.org/10.48550/arXiv.2205.12393.
29. A. Ibrahim, et al., Simple and scalable strategies to continually pre-train large language models. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2403.08763.
30. A. Robins, Catastrophic forgetting, rehearsal and pseudorehearsal. *Connect. Sci.* **7**, 123146 (1995).
31. H. Shin, J. K. Lee, J. Kim, J. Kim, Continual learning with deep generative replay. *Adv. Neural Inf. Process. Syst.* **30** (2017). https://doi.org/10.48550/arXiv.1705.08690.
32. B. Yi, Q. Liu, Y. Cheng, H. Xu, Escaping model collapse via synthetic data verification. arXiv [Preprint] (2025). https://doi.org/10.48550/arXiv.2510.16657.
33. Y. Feng, et al., Beyond model collapse: Scaling up with synthesized data requires verification. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2406.07515.
34. A. A. Rusu, et al., Progressive neural networks. arXiv [Preprint] (2016). https://doi.org/10.48550/arXiv.1606.04671.
35. D. Biderman, et al., LoRA learns less and forgets less. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2405.09673.
36. J. L. McClelland, B. L. McNaughton, R. C. O'Reilly, Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory. *Psychol. Rev.* **102**, 419457 (1995).
37. D. Kumaran, D. Hassabis, J. L. McClelland, What learning systems do intelligent agents need? Complementary learning systems theory updated. *Trends Cogn. Sci.* **20**, 512534 (2016).
38. J. Schwarz, et al., Progress & Compress: A scalable framework for continual learning. *Proc. Int. Conf. Mach. Learn.* (2018).
39. G. Ilharco, et al., Editing models with task arithmetic. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2212.04089.
40. D. Marczak, B. Twardowski, T. Trzciński, S. Cygert, MagMax: Leveraging model merging for seamless continual learning. *Proc. Eur. Conf. Comput. Vis.* (2024). https://doi.org/10.48550/arXiv.2407.06322.
41. A. Alexandrov, et al., Mitigating catastrophic forgetting in language transfer via model merging. *Findings Assoc. Comput. Linguist.: EMNLP* (2024). https://doi.org/10.48550/arXiv.2407.08699.
42. S. Dziadzio, et al., How to merge your multimodal models over time? *Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.* (2025). https://doi.org/10.48550/arXiv.2412.06712.
43. M. Toneva, et al., An empirical study of example forgetting during deep neural network learning. *Int. Conf. Learn. Represent.* (2019). https://doi.org/10.48550/arXiv.1812.05159.
44. N. Kandpal, H. Deng, A. Roberts, E. Wallace, C. Raffel, Large language models struggle to learn long-tail knowledge. *Proc. Int. Conf. Mach. Learn.* (2023). https://doi.org/10.48550/arXiv.2211.08411.
45. X. Liu, et al., Long-tailed class incremental learning. *Proc. Eur. Conf. Comput. Vis.* (2022). https://doi.org/10.48550/arXiv.2210.00266.
46. H. J. Muller, The relation of recombination to mutational advance. *Mutat. Res.* **1**, 29 (1964).
47. M. Gerstgrasser, et al., Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. *Conf. Lang. Model.* (2024). https://doi.org/10.48550/arXiv.2404.01413.
48. S. Wright, Evolution in Mendelian populations. *Genetics* **16**, 97159 (1931).
49. R. A. Fisher, *The Genetical Theory of Natural Selection* (Clarendon Press, 1930).
50. H. J. Muller, Some genetic aspects of sex. *Am. Nat.* **66**, 118138 (1932).
51. H. A. Orr, The population genetics of speciation: The evolution of hybrid incompatibilities. *Genetics* **139**, 18051813 (1995).
52. H. A. Orr, M. Turelli, The evolution of postzygotic isolation: Accumulating DobzhanskyMuller incompatibilities. *Evolution* **55**, 10851094 (2001).
53. A. Livnat, C. Papadimitriou, Sex as an algorithm: The theory of evolution under the lens of computation. *Commun. ACM* **59**, 8493 (2016).
54. L. Yu, B. Yu, H. Yu, F. Huang, Y. Li, Language models are super Mario: Absorbing abilities from homologous models as a free lunch. *Proc. Int. Conf. Mach. Learn.* (2024). https://doi.org/10.48550/arXiv.2311.03099.
55. M. Wortsman, et al., Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. *Proc. Int. Conf. Mach. Learn.* (2022). https://doi.org/10.48550/arXiv.2203.05482.
56. L. Zhou, B. Zhao, R. Yu, E. Rodolà, Demystifying mergeability: Interpretable properties to predict model merging success. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2601.22285.
57. Y. Cao, et al., An empirical study and theoretical explanation on task-level model-merging collapse. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2603.09463.
58. J. Pari, S. Jelassi, P. Agrawal, Collective model intelligence requires compatible specialization. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2411.02207.
59. E. J. Hu, et al., LoRA: Low-rank adaptation of large language models. *Int. Conf. Learn. Represent.* (2022). https://doi.org/10.48550/arXiv.2106.09685.
60. S. A. Kauffman, S. Levin, Towards a general theory of adaptive walks on rugged landscapes. *J. Theor. Biol.* **128**, 1145 (1987).
61. J. Lehman, K. O. Stanley, Abandoning objectives: Evolution through the search for novelty alone. *Evol. Comput.* **19**, 189223 (2011).
62. S. K. Ainsworth, J. Hayase, S. Srinivasa, Git Re-Basin: Merging models modulo permutation symmetries. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2209.04836.
63. T. Li, Z. Shen, Scaling linear mode connectivity and merging to billion-parameter pretrained transformers. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2606.23607.
64. E. Sharma, D. M. Roy, G. K. Dziugaite, The non-local model merging problem: Permutation symmetries and variance collapse. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2410.12766.
65. N. Kozodoi, Z. Afolabi, J. Butler, Are we merging the right models? Impact of expert training duration on model merging for LLMs. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2607.11997.
66. S. Kotha, J. M. Springer, A. Raghunathan, Understanding catastrophic forgetting in language models via implicit inference. *Int. Conf. Learn. Represent.* (2024). https://doi.org/10.48550/arXiv.2309.10105.
28. H. J. Muller, The relation of recombination to mutational advance. *Mutat. Res.* **1**, 29 (1964).
29. B. Yi, Q. Liu, Y. Cheng, H. Xu, Escaping model collapse via synthetic data verification. arXiv [Preprint] (2025). https://doi.org/10.48550/arXiv.2510.16657.
30. M. Gerstgrasser, et al., Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. *Conf. Lang. Model.* (2024). https://doi.org/10.48550/arXiv.2404.01413.
31. S. Wright, Evolution in Mendelian populations. *Genetics* **16**, 97159 (1931).
32. J. Pari, S. Jelassi, P. Agrawal, Collective model intelligence requires compatible specialization. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2411.02207.
33. R. A. Fisher, *The Genetical Theory of Natural Selection* (Clarendon Press, 1930).
34. H. J. Muller, Some genetic aspects of sex. *Am. Nat.* **66**, 118138 (1932).
35. E. J. Hu, et al., LoRA: Low-rank adaptation of large language models. *Int. Conf. Learn. Represent.* (2022). https://doi.org/10.48550/arXiv.2106.09685.
36. S. A. Kauffman, S. Levin, Towards a general theory of adaptive walks on rugged landscapes. *J. Theor. Biol.* **128**, 1145 (1987).
37. J. Lehman, K. O. Stanley, Abandoning objectives: Evolution through the search for novelty alone. *Evol. Comput.* **19**, 189223 (2011).
38. H. A. Orr, The population genetics of speciation: The evolution of hybrid incompatibilities. *Genetics* **139**, 18051813 (1995).
39. H. A. Orr, M. Turelli, The evolution of postzygotic isolation: Accumulating DobzhanskyMuller incompatibilities. *Evolution* **55**, 10851094 (2001).
40. S. K. Ainsworth, J. Hayase, S. Srinivasa, Git Re-Basin: Merging models modulo permutation symmetries. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2209.04836.
41. T. Li, Z. Shen, Scaling linear mode connectivity and merging to billion-parameter pretrained transformers. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2606.23607.
42. E. Sharma, D. M. Roy, G. K. Dziugaite, The non-local model merging problem: Permutation symmetries and variance collapse. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2410.12766.
43. N. Kozodoi, Z. Afolabi, J. Butler, Are we merging the right models? Impact of expert training duration on model merging for LLMs. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2607.11997.
44. L. Zhou, B. Zhao, R. Yu, E. Rodolà, Demystifying mergeability: Interpretable properties to predict model merging success. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2601.22285.
45. T. Scialom, T. Chakrabarty, S. Muresan, Fine-tuned language models are continual learners. *Proc. Conf. Empir. Methods Nat. Lang. Process.* (2022). https://doi.org/10.48550/arXiv.2205.12393.
46. A. Ibrahim, et al., Simple and scalable strategies to continually pre-train large language models. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2403.08763.
47. A. Robins, Catastrophic forgetting, rehearsal and pseudorehearsal. *Connect. Sci.* **7**, 123146 (1995).
48. H. Shin, J. K. Lee, J. Kim, J. Kim, Continual learning with deep generative replay. *Adv. Neural Inf. Process. Syst.* **30** (2017). https://doi.org/10.48550/arXiv.1705.08690.
49. Y. Feng, et al., Beyond model collapse: Scaling up with synthesized data requires verification. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2406.07515.
50. A. A. Rusu, et al., Progressive neural networks. arXiv [Preprint] (2016). https://doi.org/10.48550/arXiv.1606.04671.
51. D. Biderman, et al., LoRA learns less and forgets less. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2405.09673.
52. J. L. McClelland, B. L. McNaughton, R. C. O'Reilly, Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory. *Psychol. Rev.* **102**, 419457 (1995).
53. D. Kumaran, D. Hassabis, J. L. McClelland, What learning systems do intelligent agents need? Complementary learning systems theory updated. *Trends Cogn. Sci.* **20**, 512534 (2016).
54. J. Schwarz, et al., Progress & Compress: A scalable framework for continual learning. *Proc. Int. Conf. Mach. Learn.* (2018).
55. G. Ilharco, et al., Editing models with task arithmetic. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2212.04089.
56. D. Marczak, B. Twardowski, T. Trzciński, S. Cygert, MagMax: Leveraging model merging for seamless continual learning. *Proc. Eur. Conf. Comput. Vis.* (2024). https://doi.org/10.48550/arXiv.2407.06322.
57. A. Alexandrov, et al., Mitigating catastrophic forgetting in language transfer via model merging. *Findings Assoc. Comput. Linguist.: EMNLP* (2024). https://doi.org/10.48550/arXiv.2407.08699.
58. S. Dziadzio, et al., How to merge your multimodal models over time? *Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.* (2025). https://doi.org/10.48550/arXiv.2412.06712.
59. M. Toneva, et al., An empirical study of example forgetting during deep neural network learning. *Int. Conf. Learn. Represent.* (2019). https://doi.org/10.48550/arXiv.1812.05159.
60. N. Kandpal, H. Deng, A. Roberts, E. Wallace, C. Raffel, Large language models struggle to learn long-tail knowledge. *Proc. Int. Conf. Mach. Learn.* (2023). https://doi.org/10.48550/arXiv.2211.08411.
61. X. Liu, et al., Long-tailed class incremental learning. *Proc. Eur. Conf. Comput. Vis.* (2022). https://doi.org/10.48550/arXiv.2210.00266.
62. S. Kotha, J. M. Springer, A. Raghunathan, Understanding catastrophic forgetting in language models via implicit inference. *Int. Conf. Learn. Represent.* (2024). https://doi.org/10.48550/arXiv.2309.10105.
63. L. Yu, B. Yu, H. Yu, F. Huang, Y. Li, Language models are super Mario: Absorbing abilities from homologous models as a free lunch. *Proc. Int. Conf. Mach. Learn.* (2024). https://doi.org/10.48550/arXiv.2311.03099.
64. M. Wortsman, et al., Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. *Proc. Int. Conf. Mach. Learn.* (2022). https://doi.org/10.48550/arXiv.2203.05482.
65. Y. Cao, et al., An empirical study and theoretical explanation on task-level model-merging collapse. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2603.09463.
66. A. Livnat, C. Papadimitriou, Sex as an algorithm: The theory of evolution under the lens of computation. *Commun. ACM* **59**, 8493 (2016).