introduction rewritten: setting, diagnosis, question — nothing else
The Introduction is halved (1,360 -> 654 words, four paragraphs): the model-population setting; the data-coupled generations + the thesis sentence; the drift diagnosis placed in the literature; and the motivating question (the four operator decisions with no principled guidance + the continual-learning framing), closing on the value anticipation without disclosing results. Evicted and rehomed: the interpretation/explanation/prediction ladder (deleted — its content lives in the calibrated Results and ledger); the answers-list (deleted — results belong in Results); the correspondence walk-through (Muller's ratchet moved to the minimal-model section with its scope clause; immigration/Fisher-Muller/BDM citations anchored where the concepts are developed in Results; the Livnat precursor and predictor-delta moved to the Discussion ledger); the tiers-of-evidence and negative-results- prominence sentences (deleted). The continual-learning operator mapping moved into the Discussion block, retitled "Continual learning at the population scale", deduplicated against its five offers. All 66 references wholesale-renumbered to the new first-appearance order and the list reordered (invariant verified: in-text order = 1..66 = list). Main text 4.7k words; 19 pp. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
parent
02ff327b58
commit
612be58433
3 changed files with 128 additions and 176 deletions
|
|
@ -82,77 +82,17 @@ structure) and of where those mechanisms reach their limits. This paper develops
|
|||
for model populations: the arc from drift through its remedies to its limit, reproductive isolation,
|
||||
carried as one framework from closed forms to trained networks to language models.
|
||||
|
||||
In machine learning's own terms, the problem this frame addresses is the field's oldest,
|
||||
*continual learning*, reappearing one level up. Within a single network, sequential learning
|
||||
overwrites prior knowledge (catastrophic forgetting; 26, 27), and the discipline's remedies are, one
|
||||
by one, the population operators of this paper in single-model form: *rehearsal and replay* of past
|
||||
data is grounding's within-lineage counterpart, and the field's empirically settled replay fractions,
|
||||
on the order of 1% for instruction tuning (28) and 5% to 25% by distribution-shift strength in
|
||||
continual pretraining (29), sit where the minimal model's operational grounding threshold lies, a
|
||||
correspondence for which the framework supplies the missing theory (equilibrium diversity, and a
|
||||
per-capability survival law). *Pseudo-rehearsal*, the replay of the network's own generated
|
||||
samples, proposed as a cure in 1995 (30) and revived as generative replay (31), is this paper's
|
||||
ungrounded null: immigration from a drifting source, benign for one hop and compounding into
|
||||
collapse over generations; verifier-filtering (32, 33) converts it back into grounding.
|
||||
*Parameter isolation* (34, and frozen-base adapters, which forget far less; 35) is the engineered
|
||||
decorrelation our specialists use; *complementary-learning-systems consolidation* (36–38) is our
|
||||
periodic adapter-into-base merge; the recent turn to *merging as a continual-learning mechanism*
|
||||
(39–42) applies recombination within one lineage over time, where we apply it across lineages; and
|
||||
the observation that rare examples and long-tail knowledge are forgotten first (43–45) is
|
||||
tail-allele extinction observed one model at a time. One distinction is kept explicit throughout:
|
||||
catastrophic forgetting is largely deterministic interference from shifted training, whereas collapse
|
||||
is stochastic sampling drift; the two phenomena share their victims, the rare, and their remedies, but not their mechanism. To our knowledge, no prior work carries population-genetic formalism into continual
|
||||
learning itself; that bridge (replay as immigration with a survival law, merging as recombination with a
|
||||
compatibility criterion, consolidation as the slow store of a two-speed memory) is where this
|
||||
framework may matter most.
|
||||
|
||||
We are explicit about what kind of contribution each claim is, distinguishing *interpretation* (an existing result understood in population-genetic terms),
|
||||
*explanation* (the transferred mechanism accounts for observations existing accounts leave open),
|
||||
and *prediction* (the framework forecasts an unmeasured outcome). The paper is strongest on the
|
||||
first; makes concrete progress on the second (separating merge failures that are coordinate artefacts
|
||||
from those that are functional); and reports a first, bounded step on the third: a controlled
|
||||
predictive test in which pre-merge functional-disagreement measures, chosen by the framework,
|
||||
predicted merge damage on a constructed task grid while the tested weight-geometry baselines showed
|
||||
no detectable association.
|
||||
|
||||
Stated as a problem: an operator of a model population today has no principled answer to four
|
||||
recurring decisions. How much verified real data does retraining need before a lineage decays?
|
||||
Will combining two particular models compose their abilities or damage them? Can incompatibility be
|
||||
detected before paying for a failed merge? And when should specialists be kept separate rather than
|
||||
consolidated? Current practice answers these with folklore constants and trial-and-error searches.
|
||||
The framework prices each decision, and several of its answers are not the intuitive ones. Averaging,
|
||||
the default combining operator, cancels the benefit of multiple parents to first order precisely in
|
||||
the regime where that benefit matters most, the rare-capability tail. Specialisation and divergence,
|
||||
widely treated as the threat to mergeability, produced no incompatibility in any regime we tested;
|
||||
conflicting conventions always did. Weight distance, the field's default compatibility signal, carried
|
||||
no detectable predictive signal in our controlled test, while a cheap behavioural measure did. And
|
||||
where the framework's numbers can be checked against settled practice, they land on it: the replay
|
||||
fractions that continual learning converged on empirically sit at the minimal model's threshold.
|
||||
|
||||
The correspondences we develop, summarised in Table 1: single-teacher retraining is *asexual
|
||||
reproduction*, and the irreversible arm of its decay shares the defining consequence of *Muller's
|
||||
ratchet* (46): once every copy of a rare capability is gone from all parents and sources, no
|
||||
recombination can rebuild it, which is why remedies must act before fixation-by-loss (a consequence-
|
||||
level correspondence: the minimal model lacks the ratchet's recurrent deleterious-mutation mechanism,
|
||||
so irreversible loss alone does not identify that specific mechanism). Injecting verified real data is
|
||||
*immigration* from a non-drifting source (32, 47, 48). Model merging is *recombination*, and its
|
||||
central payoff, a merged model exceeding every parent, is the *Fisher–Muller effect* (49, 50).
|
||||
Merging entangled skills courts *outbreeding depression*; screening many candidate merges is
|
||||
engineered recombination with unusually flexible parent choice and pre-deployment screening (we use
|
||||
the shorthand *directed sex*); restricting who merges with whom is *population structure*. Merging's hard limit, models too diverged in function to combine, is *reproductive
|
||||
isolation*, for which the Bateson–Dobzhansky–Muller theory of incompatibilities (51, 52) supplies the
|
||||
structure. The nearest precursor to this programme reads sex as an algorithm for mixability in the
|
||||
theory of computation (53), pre-dating model merging; the model-merging literature itself has strong
|
||||
empirical operators (4, 54, 55) and emerging merge-success predictors (56, 57), to which our delta is
|
||||
mechanism: *when and why* failure is coordinate versus functional, and what moves the boundary.
|
||||
|
||||
We support the framework at three tiers of evidence, in ascending realism and descending exactness: a
|
||||
*minimal analytic model* validated against closed forms to a fraction of a percent; *small trained
|
||||
networks* (MLPs, recurrent networks, an MNIST image generator) where the operators are measured in
|
||||
real weights; and *language models* (LoRA-specialised Qwen models, 0.5B locally and 7B on a compute
|
||||
cluster) where the claims are tested as signs under seed replication. Negative results are reported with the same prominence as confirmations; they include the failure of
|
||||
an internal pre-registered prediction, a null on emergent speciation that bounds the analogy, and the
|
||||
sensitivity analyses on the predictive test.
|
||||
The question this paper addresses is what to do with that diagnosis. An operator of a model
|
||||
population faces recurring decisions for which there is no principled guidance: how much verified
|
||||
real data does retraining need before a lineage decays; will combining two particular models compose
|
||||
their abilities or damage them; can incompatibility be detected before paying for a failed merge; and
|
||||
when should specialists be kept separate rather than consolidated? In practice these are settled by
|
||||
convention and by trial-and-error search. They are also, recognisably, machine learning's oldest
|
||||
problem at a new scale: *continual learning*, the struggle to acquire new abilities without losing old
|
||||
ones (26, 27), transposed from a single network to a population whose members inherit from one
|
||||
another. Population genetics, we will argue, prices these decisions. Table 1 summarises the
|
||||
correspondences on which the argument runs; the sections that follow develop them from closed-form
|
||||
theory to experiments in trained networks and language models.
|
||||
|
||||
## The minimal model, and where its exactness ends
|
||||
|
||||
|
|
@ -174,7 +114,11 @@ refit) reproduces both. Throughout, a real learner is therefore treated as Wrigh
|
|||
estimator bias*, and the drift signs (rare-first loss; the grounding response)
|
||||
survived that bias in every architecture we tested, including a convolutional VAE retrained on its own
|
||||
generated digits, where the dry lineage collapses to a single blurred digit class while 10% grounding
|
||||
holds all thirty modes (Fig. 1).
|
||||
holds all thirty modes (Fig. 1). One consequence of drift deserves its genetic name. Retraining on a
|
||||
single parent is *asexual reproduction*, and sustained loss under it carries the defining consequence
|
||||
of *Muller's ratchet* (28): once every copy of a rare capability is gone from all parents and sources,
|
||||
no recombination can rebuild it, so remedies must act before fixation-by-loss (a consequence-level
|
||||
correspondence; the minimal model lacks the ratchet's recurrent-mutation driver).
|
||||
|
||||
**Table 1.** The dictionary. Each correspondence is stated with the level of support it currently has
|
||||
(exact = closed form in the minimal model; empirical = measured in trained systems; hypothesis =
|
||||
|
|
@ -197,8 +141,8 @@ known limits is SI Appendix, Table S1.
|
|||
|
||||
### Grounding is immigration: cheap, with a floor
|
||||
|
||||
In the minimal model, grounding from a fixed real source is immigration into a drifting population,
|
||||
and the equilibrium diversity has a closed form our simulator matches exactly. That equilibrium is
|
||||
In the minimal model, grounding from a fixed real source is *immigration* into a drifting population
|
||||
(29–31), and the equilibrium diversity has a closed form our simulator matches exactly. That equilibrium is
|
||||
*smooth* in the grounding fraction (there is no phase transition in aggregate diversity), so the
|
||||
practical number is an operational threshold, and we define it as such: under the tested population
|
||||
size and Zipf source distribution, `g ≈ 0.05` retained most (≥95%) of equilibrium diversity
|
||||
|
|
@ -238,7 +182,7 @@ contrasting union operator (keep each item's strongest source, then renormalise,
|
|||
redistributes mass and presupposes a verifier or oracle to identify the strongest source) increases
|
||||
expected retention with K in all regimes in the minimal model. The practically important
|
||||
operators, *weight averaging* (a nonlinear network's weight-mean does not compute its parents'
|
||||
output-mean) and *routing among intact specialists* (58) (different storage and inference budgets from a
|
||||
output-mean) and *routing among intact specialists* (32) (different storage and inference budgets from a
|
||||
single child), are its empirical cousins, and the measured bridge is a *headroom rule*, stated qualitatively: in language models,
|
||||
union-preserving operators beat the weight-average where that average falls short of attainable
|
||||
performance, and add nothing where it does not (easy-versus-hard contrasts at two scales; a
|
||||
|
|
@ -246,15 +190,15 @@ quantitative form of the relationship is untested). On easy tasks a capable base
|
|||
add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent
|
||||
and routing wins by a wide margin (Fig. 6A–B).
|
||||
|
||||
The generative payoff is the Fisher–Muller effect: recombination assembles, in one offspring,
|
||||
The generative payoff is the *Fisher–Muller effect* (33, 34): recombination assembles, in one offspring,
|
||||
complementary variants that arose in different lineages, producing a genotype fitter than any parent.
|
||||
In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum, a
|
||||
genotype no parent held, while the best single parent and the blended average both plateau below
|
||||
(Fig. 2). In real language models the signature replicates under seed replication: merges of three
|
||||
LoRA (59) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper
|
||||
LoRA (35) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper
|
||||
worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 6A).
|
||||
|
||||
Sex has risks and, for AI, an unfair advantage, both quantified on rugged (epistatic) NK landscapes (60)
|
||||
Sex has risks and, for AI, an unfair advantage, both quantified on rugged (epistatic) NK landscapes (36)
|
||||
(Fig. 3). When skills are entangled, blind recombination produces offspring *below* their parents
|
||||
(outbreeding depression), worsening with ruggedness, and the optimal recombination rate shrinks as
|
||||
entanglement grows. But an engineered population can do what biology cannot: recombine unbounded
|
||||
|
|
@ -283,7 +227,7 @@ the analogue of scoring models by the crowd's approval (the fitness channel). Th
|
|||
ideas, since both couple the lineage to a non-drifting external signal, but they are different operators,
|
||||
and we name them separately. In the tested society (a finite agent population on a rugged NK
|
||||
landscape), a four-arm ablation separates the failure modes: the full system (grounded evaluation +
|
||||
directed recombination + diversity-preserving selection (61)) climbs to near the global optimum while
|
||||
directed recombination + diversity-preserving selection (37)) climbs to near the global optimum while
|
||||
keeping its specialists; removing grounded evaluation converges the population confidently on an unfit
|
||||
consensus (self-consumption); removing recombination strands it on local optima; removing diversity
|
||||
converges it prematurely to a worse answer. Each removal fails differently; the three implementations
|
||||
|
|
@ -296,18 +240,18 @@ language-model scale this composed loop remains unbuilt; it is the paper's large
|
|||
### The limit of sex: model speciation
|
||||
|
||||
Recombination presupposes compatible parents. In biology, lineages pushed far enough apart become
|
||||
separate species (reproductive isolation) through Bateson–Dobzhansky–Muller incompatibilities:
|
||||
separate species (*reproductive isolation*) through Bateson–Dobzhansky–Muller incompatibilities (38, 39):
|
||||
changes harmless on their own background but deleterious in combination. A merged model is exactly the
|
||||
exposed hybrid. We built the analytic model (Fig. 5A): hybrid fitness tracks the parents while
|
||||
compatible, then peels off and crashes below the ancestor; the isolation cliff arrives earlier the
|
||||
denser the incompatibilities; and the incompatibility *count* snowballs quadratically with divergence
|
||||
(52). We note that a super-linear count does not by itself entail a sharp performance cliff without
|
||||
(39). We note that a super-linear count does not by itself entail a sharp performance cliff without
|
||||
the count-to-effect-size link, which the analytic model supplies under its assumptions and any neural
|
||||
test must establish separately.
|
||||
|
||||
In trained networks, the claim must survive a known alternative: merge barriers between independently
|
||||
trained networks are famously *coordinate artefacts*, removable by re-aligning hidden units (62);
|
||||
richer symmetry groups remove more (63), with known failures beyond the shared-data regime (64). We therefore aligned under the composition of
|
||||
trained networks are famously *coordinate artefacts*, removable by re-aligning hidden units (40);
|
||||
richer symmetry groups remove more (41), with known failures beyond the shared-data regime (42). We therefore aligned under the composition of
|
||||
permutation matching and exact per-unit rescaling (the unit symmetry group of plain ReLU MLPs, as the
|
||||
search space) and decomposed the barrier (Fig. 5 C and D): two networks trained from different
|
||||
initialisations on the *same* task have a barrier that this alignment removes essentially entirely
|
||||
|
|
@ -330,7 +274,7 @@ catastrophically-forgetting specialists (parents ≈ 0.50, merge ≈ 0.955, a su
|
|||
rescue). The same double result appears at the language-model tier (Fig. 5 E and F): conflicting conventions
|
||||
produce *function-specific* hybrid breakdown (the merge scores below both parents on the conflicted
|
||||
function, while a budget-controlled design shows the disjoint skills merge unharmed), and over-training
|
||||
disjoint specialists 1→12 epochs (cf. the merging literature's expert-duration effect; 65) produces
|
||||
disjoint specialists 1→12 epochs (cf. the merging literature's expert-duration effect; 43) produces
|
||||
no isolation at all — the merge improves. Across every tier
|
||||
tested, isolation had to be provoked by functional conflict; specialisation alone did not speciate
|
||||
— a bound on the analogy that sharpens the design rule: what breaks merging is conflicting conventions
|
||||
|
|
@ -350,7 +294,7 @@ divergence with zero conflict). Before merging, six predictors are computed: *co
|
|||
functional conflict* (bilateral confident disagreement on probes drawn blind to where conflict lives —
|
||||
a proposed proxy for merge-relevant interactions, motivated by the observation that raw disagreement
|
||||
counts harmless complementation, one parent merely ignorant, as conflict), raw disagreement, gradient
|
||||
alignment at the shared base (56), LoRA-delta cosine and distance, and a cross-task performance
|
||||
alignment at the shared base (44), LoRA-delta cosine and distance, and a cross-task performance
|
||||
baseline. The pre-registered outcome is the merge penalty against oracle parent potential (the
|
||||
hybrid-load analogue), also reported against best- and mean-parent references because the predictor
|
||||
ordering is sensitive to that choice.
|
||||
|
|
@ -407,15 +351,29 @@ tested weight-distance baselines were not; and *do not treat divergence or speci
|
|||
evidence of incompatibility* — in every regime we tested, what broke merging was conflicting
|
||||
conventions on shared circuitry, which is the thing to detect.
|
||||
|
||||
**What this offers continual learning.** Read into the field where these results most directly land:
|
||||
(i) a first-principles account of the *replay ratio*: the field's constants (≈1%, 5%, 25%; 28, 29)
|
||||
acquire an equilibrium theory and a sharper prediction, that the required fraction is set by the
|
||||
rarest capability one refuses to lose (the `1 − e^{−m·p}` law) rather than by average loss, which is
|
||||
testable against published replay sweeps; (ii) a *failure theory for generative replay*:
|
||||
**Continual learning at the population scale.** Within a single network, the discipline's remedies
|
||||
for forgetting are this framework's operators writ small. Rehearsal and replay of stored data (26, 27)
|
||||
is grounded inheritance within one lineage, and the replay fractions the field settled on empirically,
|
||||
on the order of 1% for instruction tuning (45) and 5% to 25% by distribution-shift strength in
|
||||
continual pretraining (46), sit where the minimal model's operational threshold lies.
|
||||
*Pseudo-rehearsal*, the replay of a network's own generated samples, proposed as a cure in 1995 (47)
|
||||
and revived as generative replay (48), is precisely the ungrounded null studied here: immigration from
|
||||
a drifting source, benign for one hop, compounding over generations, with verifier-filtering (29, 49)
|
||||
converting it back into grounding. Parameter isolation (50), including frozen-base adapters, which forget far
|
||||
less (51), is engineered decorrelation; complementary-learning-systems consolidation (52–54) is the periodic
|
||||
adapter-into-base merge; the recent turn to merging as a continual-learning mechanism (55–58) applies
|
||||
recombination within one lineage over time, where this paper applies it across lineages; and the
|
||||
observation that rare examples and long-tail knowledge are forgotten first (59–61) is tail extinction
|
||||
seen one model at a time. The mechanisms differ (forgetting is largely deterministic interference,
|
||||
collapse is sampling drift) but the victims and the remedies coincide, and to our knowledge no prior
|
||||
work carries population-genetic formalism into continual learning. Read into that field, the results
|
||||
offer: (i) an equilibrium theory for the replay ratio, with the sharper prediction that the required
|
||||
fraction is set by the rarest capability one refuses to lose (the `1 − e^{−m·p}` law) rather than by
|
||||
average loss, testable against published replay sweeps; (ii) a *failure theory for generative replay*:
|
||||
self-generated rehearsal is safe for short horizons and compounds into collapse across generations
|
||||
unless verifier-filtered back into grounding (30–33); (iii) *pre-merge interference
|
||||
unless verifier-filtered back into grounding (29, 47–49); (iii) *pre-merge interference
|
||||
prediction with a mechanism*: where the current state of the art fits regressions over candidate
|
||||
metrics (56), the functional-conflict measure arrives at a convergent signal from principle and comes
|
||||
metrics (44), the functional-conflict measure arrives at a convergent signal from principle and comes
|
||||
with an operator prescription — when conflict is high, do not average; route or breed-and-screen;
|
||||
(iv) a candidate *decision rule for the consolidate-versus-stay-modular question* that currently
|
||||
splits the field's practice (keep adapters separate vs merge them; 54–58): union-preserving operators
|
||||
|
|
@ -423,14 +381,16 @@ where headroom exists, fusion where the base composes, consolidation as the slow
|
|||
*tail monitoring as the leading indicator*: continual-learning evaluation that averages over
|
||||
capabilities hides exactly the losses that drift theory says come first and, past a threshold, become
|
||||
irreversible. On that last point we note the standing objection that apparent forgetting can be
|
||||
skewed task-inference over latent capability rather than erasure (66); our irreversibility results
|
||||
skewed task-inference over latent capability rather than erasure (62); our irreversibility results
|
||||
concern oracle-measured behavioural distributions, and distinguishing latent from extinct capability
|
||||
at language-model scale is an open experiment whose outcome would be decisive for both readings.
|
||||
|
||||
**What is borrowed and what is ours.** The collapse-as-drift diagnosis is established prior work
|
||||
(21–25); so are the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
|
||||
naive averaging loses to interference-aware or routed merges (4, 54, 55), that model populations can
|
||||
climb (5, 8–10), and that merge success admits ML-native predictors (56, 57). Ours is the framework-level
|
||||
naive averaging loses to interference-aware or routed merges (4, 63, 64), that model populations can
|
||||
climb (5, 8–10), and that merge success admits ML-native predictors (44, 65), correlational where this framework
|
||||
supplies mechanism; the reading of sex as an algorithm for mixability in the theory of computation
|
||||
(66) anticipated the transfer before model merging existed. Ours is the framework-level
|
||||
synthesis — inheritance, diversity, and compatibility as managed quantities — together with: the
|
||||
conservation law for blending inheritance and its operator boundaries; the per-item grounding floor;
|
||||
the society ablation with its complementary failure modes; model speciation as a named, tested question, with the
|
||||
|
|
@ -463,7 +423,7 @@ our tested regimes found, freely recombinable in the absence of conflicting conv
|
|||
ecosystem evolves as one interbreeding population, and the levers that matter are grounding budgets
|
||||
priced per rare capability and diversity preserved deliberately. If instead long-horizon
|
||||
specialisation at scale begins to produce emergent incompatibility, as the expert-training-duration
|
||||
observations hint (65) and our small-scale null does not rule out, then lineages will begin to
|
||||
observations hint (43) and our small-scale null does not rule out, then lineages will begin to
|
||||
speciate, and the ecosystem's future is a set of diverging species connected by routing rather than by
|
||||
merging. Which of the two it will be is measurable now, with the pre-merge conflict instruments this
|
||||
paper tested.
|
||||
|
|
@ -531,42 +491,42 @@ publication; every figure in this paper regenerates from committed artifacts wit
|
|||
25. Y. Yoon, D. Hu, I. Weissburg, Y. Qin, H. Jeong, Model collapse in the self-consuming chain of diffusion finetuning: A novel perspective from quantitative trait modeling. *Int. Conf. Learn. Represent.* (2025). https://doi.org/10.48550/arXiv.2407.17493.
|
||||
26. M. McCloskey, N. J. Cohen, Catastrophic interference in connectionist networks: The sequential learning problem. *Psychol. Learn. Motiv.* **24**, 109–165 (1989).
|
||||
27. R. M. French, Catastrophic forgetting in connectionist networks. *Trends Cogn. Sci.* **3**, 128–135 (1999).
|
||||
28. T. Scialom, T. Chakrabarty, S. Muresan, Fine-tuned language models are continual learners. *Proc. Conf. Empir. Methods Nat. Lang. Process.* (2022). https://doi.org/10.48550/arXiv.2205.12393.
|
||||
29. A. Ibrahim, et al., Simple and scalable strategies to continually pre-train large language models. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2403.08763.
|
||||
30. A. Robins, Catastrophic forgetting, rehearsal and pseudorehearsal. *Connect. Sci.* **7**, 123–146 (1995).
|
||||
31. H. Shin, J. K. Lee, J. Kim, J. Kim, Continual learning with deep generative replay. *Adv. Neural Inf. Process. Syst.* **30** (2017). https://doi.org/10.48550/arXiv.1705.08690.
|
||||
32. B. Yi, Q. Liu, Y. Cheng, H. Xu, Escaping model collapse via synthetic data verification. arXiv [Preprint] (2025). https://doi.org/10.48550/arXiv.2510.16657.
|
||||
33. Y. Feng, et al., Beyond model collapse: Scaling up with synthesized data requires verification. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2406.07515.
|
||||
34. A. A. Rusu, et al., Progressive neural networks. arXiv [Preprint] (2016). https://doi.org/10.48550/arXiv.1606.04671.
|
||||
35. D. Biderman, et al., LoRA learns less and forgets less. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2405.09673.
|
||||
36. J. L. McClelland, B. L. McNaughton, R. C. O'Reilly, Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory. *Psychol. Rev.* **102**, 419–457 (1995).
|
||||
37. D. Kumaran, D. Hassabis, J. L. McClelland, What learning systems do intelligent agents need? Complementary learning systems theory updated. *Trends Cogn. Sci.* **20**, 512–534 (2016).
|
||||
38. J. Schwarz, et al., Progress & Compress: A scalable framework for continual learning. *Proc. Int. Conf. Mach. Learn.* (2018).
|
||||
39. G. Ilharco, et al., Editing models with task arithmetic. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2212.04089.
|
||||
40. D. Marczak, B. Twardowski, T. Trzciński, S. Cygert, MagMax: Leveraging model merging for seamless continual learning. *Proc. Eur. Conf. Comput. Vis.* (2024). https://doi.org/10.48550/arXiv.2407.06322.
|
||||
41. A. Alexandrov, et al., Mitigating catastrophic forgetting in language transfer via model merging. *Findings Assoc. Comput. Linguist.: EMNLP* (2024). https://doi.org/10.48550/arXiv.2407.08699.
|
||||
42. S. Dziadzio, et al., How to merge your multimodal models over time? *Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.* (2025). https://doi.org/10.48550/arXiv.2412.06712.
|
||||
43. M. Toneva, et al., An empirical study of example forgetting during deep neural network learning. *Int. Conf. Learn. Represent.* (2019). https://doi.org/10.48550/arXiv.1812.05159.
|
||||
44. N. Kandpal, H. Deng, A. Roberts, E. Wallace, C. Raffel, Large language models struggle to learn long-tail knowledge. *Proc. Int. Conf. Mach. Learn.* (2023). https://doi.org/10.48550/arXiv.2211.08411.
|
||||
45. X. Liu, et al., Long-tailed class incremental learning. *Proc. Eur. Conf. Comput. Vis.* (2022). https://doi.org/10.48550/arXiv.2210.00266.
|
||||
46. H. J. Muller, The relation of recombination to mutational advance. *Mutat. Res.* **1**, 2–9 (1964).
|
||||
47. M. Gerstgrasser, et al., Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. *Conf. Lang. Model.* (2024). https://doi.org/10.48550/arXiv.2404.01413.
|
||||
48. S. Wright, Evolution in Mendelian populations. *Genetics* **16**, 97–159 (1931).
|
||||
49. R. A. Fisher, *The Genetical Theory of Natural Selection* (Clarendon Press, 1930).
|
||||
50. H. J. Muller, Some genetic aspects of sex. *Am. Nat.* **66**, 118–138 (1932).
|
||||
51. H. A. Orr, The population genetics of speciation: The evolution of hybrid incompatibilities. *Genetics* **139**, 1805–1813 (1995).
|
||||
52. H. A. Orr, M. Turelli, The evolution of postzygotic isolation: Accumulating Dobzhansky–Muller incompatibilities. *Evolution* **55**, 1085–1094 (2001).
|
||||
53. A. Livnat, C. Papadimitriou, Sex as an algorithm: The theory of evolution under the lens of computation. *Commun. ACM* **59**, 84–93 (2016).
|
||||
54. L. Yu, B. Yu, H. Yu, F. Huang, Y. Li, Language models are super Mario: Absorbing abilities from homologous models as a free lunch. *Proc. Int. Conf. Mach. Learn.* (2024). https://doi.org/10.48550/arXiv.2311.03099.
|
||||
55. M. Wortsman, et al., Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. *Proc. Int. Conf. Mach. Learn.* (2022). https://doi.org/10.48550/arXiv.2203.05482.
|
||||
56. L. Zhou, B. Zhao, R. Yu, E. Rodolà, Demystifying mergeability: Interpretable properties to predict model merging success. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2601.22285.
|
||||
57. Y. Cao, et al., An empirical study and theoretical explanation on task-level model-merging collapse. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2603.09463.
|
||||
58. J. Pari, S. Jelassi, P. Agrawal, Collective model intelligence requires compatible specialization. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2411.02207.
|
||||
59. E. J. Hu, et al., LoRA: Low-rank adaptation of large language models. *Int. Conf. Learn. Represent.* (2022). https://doi.org/10.48550/arXiv.2106.09685.
|
||||
60. S. A. Kauffman, S. Levin, Towards a general theory of adaptive walks on rugged landscapes. *J. Theor. Biol.* **128**, 11–45 (1987).
|
||||
61. J. Lehman, K. O. Stanley, Abandoning objectives: Evolution through the search for novelty alone. *Evol. Comput.* **19**, 189–223 (2011).
|
||||
62. S. K. Ainsworth, J. Hayase, S. Srinivasa, Git Re-Basin: Merging models modulo permutation symmetries. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2209.04836.
|
||||
63. T. Li, Z. Shen, Scaling linear mode connectivity and merging to billion-parameter pretrained transformers. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2606.23607.
|
||||
64. E. Sharma, D. M. Roy, G. K. Dziugaite, The non-local model merging problem: Permutation symmetries and variance collapse. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2410.12766.
|
||||
65. N. Kozodoi, Z. Afolabi, J. Butler, Are we merging the right models? Impact of expert training duration on model merging for LLMs. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2607.11997.
|
||||
66. S. Kotha, J. M. Springer, A. Raghunathan, Understanding catastrophic forgetting in language models via implicit inference. *Int. Conf. Learn. Represent.* (2024). https://doi.org/10.48550/arXiv.2309.10105.
|
||||
28. H. J. Muller, The relation of recombination to mutational advance. *Mutat. Res.* **1**, 2–9 (1964).
|
||||
29. B. Yi, Q. Liu, Y. Cheng, H. Xu, Escaping model collapse via synthetic data verification. arXiv [Preprint] (2025). https://doi.org/10.48550/arXiv.2510.16657.
|
||||
30. M. Gerstgrasser, et al., Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. *Conf. Lang. Model.* (2024). https://doi.org/10.48550/arXiv.2404.01413.
|
||||
31. S. Wright, Evolution in Mendelian populations. *Genetics* **16**, 97–159 (1931).
|
||||
32. J. Pari, S. Jelassi, P. Agrawal, Collective model intelligence requires compatible specialization. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2411.02207.
|
||||
33. R. A. Fisher, *The Genetical Theory of Natural Selection* (Clarendon Press, 1930).
|
||||
34. H. J. Muller, Some genetic aspects of sex. *Am. Nat.* **66**, 118–138 (1932).
|
||||
35. E. J. Hu, et al., LoRA: Low-rank adaptation of large language models. *Int. Conf. Learn. Represent.* (2022). https://doi.org/10.48550/arXiv.2106.09685.
|
||||
36. S. A. Kauffman, S. Levin, Towards a general theory of adaptive walks on rugged landscapes. *J. Theor. Biol.* **128**, 11–45 (1987).
|
||||
37. J. Lehman, K. O. Stanley, Abandoning objectives: Evolution through the search for novelty alone. *Evol. Comput.* **19**, 189–223 (2011).
|
||||
38. H. A. Orr, The population genetics of speciation: The evolution of hybrid incompatibilities. *Genetics* **139**, 1805–1813 (1995).
|
||||
39. H. A. Orr, M. Turelli, The evolution of postzygotic isolation: Accumulating Dobzhansky–Muller incompatibilities. *Evolution* **55**, 1085–1094 (2001).
|
||||
40. S. K. Ainsworth, J. Hayase, S. Srinivasa, Git Re-Basin: Merging models modulo permutation symmetries. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2209.04836.
|
||||
41. T. Li, Z. Shen, Scaling linear mode connectivity and merging to billion-parameter pretrained transformers. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2606.23607.
|
||||
42. E. Sharma, D. M. Roy, G. K. Dziugaite, The non-local model merging problem: Permutation symmetries and variance collapse. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2410.12766.
|
||||
43. N. Kozodoi, Z. Afolabi, J. Butler, Are we merging the right models? Impact of expert training duration on model merging for LLMs. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2607.11997.
|
||||
44. L. Zhou, B. Zhao, R. Yu, E. Rodolà, Demystifying mergeability: Interpretable properties to predict model merging success. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2601.22285.
|
||||
45. T. Scialom, T. Chakrabarty, S. Muresan, Fine-tuned language models are continual learners. *Proc. Conf. Empir. Methods Nat. Lang. Process.* (2022). https://doi.org/10.48550/arXiv.2205.12393.
|
||||
46. A. Ibrahim, et al., Simple and scalable strategies to continually pre-train large language models. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2403.08763.
|
||||
47. A. Robins, Catastrophic forgetting, rehearsal and pseudorehearsal. *Connect. Sci.* **7**, 123–146 (1995).
|
||||
48. H. Shin, J. K. Lee, J. Kim, J. Kim, Continual learning with deep generative replay. *Adv. Neural Inf. Process. Syst.* **30** (2017). https://doi.org/10.48550/arXiv.1705.08690.
|
||||
49. Y. Feng, et al., Beyond model collapse: Scaling up with synthesized data requires verification. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2406.07515.
|
||||
50. A. A. Rusu, et al., Progressive neural networks. arXiv [Preprint] (2016). https://doi.org/10.48550/arXiv.1606.04671.
|
||||
51. D. Biderman, et al., LoRA learns less and forgets less. *Trans. Mach. Learn. Res.* (2024). https://doi.org/10.48550/arXiv.2405.09673.
|
||||
52. J. L. McClelland, B. L. McNaughton, R. C. O'Reilly, Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory. *Psychol. Rev.* **102**, 419–457 (1995).
|
||||
53. D. Kumaran, D. Hassabis, J. L. McClelland, What learning systems do intelligent agents need? Complementary learning systems theory updated. *Trends Cogn. Sci.* **20**, 512–534 (2016).
|
||||
54. J. Schwarz, et al., Progress & Compress: A scalable framework for continual learning. *Proc. Int. Conf. Mach. Learn.* (2018).
|
||||
55. G. Ilharco, et al., Editing models with task arithmetic. *Int. Conf. Learn. Represent.* (2023). https://doi.org/10.48550/arXiv.2212.04089.
|
||||
56. D. Marczak, B. Twardowski, T. Trzciński, S. Cygert, MagMax: Leveraging model merging for seamless continual learning. *Proc. Eur. Conf. Comput. Vis.* (2024). https://doi.org/10.48550/arXiv.2407.06322.
|
||||
57. A. Alexandrov, et al., Mitigating catastrophic forgetting in language transfer via model merging. *Findings Assoc. Comput. Linguist.: EMNLP* (2024). https://doi.org/10.48550/arXiv.2407.08699.
|
||||
58. S. Dziadzio, et al., How to merge your multimodal models over time? *Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.* (2025). https://doi.org/10.48550/arXiv.2412.06712.
|
||||
59. M. Toneva, et al., An empirical study of example forgetting during deep neural network learning. *Int. Conf. Learn. Represent.* (2019). https://doi.org/10.48550/arXiv.1812.05159.
|
||||
60. N. Kandpal, H. Deng, A. Roberts, E. Wallace, C. Raffel, Large language models struggle to learn long-tail knowledge. *Proc. Int. Conf. Mach. Learn.* (2023). https://doi.org/10.48550/arXiv.2211.08411.
|
||||
61. X. Liu, et al., Long-tailed class incremental learning. *Proc. Eur. Conf. Comput. Vis.* (2022). https://doi.org/10.48550/arXiv.2210.00266.
|
||||
62. S. Kotha, J. M. Springer, A. Raghunathan, Understanding catastrophic forgetting in language models via implicit inference. *Int. Conf. Learn. Represent.* (2024). https://doi.org/10.48550/arXiv.2309.10105.
|
||||
63. L. Yu, B. Yu, H. Yu, F. Huang, Y. Li, Language models are super Mario: Absorbing abilities from homologous models as a free lunch. *Proc. Int. Conf. Mach. Learn.* (2024). https://doi.org/10.48550/arXiv.2311.03099.
|
||||
64. M. Wortsman, et al., Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. *Proc. Int. Conf. Mach. Learn.* (2022). https://doi.org/10.48550/arXiv.2203.05482.
|
||||
65. Y. Cao, et al., An empirical study and theoretical explanation on task-level model-merging collapse. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2603.09463.
|
||||
66. A. Livnat, C. Papadimitriou, Sex as an algorithm: The theory of evolution under the lens of computation. *Commun. ACM* **59**, 84–93 (2016).
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue