narrative: remove internal-deliberation ghosts from the diagnosis passages

The convergence paragraph rewritten as a natural literature entry: the
drift identification is stated as a fact of the field, made repeatedly
and independently (pre-deep-learning inference chains; LLM text
ecosystems; the first-extinction law; quantitative-genetic form), its
multiplicity presented as a property of the idea rather than a claim
about us; the pivot is positive (population genetics is a theory of what
maintains populations despite decay, and this paper develops that fuller
structure) instead of defensive ("what none of that parallel work
develops"). "We reached independently", "priority of publication", and
"convergence we take as support" removed from the abstract and the
Discussion ledger as well. Refs 22-25 renumbered to the new textual
(chronological) order; citation invariant re-verified (1..66). Lesson
recorded: internal strategic deliberations must not surface in
reader-facing prose — confident papers situate, they do not litigate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
Giorgio Gilestro 2026-09-07 10:03:28 +01:00
parent e2981c845d
commit f82d11ccce
4 changed files with 35 additions and 26 deletions

View file

@ -23,8 +23,8 @@ model output, and recombined by weight merging, and the practice is described in
vocabulary with little use of evolutionary theory. We treat multigenerational model populations as
systems whose inheritance, diversity, and compatibility must be managed, and we transfer the
quantitative framework of the evolution of sex. Its starting point, that training on model output is
genetic drift and model collapse its signature, we reached independently; parallel work has
formalised the same diagnosis, a convergence we take as support for the frame. In a minimal
genetic drift and model collapse its signature, is by now established from several independent
directions; we develop the structure that follows from it. In a minimal
inheritance model that is exactly WrightFisher, and measurably WrightFisher plus estimator bias in
trained networks, we derive and test remedies. Grounding acts as immigration: a real-data fraction
far below one retained most equilibrium diversity, with a per-capability observation floor that
@ -67,19 +67,20 @@ diversity, and compatibility must be managed, not merely collections of models t
branch of biology that studies exactly this problem, the population genetics of the evolution of sex,
transfers as a quantitative framework.
The frame's entry point is the diagnosis. Training each generation of a model on the previous
generation's output degrades it (*model collapse*): rare capabilities vanish first and the lineage
drifts toward its own most common behaviour (21). That this is the mathematics of *genetic drift* in
a finite population is a conclusion we reached independently in building the present framework, and
one that has been derived in parallel from several other directions (2224), including a closed-form
first-extinction law placing collapse onset at the WrightFisher first-extinction time (23), and that
was anticipated, before deep learning, in an analysis of sequential inference chains as generalised
genetic drift (25). We cite these works for priority of publication on the diagnosis and read the
convergence, independent arrivals at the same population-genetic account by different routes and in
different decades, as corroboration that the frame is the natural one. What none of that parallel work develops, and what this paper is about, is the
structure the diagnosis opens: the full arc from drift through its remedies (immigration,
recombination, selection, population structure) to its limit (reproductive isolation), carried as one
framework from closed forms to trained networks to language models.
The diagnosis comes first. Training each generation of a model on the previous generation's output
degrades it (*model collapse*): rare capabilities vanish first, and the lineage drifts toward its own
most common behaviour (21). That degradation is, mathematically, *genetic drift*, the loss of rare
variants that any finite population suffers when each generation is a finite sample of the last. The
identification has been made repeatedly and independently: for sequential inference chains before deep
learning (22), for language-model text ecosystems (23), as a closed-form first-extinction law placing
collapse onset at the WrightFisher first-extinction time (24), and in quantitative-genetic form for
self-consuming diffusion models (25). A diagnosis reached so often, from such different starting
points, marks population genetics as the natural mathematics of the setting. It is also only the entry
point. Population genetics is not, at heart, a theory of decay; it is a theory of the mechanisms that
maintain and build populations despite decay (immigration, recombination, selection, population
structure) and of where those mechanisms reach their limits. This paper develops that fuller structure
for model populations: the arc from drift through its remedies to its limit, reproductive isolation,
carried as one framework from closed forms to trained networks to language models.
In machine learning's own terms, the problem this frame addresses is the field's oldest,
*continual learning*, reappearing one level up. Within a single network, sequential learning
@ -426,9 +427,8 @@ skewed task-inference over latent capability rather than erasure (66); our irrev
concern oracle-measured behavioural distributions, and distinguishing latent from extinct capability
at language-model scale is an open experiment whose outcome would be decisive for both readings.
**What is borrowed and what is ours.** The diagnosis — collapse as drift — was published first by
others and we cite it so (2124), while noting the derivations are independent and convergent; prior art
in the strict sense are the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
**What is borrowed and what is ours.** The collapse-as-drift diagnosis is established prior work
(2125); so are the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
naive averaging loses to interference-aware or routed merges (4, 54, 55), that model populations can
climb (5, 810), and that merge success admits ML-native predictors (56, 57). Ours is the framework-level
synthesis — inheritance, diversity, and compatibility as managed quantities — together with: the
@ -525,10 +525,10 @@ publication; every figure in this paper regenerates from committed artifacts wit
19. T. Guo, et al., Large language model based multi-agents: A survey of progress and challenges. *Proc. Int. Joint Conf. Artif. Intell.* (2024). https://doi.org/10.48550/arXiv.2402.01680.
20. N. Tomasev, et al., Virtual agent economies. arXiv [Preprint] (2025). https://doi.org/10.48550/arXiv.2509.10147.
21. I. Shumailov, et al., AI models collapse when trained on recursively generated data. *Nature* **631**, 755759 (2024).
22. S. Riis, Drift and selection in LLM text ecosystems. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2604.08554.
23. M. Benati, A. Londei, D. Lanzieri, V. Loreto, First-extinction law for resampling processes. arXiv [Preprint] (2025). https://doi.org/10.48550/arXiv.2509.20101.
24. Y. Yoon, D. Hu, I. Weissburg, Y. Qin, H. Jeong, Model collapse in the self-consuming chain of diffusion finetuning: A novel perspective from quantitative trait modeling. *Int. Conf. Learn. Represent.* (2025). https://doi.org/10.48550/arXiv.2407.17493.
25. J. P. Crutchfield, S. Whalen, Structural drift: The population dynamics of sequential learning. *PLOS Comput. Biol.* **8**, e1002510 (2012).
22. J. P. Crutchfield, S. Whalen, Structural drift: The population dynamics of sequential learning. *PLOS Comput. Biol.* **8**, e1002510 (2012).
23. S. Riis, Drift and selection in LLM text ecosystems. arXiv [Preprint] (2026). https://doi.org/10.48550/arXiv.2604.08554.
24. M. Benati, A. Londei, D. Lanzieri, V. Loreto, First-extinction law for resampling processes. arXiv [Preprint] (2025). https://doi.org/10.48550/arXiv.2509.20101.
25. Y. Yoon, D. Hu, I. Weissburg, Y. Qin, H. Jeong, Model collapse in the self-consuming chain of diffusion finetuning: A novel perspective from quantitative trait modeling. *Int. Conf. Learn. Represent.* (2025). https://doi.org/10.48550/arXiv.2407.17493.
26. M. McCloskey, N. J. Cohen, Catastrophic interference in connectionist networks: The sequential learning problem. *Psychol. Learn. Motiv.* **24**, 109165 (1989).
27. R. M. French, Catastrophic forgetting in connectionist networks. *Trends Cogn. Sci.* **3**, 128135 (1999).
28. T. Scialom, T. Chakrabarty, S. Muresan, Fine-tuned language models are continual learners. *Proc. Conf. Empir. Methods Nat. Lang. Process.* (2022). https://doi.org/10.48550/arXiv.2205.12393.