New Fig. 1 (experimental-programme schematic); Table 2 to SI; figures in citation order; Fig. 2B legible labels
Replaces the results table with a pipeline figure: five questions x three architecture tiers (exact Wright-Fisher simulator, trained networks, language models), filled cells naming the experiments, dashed cells the honest gaps. Table 1 (the dictionary) stays; Table 2 moves to SI Appendix Table S2. The renumber surfaced a pre-existing citation-order violation (the LLM figure was cited in the recombination section before Figs. 3-6), so figures are renumbered to strict first-citation order (LLM tier is now Fig. 3). Fig. 2B: the montage's baked-in raster labels are cropped away and replaced with vector row numbers under a rotated "generation" header. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
parent
073fc33509
commit
0159e2839a
13 changed files with 208 additions and 109 deletions
|
|
@ -16,13 +16,18 @@ The generations are coupled through data as well as through weights. Successive
|
|||
|
||||
Training each generation of a model on the previous generation's output degrades it (\emph{model collapse}): rare capabilities vanish first, and the lineage drifts toward its own most common behaviour (21). That degradation is, mathematically, \emph{genetic drift}, the loss of rare variants that any finite population suffers when each generation is a finite sample of the last --- the same sampling accident by which rare surnames vanish from small villages and rare alleles (gene variants) drift out of island populations with no selection against them. The identification has been made repeatedly and independently: for sequential inference chains before deep learning (22), for language-model text ecosystems (23), as a closed-form first-extinction law placing collapse onset at the Wright--Fisher first-extinction time (24), and in quantitative-genetic form for self-consuming diffusion models (25). A diagnosis reached so often, from such different starting points, marks population genetics as the natural mathematics of the setting, though only as its entry point: population genetics is not, at heart, a theory of decay; it is a theory of the mechanisms that maintain and build populations despite decay (immigration, recombination, selection, population structure) and of where those mechanisms reach their limits. This paper develops that fuller structure for model populations: the arc from drift through its remedies to its limit, reproductive isolation --- the point at which diverged lineages can no longer produce working offspring, biology's boundary between species --- carried as one framework from closed forms to trained networks to language models.
|
||||
|
||||
An operator of a model population faces recurring decisions for which there is no principled guidance: how much verified real data does retraining need before a lineage decays; will combining two particular models compose their abilities or damage them; can incompatibility be detected before paying for a failed merge; and when should specialists be kept separate rather than consolidated? In practice these are settled by convention and by trial-and-error search. They are also, recognisably, machine learning's oldest problem at a new scale: \emph{continual learning}, the struggle to acquire new abilities without losing old ones (26, 27), transposed from a single network to a population whose members inherit from one another. Population genetics, I will argue, prices these decisions. Table 1 summarises the correspondences on which the argument runs; the sections that follow develop them from closed-form theory to experiments in trained networks and language models.
|
||||
An operator of a model population faces recurring decisions for which there is no principled guidance: how much verified real data does retraining need before a lineage decays; will combining two particular models compose their abilities or damage them; can incompatibility be detected before paying for a failed merge; and when should specialists be kept separate rather than consolidated? In practice these are settled by convention and by trial-and-error search. They are also, recognisably, machine learning's oldest problem at a new scale: \emph{continual learning}, the struggle to acquire new abilities without losing old ones (26, 27), transposed from a single network to a population whose members inherit from one another. Population genetics, I will argue, prices these decisions. Table 1 summarises the correspondences on which the argument runs, and Fig. 1 maps the experimental programme built on them: the same abstractions tested at three tiers of model architecture --- an exact simulator, trained neural networks, and language models --- with the sections that follow climbing that ladder.
|
||||
|
||||
\begin{figure*}[p]\centering % fig1
|
||||
\includegraphics[width=\textwidth]{figs/fig1.pdf}
|
||||
\caption{The experimental programme. Each population-genetic abstraction (Table 1) is tested at up to three tiers of model architecture, ordered left to right by increasing realism: an exact Wright--Fisher simulator over knowledge distributions (closed forms; bitwise-reproducible), trained neural networks measured against exact oracles (recurrent, feedforward, and variational-autoencoder generators on a synthetic mode universe, and a convolutional VAE on MNIST), and language models (LoRA specialists on Qwen bases at 0.5B and 7B, scored by an exact-match verifier). Rows are the paper's five questions; filled cells name the experiments run at each tier; dashed cells were not tested, and the composed society at language-model scale is the paper's stated gap.}\label{fig1}
|
||||
\end{figure*}
|
||||
|
||||
\section*{The minimal model, and where its exactness ends}
|
||||
|
||||
Knowledge is modelled as a distribution \texttt{p\_t} over \texttt{K} discrete items (capabilities, facts, modes of behaviour), with a fixed true distribution \texttt{p*} whose rare tail carries the knowledge most at risk. One generation is: *draw \texttt{n} samples from the parent's distribution, optionally mix in \texttt{m} verified real samples (``grounding'', \texttt{g = m/(n+m)}), and refit the child\emph{. In this minimal inheritance model the resampling step }is* the Wright--Fisher process, population genetics' canonical model of neutral evolution, in which each new generation is a random sample of size \texttt{n} from the previous one and every statistical property of drift follows from that one step. Diversity throughout this paper is \emph{heterozygosity}, \texttt{H = 1 − Σ p\_i²}: the probability that two random draws differ (one minus a collision probability), high when many items share the mass, zero at total collapse. The identity is exploited as an engineering gate: the simulator reproduces the classical closed forms (heterozygosity decay \texttt{E[H\_t] = H\_0(1 − 1/n)\textasciicircum{}t}; the exact immigration--drift equilibrium; the closed-form multi-teacher union) to within 0.5\%, and these are standing tests in the codebase, not one-off checks.
|
||||
|
||||
The boundary of the exactness matters, and I measured it rather than assumed it. Real training adds approximation, optimisation noise, and inductive bias, and when trained networks are fit against the exact drift null they deviate in \emph{opposite, architecture-specific} directions: a smoothing recurrent network resists collapse (keeping spurious variants alive), while a sharpening image generator accelerates it. A one-parameter \emph{learning kernel} (a smoothing knob and a sharpening knob on the refit) reproduces both. Throughout, a real learner is therefore treated as Wright--Fisher \emph{plus a signed, measurable estimator bias}, and the drift signs (rare-first loss; the grounding response) survived that bias in every architecture I tested, including a convolutional VAE retrained on its own generated digits, where the dry lineage collapses to a single blurred digit class while 10\% grounding holds all thirty modes (Fig. 1). Retraining on a single parent is \emph{asexual reproduction}, and sustained loss under it carries the defining consequence of \emph{Muller's ratchet} (28), the mechanism by which lineages that never recombine decay irreversibly --- the reason non-recombining genomes such as the Y chromosome have shed most of their ancestral genes. Once every copy of a rare capability is gone from all parents and sources, no recombination can rebuild it: each such loss is a click of the ratchet, and remedies must act while copies still survive somewhere (a consequence-level correspondence; the minimal model lacks the ratchet's recurrent-mutation driver).
|
||||
The boundary of the exactness matters, and I measured it rather than assumed it. Real training adds approximation, optimisation noise, and inductive bias, and when trained networks are fit against the exact drift null they deviate in \emph{opposite, architecture-specific} directions: a smoothing recurrent network resists collapse (keeping spurious variants alive), while a sharpening image generator accelerates it. A one-parameter \emph{learning kernel} (a smoothing knob and a sharpening knob on the refit) reproduces both. Throughout, a real learner is therefore treated as Wright--Fisher \emph{plus a signed, measurable estimator bias}, and the drift signs (rare-first loss; the grounding response) survived that bias in every architecture I tested, including a convolutional VAE retrained on its own generated digits, where the dry lineage collapses to a single blurred digit class while 10\% grounding holds all thirty modes (Fig. 2). Retraining on a single parent is \emph{asexual reproduction}, and sustained loss under it carries the defining consequence of \emph{Muller's ratchet} (28), the mechanism by which lineages that never recombine decay irreversibly --- the reason non-recombining genomes such as the Y chromosome have shed most of their ancestral genes. Once every copy of a rare capability is gone from all parents and sources, no recombination can rebuild it: each such loss is a click of the ratchet, and remedies must act while copies still survive somewhere (a consequence-level correspondence; the minimal model lacks the ratchet's recurrent-mutation driver).
|
||||
|
||||
\textbf{Table 1.} The dictionary. Each biological term is introduced in the section that develops it; each correspondence is stated with the level of support it currently has (exact = closed form in the minimal model; empirical = measured in trained systems; hypothesis = stated with a falsifier, untested or unconfirmed). The full claim-by-claim ledger with assumptions and known limits is SI Appendix, Table S1.
|
||||
|
||||
|
|
@ -45,79 +50,65 @@ Selection on a fitness function & Verifier-anchored selection (``reality that ca
|
|||
|
||||
\subsection*{Grounding is immigration: cheap, with a floor}
|
||||
|
||||
In the minimal model, grounding from a fixed real source is \emph{immigration} into a drifting population (29--31). Immigration is what conservation managers prescribe when a fragmented reserve loses diversity, and its striking property there is how little is needed --- the field's rule of thumb is that one migrant per generation holds an isolated population's diversity (32). The same economy appears here: the equilibrium diversity has a closed form the simulator matches exactly. That equilibrium is \emph{smooth} in the grounding fraction (there is no phase transition in aggregate diversity), so the practical number is an operational threshold, and I define it as such: under the tested population size and Zipf source distribution, \texttt{g \(\approx\) 0.05} retained most (\(\geq\)95\%) of equilibrium diversity indefinitely, with the required fraction depending on sample size, source distribution, and the chosen retention target (dependencies in SI). Verified real data remains, on any of these definitions, cheap insurance at fractions far below one. But the same analysis yields a floor the field's average-loss framing misses: under unstratified sampling from the source, a capability of rarity \texttt{p} appears in a real-data batch of size \texttt{m} with probability \texttt{1 − e\textasciicircum{}{−m\(\cdot\)p}}, so \texttt{m\(\cdot\)p \(\approx\) 1} marks roughly a 63\% chance of one example per batch: a soft observation floor, with higher confidence priced accordingly, and with distinct consequences for continuous retention, stationary occupancy, and reintroduction after loss (immigration can restore an absent item; SI separates these). Protecting the rarest knowledge under unstratified grounding is therefore priced per item at cost \texttt{\(\propto\) 1/p}; targeted or stratified sampling changes that cost, and recombination can recover rare capabilities \emph{that are still retained across complementary parents} (next section). In trained networks the \emph{sign} of the grounding response transfers everywhere I looked, with two deviations, both traced to the estimator bias above: sharp thresholds soften, and support-counting metrics decouple from truth (forward-KL is the operative collapse metric for a smoothing learner). On real images (Fig. 1B), dry self-training collapses a convolutional VAE to one mode while \textasciitilde{}10\% grounding holds all thirty (the trained model needs roughly twice the exact-operator fraction, the measured price of the estimator bias).
|
||||
In the minimal model, grounding from a fixed real source is \emph{immigration} into a drifting population (29--31). Immigration is what conservation managers prescribe when a fragmented reserve loses diversity, and its striking property there is how little is needed --- the field's rule of thumb is that one migrant per generation holds an isolated population's diversity (32). The same economy appears here: the equilibrium diversity has a closed form the simulator matches exactly. That equilibrium is \emph{smooth} in the grounding fraction (there is no phase transition in aggregate diversity), so the practical number is an operational threshold, and I define it as such: under the tested population size and Zipf source distribution, \texttt{g \(\approx\) 0.05} retained most (\(\geq\)95\%) of equilibrium diversity indefinitely, with the required fraction depending on sample size, source distribution, and the chosen retention target (dependencies in SI). Verified real data remains, on any of these definitions, cheap insurance at fractions far below one. But the same analysis yields a floor the field's average-loss framing misses: under unstratified sampling from the source, a capability of rarity \texttt{p} appears in a real-data batch of size \texttt{m} with probability \texttt{1 − e\textasciicircum{}{−m\(\cdot\)p}}, so \texttt{m\(\cdot\)p \(\approx\) 1} marks roughly a 63\% chance of one example per batch: a soft observation floor, with higher confidence priced accordingly, and with distinct consequences for continuous retention, stationary occupancy, and reintroduction after loss (immigration can restore an absent item; SI separates these). Protecting the rarest knowledge under unstratified grounding is therefore priced per item at cost \texttt{\(\propto\) 1/p}; targeted or stratified sampling changes that cost, and recombination can recover rare capabilities \emph{that are still retained across complementary parents} (next section). In trained networks the \emph{sign} of the grounding response transfers everywhere I looked, with two deviations, both traced to the estimator bias above: sharp thresholds soften, and support-counting metrics decouple from truth (forward-KL is the operative collapse metric for a smoothing learner). On real images (Fig. 2B), dry self-training collapses a convolutional VAE to one mode while \textasciitilde{}10\% grounding holds all thirty (the trained model needs roughly twice the exact-operator fraction, the measured price of the estimator bias).
|
||||
|
||||
\begin{figure*}[p]\centering % fig1
|
||||
\includegraphics[width=\textwidth]{figs/fig1.pdf}
|
||||
\caption{Grounding is immigration. (A) Stationary diversity against the grounding fraction in the minimal inheritance model: simulation (points, 95\% CI) matches the exact immigration--drift equilibrium (dashed). The equilibrium is smooth in $g$; $g \approx 0.05$ marks the operational threshold retaining 95\% of source diversity in this setting (red line, bootstrap CI shaded); the hollow point at $g=0$ is a finite-time value (the true equilibrium is zero). (B) The same signs on real images: samples from a convolutional VAE retrained each generation on its own output (rows: generations 0--15 of an ungrounded lineage) collapse toward a single blurred mode; 10\% grounding holds all thirty modes (quantified in SI).}\label{fig1}
|
||||
\begin{figure*}[p]\centering % fig2
|
||||
\includegraphics[width=\textwidth]{figs/fig2.pdf}
|
||||
\caption{Grounding is immigration. (A) Stationary diversity against the grounding fraction in the minimal inheritance model: simulation (points, 95\% CI) matches the exact immigration--drift equilibrium (dashed). The equilibrium is smooth in $g$; $g \approx 0.05$ marks the operational threshold retaining 95\% of source diversity in this setting (red line, bootstrap CI shaded); the hollow point at $g=0$ is a finite-time value (the true equilibrium is zero). (B) The same signs on real images: samples from a convolutional VAE retrained each generation on its own output (rows: generations 0--15 of an ungrounded lineage) collapse toward a single blurred mode; 10\% grounding holds all thirty modes (quantified in SI).}\label{fig2}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{Recombination: a conservation law, its operators, and offspring that exceed every parent}
|
||||
|
||||
The largest returns from the transfer concern merging. \emph{Blending inheritance} --- offspring as the average of their parents --- is the failure mode at the root of population genetics' founding controversy: the swamping argument pressed in Jenkin's 1867 review of \emph{The Origin of Species}, that under blending a rare advantageous variant is diluted toward the common type faster than selection can multiply it (33), an objection dissolved only by Mendel's particulate inheritance, in which discrete variants pass through generations undiluted. Refitting a child model to the mean of its parents' output distributions is blending inheritance, and the proposition below is Jenkin's dilution made exact. \textbf{Proposition (blending inheritance, rare-item regime).} Let K parents independently retain a rare item (mass \texttt{p} when retained), and let the child draw \texttt{n} samples either from one parent chosen at random or from the \emph{mean of the parents' output distributions}. Expected item mass is identical under the two schemes; and in the rare-item regime \texttt{n\(\cdot\)p/K \(\ll\) 1}, where per-item survival is first-order in sampled mass, expected \emph{survival} is also identical: the 1/K dilution of averaging cancels the K-parent union gain to first order, so in this regime adding parents through the output-mean does not increase expected tail retention. Two boundaries: outside that regime, survival is a convex function of mixed mass, so the variance reduction from averaging can \emph{reduce} extinction relative to a randomly chosen single parent; the cancellation is a first-order result about rare items, not a universal impossibility; and the contrasting union operator (keep each item's strongest source, then renormalise, which itself redistributes mass and presupposes a verifier or oracle to identify the strongest source) increases expected retention with K in all regimes in the minimal model. The practically important operators, \emph{weight averaging} (a nonlinear network's weight-mean does not compute its parents' output-mean) and \emph{routing among intact specialists} (34) (different storage and inference budgets from a single child), are its empirical cousins, and the measured bridge is a \emph{headroom rule}, stated qualitatively: in language models, union-preserving operators beat the weight-average where that average falls short of attainable performance, and add nothing where it does not (easy-versus-hard contrasts at two scales; a quantitative form of the relationship is untested). On easy tasks a capable base's average is already at ceiling and refinements add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent and routing wins by a wide margin (Fig. 6A--B).
|
||||
The largest returns from the transfer concern merging. \emph{Blending inheritance} --- offspring as the average of their parents --- is the failure mode at the root of population genetics' founding controversy: the swamping argument pressed in Jenkin's 1867 review of \emph{The Origin of Species}, that under blending a rare advantageous variant is diluted toward the common type faster than selection can multiply it (33), an objection dissolved only by Mendel's particulate inheritance, in which discrete variants pass through generations undiluted. Refitting a child model to the mean of its parents' output distributions is blending inheritance, and the proposition below is Jenkin's dilution made exact. \textbf{Proposition (blending inheritance, rare-item regime).} Let K parents independently retain a rare item (mass \texttt{p} when retained), and let the child draw \texttt{n} samples either from one parent chosen at random or from the \emph{mean of the parents' output distributions}. Expected item mass is identical under the two schemes; and in the rare-item regime \texttt{n\(\cdot\)p/K \(\ll\) 1}, where per-item survival is first-order in sampled mass, expected \emph{survival} is also identical: the 1/K dilution of averaging cancels the K-parent union gain to first order, so in this regime adding parents through the output-mean does not increase expected tail retention. Two boundaries: outside that regime, survival is a convex function of mixed mass, so the variance reduction from averaging can \emph{reduce} extinction relative to a randomly chosen single parent; the cancellation is a first-order result about rare items, not a universal impossibility; and the contrasting union operator (keep each item's strongest source, then renormalise, which itself redistributes mass and presupposes a verifier or oracle to identify the strongest source) increases expected retention with K in all regimes in the minimal model. The practically important operators, \emph{weight averaging} (a nonlinear network's weight-mean does not compute its parents' output-mean) and \emph{routing among intact specialists} (34) (different storage and inference budgets from a single child), are its empirical cousins, and the measured bridge is a \emph{headroom rule}, stated qualitatively: in language models, union-preserving operators beat the weight-average where that average falls short of attainable performance, and add nothing where it does not (easy-versus-hard contrasts at two scales; a quantitative form of the relationship is untested). On easy tasks a capable base's average is already at ceiling and refinements add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent and routing wins by a wide margin (Fig. 3A--B).
|
||||
|
||||
The generative payoff is the \emph{Fisher--Muller effect} (35, 36), the classical account of why sex speeds adaptation: in an asexual population, beneficial variants arising in different individuals can only compete until all but one lineage is lost, whereas recombination assembles them in one offspring, producing a \emph{genotype} (an individual's combination of variants, one at each \emph{locus}, or position) fitter than any parent. In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum, a genotype no parent held, while the best single parent and the blended average both plateau below (Fig. 2). In real language models the signature replicates under seed replication: merges of three LoRA (37) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 6A).
|
||||
The generative payoff is the \emph{Fisher--Muller effect} (35, 36), the classical account of why sex speeds adaptation: in an asexual population, beneficial variants arising in different individuals can only compete until all but one lineage is lost, whereas recombination assembles them in one offspring, producing a \emph{genotype} (an individual's combination of variants, one at each \emph{locus}, or position) fitter than any parent. In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum, a genotype no parent held, while the best single parent and the blended average both plateau below (Fig. 4). In real language models the signature replicates under seed replication: merges of three LoRA (37) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 3A).
|
||||
|
||||
Sex has risks and, for AI, an unfair advantage, both quantified on Kauffman's NK fitness landscapes (38), the standard model of \emph{epistasis}, biology's term for interaction between genes: the fitness contribution of a variant depends on which variants occupy the other loci, much as a component's value in an ML system depends on the components around it. Each of the landscape's \texttt{N} sites interacts with \texttt{K} others (the model's eponymous parameters), and raising that interaction count tunes the landscape from smooth and additive to rugged and many-peaked (Fig. 3). When skills are entangled, blind recombination produces offspring \emph{below} their parents, worsening with ruggedness, and the optimal recombination rate shrinks as entanglement grows. Biology knows this failure as \emph{outbreeding depression}, the reason conservation practice warns against crossing locally adapted populations: in the textbook case, an ibex herd in the Tatra Mountains restocked with animals from Turkey and Sinai produced fertile hybrids that bore their young in the coldest month of winter, and the herd died out (39). But an engineered population can do what biology cannot: recombine unbounded parents, choose complementary mates, and \emph{screen many candidate offspring against a verifier before keeping one}. This directed sex converts the outbreeding catastrophe into a reliable gain in the model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks, including one seed where the blend failed catastrophically and selection was immune (Fig. 6A). Finally, population \emph{structure} is itself a knob: sweeping the mate-pool breadth from monogamous (repeated local pairings) to promiscuous (\emph{panmixia}: any model may merge with any other) against ruggedness, wide mixing maximises the population mean while monotonically destroying diversity, and the best \emph{champion} shifts from wide breadth on smooth landscapes to intermediate breadth on rugged ones (Fig. 3C), the mating-system phenomenon known to structured-population search, mapped onto merging populations.
|
||||
|
||||
\begin{figure*}[p]\centering % fig2
|
||||
\includegraphics[width=\textwidth]{figs/fig2.pdf}
|
||||
\caption{Recombination in the minimal model: blending inheritance and the Fisher--Muller effect. (A) Expected rare-capability survival in a child refit from $K$ uncorrelated parents: the output-mean (blending) stays at the single-parent level --- the first-order cancellation --- while the union operator (strongest source per item, renormalised, oracle-identified) rises with parent count. (B) Multi-locus recombination of decorrelated specialists produces offspring fitter than any parent, approaching the optimum as parents are added; the best single parent and the blended average plateau below (mean $\pm$ 95\% CI).}\label{fig2}
|
||||
\end{figure*}
|
||||
Sex has risks and, for AI, an unfair advantage, both quantified on Kauffman's NK fitness landscapes (38), the standard model of \emph{epistasis}, biology's term for interaction between genes: the fitness contribution of a variant depends on which variants occupy the other loci, much as a component's value in an ML system depends on the components around it. Each of the landscape's \texttt{N} sites interacts with \texttt{K} others (the model's eponymous parameters), and raising that interaction count tunes the landscape from smooth and additive to rugged and many-peaked (Fig. 5). When skills are entangled, blind recombination produces offspring \emph{below} their parents, worsening with ruggedness, and the optimal recombination rate shrinks as entanglement grows. Biology knows this failure as \emph{outbreeding depression}, the reason conservation practice warns against crossing locally adapted populations: in the textbook case, an ibex herd in the Tatra Mountains restocked with animals from Turkey and Sinai produced fertile hybrids that bore their young in the coldest month of winter, and the herd died out (39). But an engineered population can do what biology cannot: recombine unbounded parents, choose complementary mates, and \emph{screen many candidate offspring against a verifier before keeping one}. This directed sex converts the outbreeding catastrophe into a reliable gain in the model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks, including one seed where the blend failed catastrophically and selection was immune (Fig. 3A). Finally, population \emph{structure} is itself a knob: sweeping the mate-pool breadth from monogamous (repeated local pairings) to promiscuous (\emph{panmixia}: any model may merge with any other) against ruggedness, wide mixing maximises the population mean while monotonically destroying diversity, and the best \emph{champion} shifts from wide breadth on smooth landscapes to intermediate breadth on rugged ones (Fig. 5C), the mating-system phenomenon known to structured-population search, mapped onto merging populations.
|
||||
|
||||
\begin{figure*}[p]\centering % fig3
|
||||
\includegraphics[width=\textwidth]{figs/fig3.pdf}
|
||||
\caption{Rugged (epistatic) landscapes: risk, remedy, and population structure. (A) Outbreeding depression: the mean offspring of blindly recombined specialist parents falls below the best parent, more steeply the more rugged the landscape (NK ruggedness $K$) and the higher the recombination rate. (B) Screening candidate offspring against a verifier (directed recombination) restores the gain at every ruggedness where blind recombination fails. (C) Mating structure: the best champion arises at wide mate-pool breadth on smooth landscapes and at intermediate breadth on rugged ones. (D) Wide breadth monotonically erodes population diversity at every ruggedness (mean $\pm$ 95\% CI, 20 replicates).}\label{fig3}
|
||||
\caption{The language-model tier. (A) Seed-replicated merging (0.5B, five seeds, fixed test sets; mean $\pm$ 95\% CI): merged specialists exceed the best single specialist overall, and only merged models are competent on every task family. (B) Hard, unsaturated tasks at 7B (single run): the weight-average dilutes a fragile specialist below the best single parent; routing among intact specialists preserves it. (C) The controlled predictive test (13 conditions $\times$ 3 seeds): pre-merge confidence-weighted functional conflict against merge penalty, coloured by grid axis --- penalty concentrates on the conflict axis. (D) Predictor comparison, $|$Spearman $\rho|$ against merge penalty over the full grid: functional measures carry signal, the tested weight-geometry baselines do not; paired differences between predictors are not individually significant.}\label{fig3}
|
||||
\end{figure*}
|
||||
|
||||
\begin{figure*}[p]\centering % fig4
|
||||
\includegraphics[width=\textwidth]{figs/fig4.pdf}
|
||||
\caption{Recombination in the minimal model: blending inheritance and the Fisher--Muller effect. (A) Expected rare-capability survival in a child refit from $K$ uncorrelated parents: the output-mean (blending) stays at the single-parent level --- the first-order cancellation --- while the union operator (strongest source per item, renormalised, oracle-identified) rises with parent count. (B) Multi-locus recombination of decorrelated specialists produces offspring fitter than any parent, approaching the optimum as parents are added; the best single parent and the blended average plateau below (mean $\pm$ 95\% CI).}\label{fig4}
|
||||
\end{figure*}
|
||||
|
||||
\begin{figure*}[p]\centering % fig5
|
||||
\includegraphics[width=\textwidth]{figs/fig5.pdf}
|
||||
\caption{Rugged (epistatic) landscapes: risk, remedy, and population structure. (A) Outbreeding depression: the mean offspring of blindly recombined specialist parents falls below the best parent, more steeply the more rugged the landscape (NK ruggedness $K$) and the higher the recombination rate. (B) Screening candidate offspring against a verifier (directed recombination) restores the gain at every ruggedness where blind recombination fails. (C) Mating structure: the best champion arises at wide mate-pool breadth on smooth landscapes and at intermediate breadth on rugged ones. (D) Wide breadth monotonically erodes population diversity at every ruggedness (mean $\pm$ 95\% CI, 20 replicates).}\label{fig5}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{The society: grounding, recombination, and diversity make complementary contributions}
|
||||
|
||||
Composing the operators (Fig. 4) requires one definitional distinction first. In the inheritance model, grounding is \emph{grounded inheritance}: external samples added to the reproduction process (the data channel). In the society model, grounding is \emph{grounded evaluation}: selection weights true fitness against conformity to the population's own consensus, \texttt{g}\(\cdot\)true-fitness + (1−g)\(\cdot\)conformity, the analogue of scoring models by the crowd's approval (the fitness channel). These are related design ideas, since both couple the lineage to a non-drifting external signal, but they are different operators, and I name them separately. In the tested society (a finite agent population on a rugged NK landscape), a four-arm ablation separates the failure modes: the full system (grounded evaluation + directed recombination + diversity-preserving selection (40)) climbs to near the global optimum while keeping its specialists; removing grounded evaluation converges the population confidently on an unfit consensus (self-consumption); removing recombination strands it on local optima; removing diversity converges it prematurely to a worse answer. Each removal fails differently; the three implementations make complementary contributions \emph{under the tested conditions}; general joint necessity is not established (alternative mutation, restart, archive, or selection schemes could alter the picture). At language-model scale this composed loop remains unbuilt; it is the paper's largest stated gap.
|
||||
Composing the operators (Fig. 6) requires one definitional distinction first. In the inheritance model, grounding is \emph{grounded inheritance}: external samples added to the reproduction process (the data channel). In the society model, grounding is \emph{grounded evaluation}: selection weights true fitness against conformity to the population's own consensus, \texttt{g}\(\cdot\)true-fitness + (1−g)\(\cdot\)conformity, the analogue of scoring models by the crowd's approval (the fitness channel). These are related design ideas, since both couple the lineage to a non-drifting external signal, but they are different operators, and I name them separately. In the tested society (a finite agent population on a rugged NK landscape), a four-arm ablation separates the failure modes: the full system (grounded evaluation + directed recombination + diversity-preserving selection (40)) climbs to near the global optimum while keeping its specialists; removing grounded evaluation converges the population confidently on an unfit consensus (self-consumption); removing recombination strands it on local optima; removing diversity converges it prematurely to a worse answer. Each removal fails differently; the three implementations make complementary contributions \emph{under the tested conditions}; general joint necessity is not established (alternative mutation, restart, archive, or selection schemes could alter the picture). At language-model scale this composed loop remains unbuilt; it is the paper's largest stated gap.
|
||||
|
||||
\begin{figure*}[p]\centering % fig4
|
||||
\includegraphics[width=\textwidth]{figs/fig4.pdf}
|
||||
\caption{The tested society: grounded evaluation, recombination, and diversity preservation make complementary contributions. A finite agent population on a rugged NK landscape; selection weights true fitness against conformity to the population consensus. (A) Best real fitness: the full system approaches the global optimum; removing grounded evaluation collapses the population onto a confident, unfit consensus; removing recombination or diversity preservation strands it lower. (B) Population diversity. (C) The self-consumption signature: conformity minus true fitness (mean $\pm$ 95\% CI, 12 replicates).}\label{fig4}
|
||||
\begin{figure*}[p]\centering % fig6
|
||||
\includegraphics[width=\textwidth]{figs/fig6.pdf}
|
||||
\caption{The tested society: grounded evaluation, recombination, and diversity preservation make complementary contributions. A finite agent population on a rugged NK landscape; selection weights true fitness against conformity to the population consensus. (A) Best real fitness: the full system approaches the global optimum; removing grounded evaluation collapses the population onto a confident, unfit consensus; removing recombination or diversity preservation strands it lower. (B) Population diversity. (C) The self-consumption signature: conformity minus true fitness (mean $\pm$ 95\% CI, 12 replicates).}\label{fig6}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{The limit of sex: model speciation}
|
||||
|
||||
Recombination presupposes compatible parents. In biology, lineages pushed far enough apart become separate species (\emph{reproductive isolation}) through Bateson--Dobzhansky--Muller incompatibilities (41, 42): changes harmless on their own genetic background but deleterious in combination --- the mechanism behind the mule's sterility and the inviability of many between-species crosses, in which two genomes that each work perfectly cannot run in the same cell. A merged model is exactly the exposed hybrid. I built the analytic model (Fig. 5A): hybrid fitness tracks the parents while compatible, then peels off and crashes below the ancestor; the isolation cliff arrives earlier the denser the incompatibilities; and the incompatibility \emph{count} snowballs quadratically with divergence (42). Note that a super-linear count does not by itself entail a sharp performance cliff without the count-to-effect-size link, which the analytic model supplies under its assumptions and any neural test must establish separately.
|
||||
Recombination presupposes compatible parents. In biology, lineages pushed far enough apart become separate species (\emph{reproductive isolation}) through Bateson--Dobzhansky--Muller incompatibilities (41, 42): changes harmless on their own genetic background but deleterious in combination --- the mechanism behind the mule's sterility and the inviability of many between-species crosses, in which two genomes that each work perfectly cannot run in the same cell. A merged model is exactly the exposed hybrid. I built the analytic model (Fig. 7A): hybrid fitness tracks the parents while compatible, then peels off and crashes below the ancestor; the isolation cliff arrives earlier the denser the incompatibilities; and the incompatibility \emph{count} snowballs quadratically with divergence (42). Note that a super-linear count does not by itself entail a sharp performance cliff without the count-to-effect-size link, which the analytic model supplies under its assumptions and any neural test must establish separately.
|
||||
|
||||
In trained networks, the claim must survive a known alternative: merge barriers between independently trained networks are famously \emph{coordinate artefacts}, removable by re-aligning hidden units (43); richer symmetry groups remove more (44), with known failures beyond the shared-data regime (45). I therefore aligned under the composition of permutation matching and exact per-unit rescaling (the unit symmetry group of plain ReLU MLPs, as the search space) and decomposed the barrier (Fig. 5 C and D): two networks trained from different initialisations on the \emph{same} task have a barrier that this alignment removes essentially entirely (residual \(\approx\) 0.001, the aligned merge performing at parent level): coordinate, not functional; two networks trained on \emph{conflicting} label maps have a barrier the same alignment leaves largely unchanged (0.502 \(\rightarrow\) 0.497), with the merged model functionally dead. The tested alignment removes the same-task barrier but leaves the conflict-associated barrier intact, supporting a functional-conflict interpretation without proving optimal alignment: exact recovery of a permuted-and-rescaled copy validates a special case, so the removable share is a lower bound and the residual an upper bound. Sweeping conflict traces the cliff as hybrid fitness, 0.97 \(\rightarrow\) 0.03. The conflict floor itself is information-theoretic (no single model can satisfy contradictory conventions; SI Appendix, Proposition S2), with the framework's role being the \emph{structure around it}: which divergences generate conflict, and what moves the cliff.
|
||||
In trained networks, the claim must survive a known alternative: merge barriers between independently trained networks are famously \emph{coordinate artefacts}, removable by re-aligning hidden units (43); richer symmetry groups remove more (44), with known failures beyond the shared-data regime (45). I therefore aligned under the composition of permutation matching and exact per-unit rescaling (the unit symmetry group of plain ReLU MLPs, as the search space) and decomposed the barrier (Fig. 7 C and D): two networks trained from different initialisations on the \emph{same} task have a barrier that this alignment removes essentially entirely (residual \(\approx\) 0.001, the aligned merge performing at parent level): coordinate, not functional; two networks trained on \emph{conflicting} label maps have a barrier the same alignment leaves largely unchanged (0.502 \(\rightarrow\) 0.497), with the merged model functionally dead. The tested alignment removes the same-task barrier but leaves the conflict-associated barrier intact, supporting a functional-conflict interpretation without proving optimal alignment: exact recovery of a permuted-and-rescaled copy validates a special case, so the removable share is a lower bound and the residual an upper bound. Sweeping conflict traces the cliff as hybrid fitness, 0.97 \(\rightarrow\) 0.03. The conflict floor itself is information-theoretic (no single model can satisfy contradictory conventions; SI Appendix, Proposition S2), with the framework's role being the \emph{structure around it}: which divergences generate conflict, and what moves the cliff.
|
||||
|
||||
The pre-registered \emph{emergent test} constrains the claim most: true BDM incompatibilities are emergent (each lineage's changes harmless alone), so I let children diverge with \emph{no conflicting signal anywhere}, using complementary class specialists and divergent input conventions, to 6.4\(\times\) the base training. No isolation emerged (residual 0.000 throughout); instead the merge \emph{rescued} the two catastrophically-forgetting specialists (parents \(\approx\) 0.50, merge \(\approx\) 0.955, a sustained Fisher--Muller rescue). The same double result appears at the language-model tier (Fig. 5 E and F): conflicting conventions produce \emph{function-specific} hybrid breakdown (the merge scores below both parents on the conflicted function, while a budget-controlled design shows the disjoint skills merge unharmed), and over-training disjoint specialists 1\(\rightarrow\)12 epochs (cf. the merging literature's expert-duration effect; 46) produces no isolation at all --- the merge improves. Across every tier tested, isolation had to be provoked by functional conflict; specialisation alone did not speciate --- a bound on the analogy that sharpens the design rule: what breaks merging is conflicting conventions on shared circuitry, not divergence per se.
|
||||
The pre-registered \emph{emergent test} constrains the claim most: true BDM incompatibilities are emergent (each lineage's changes harmless alone), so I let children diverge with \emph{no conflicting signal anywhere}, using complementary class specialists and divergent input conventions, to 6.4\(\times\) the base training. No isolation emerged (residual 0.000 throughout); instead the merge \emph{rescued} the two catastrophically-forgetting specialists (parents \(\approx\) 0.50, merge \(\approx\) 0.955, a sustained Fisher--Muller rescue). The same double result appears at the language-model tier (Fig. 7 E and F): conflicting conventions produce \emph{function-specific} hybrid breakdown (the merge scores below both parents on the conflicted function, while a budget-controlled design shows the disjoint skills merge unharmed), and over-training disjoint specialists 1\(\rightarrow\)12 epochs (cf. the merging literature's expert-duration effect; 46) produces no isolation at all --- the merge improves. Across every tier tested, isolation had to be provoked by functional conflict; specialisation alone did not speciate --- a bound on the analogy that sharpens the design rule: what breaks merging is conflicting conventions on shared circuitry, not divergence per se.
|
||||
|
||||
\begin{figure*}[p]\centering % fig5
|
||||
\includegraphics[width=\textwidth]{figs/fig5.pdf}
|
||||
\caption{Model speciation at three tiers. (A) Analytic model: hybrid fitness tracks the parents while lineages are compatible, then falls to inviability; the denser the incompatibilities, the earlier the fall. (B) The isolation cliff: probability of hybrid inviability against divergence, by incompatibility density. (C) Trained networks: the merge error barrier between two MLPs before and after permutation-and-rescaling alignment --- the same-task/different-start barrier is a coordinate artefact (removed by alignment); the conflicting-task barrier is left essentially unchanged. (D) Sweeping the fraction of conflicting classes: the residual barrier rises while merged-model accuracy falls from 0.97 to 0.03. (E) Language models (0.5B LoRA children of a shared base): on shared ambiguous prompts each parent performs under its own convention while the merged model falls below both --- function-specific hybrid breakdown. (F) Divergence without conflict: over-training disjoint specialists from 1 to 12 epochs produces no isolation; the merged model tracks or exceeds the parents throughout.}\label{fig5}
|
||||
\begin{figure*}[p]\centering % fig7
|
||||
\includegraphics[width=\textwidth]{figs/fig7.pdf}
|
||||
\caption{Model speciation at three tiers. (A) Analytic model: hybrid fitness tracks the parents while lineages are compatible, then falls to inviability; the denser the incompatibilities, the earlier the fall. (B) The isolation cliff: probability of hybrid inviability against divergence, by incompatibility density. (C) Trained networks: the merge error barrier between two MLPs before and after permutation-and-rescaling alignment --- the same-task/different-start barrier is a coordinate artefact (removed by alignment); the conflicting-task barrier is left essentially unchanged. (D) Sweeping the fraction of conflicting classes: the residual barrier rises while merged-model accuracy falls from 0.97 to 0.03. (E) Language models (0.5B LoRA children of a shared base): on shared ambiguous prompts each parent performs under its own convention while the merged model falls below both --- function-specific hybrid breakdown. (F) Divergence without conflict: over-training disjoint specialists from 1 to 12 epochs produces no isolation; the merged model tracks or exceeds the parents throughout.}\label{fig7}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{A controlled predictive test: functional conflict, measured pre-merge, predicts merge damage}
|
||||
|
||||
The framework's prediction-level claim was put to a designed test (Fig. 6C). Thirty-nine parent pairs (13 conditions \(\times\) 3 seeds; rows are not independent --- parents share task-data seeds across conditions, so inference is condition-clustered, and because shared seeds also couple rows \emph{across} conditions I report per-seed and leave-one-seed-out sensitivity alongside) span three axes decorrelated by construction: \emph{conflict} (contradictory conventions on shared prompts, private budgets fixed), \emph{compatible overlap} (the same shared prompts under the same convention --- overlap and volume without conflict), and \emph{duration} (weight divergence with zero conflict). Before merging, six predictors are computed: \emph{confidence-weighted functional conflict} (bilateral confident disagreement on probes drawn blind to where conflict lives --- a proposed proxy for merge-relevant interactions, motivated by the observation that raw disagreement counts harmless complementation, one parent merely ignorant, as conflict), raw disagreement, gradient alignment at the shared base (47), LoRA-delta cosine and distance, and a cross-task performance baseline. The pre-registered outcome is the merge penalty against oracle parent potential (the analogue of \emph{hybrid load}, the fitness a hybrid loses relative to what its parents' genes could jointly supply), also reported against best- and mean-parent references because the predictor ordering is sensitive to that choice.
|
||||
The framework's prediction-level claim was put to a designed test (Fig. 3C). Thirty-nine parent pairs (13 conditions \(\times\) 3 seeds; rows are not independent --- parents share task-data seeds across conditions, so inference is condition-clustered, and because shared seeds also couple rows \emph{across} conditions I report per-seed and leave-one-seed-out sensitivity alongside) span three axes decorrelated by construction: \emph{conflict} (contradictory conventions on shared prompts, private budgets fixed), \emph{compatible overlap} (the same shared prompts under the same convention --- overlap and volume without conflict), and \emph{duration} (weight divergence with zero conflict). Before merging, six predictors are computed: \emph{confidence-weighted functional conflict} (bilateral confident disagreement on probes drawn blind to where conflict lives --- a proposed proxy for merge-relevant interactions, motivated by the observation that raw disagreement counts harmless complementation, one parent merely ignorant, as conflict), raw disagreement, gradient alignment at the shared base (47), LoRA-delta cosine and distance, and a cross-task performance baseline. The pre-registered outcome is the merge penalty against oracle parent potential (the analogue of \emph{hybrid load}, the fitness a hybrid loses relative to what its parents' genes could jointly supply), also reported against best- and mean-parent references because the predictor ordering is sensitive to that choice.
|
||||
|
||||
Across this controlled grid, pre-merge functional disagreement predicted merge penalties (clustered bootstrap CIs excluding zero; held-out leave-one-condition-out \(\rho\) \(\approx\) 0.35--0.40), whereas LoRA-delta cosine and L2 showed no statistically detectable association; gradient alignment carried intermediate signal. Head-to-head predictor differences are not individually significant at this sample size; only these baselines were tested; and with three seeds, uncertainty about seed generalisation remains substantial --- though the seed sensitivity favours the functional measures (per-seed \(\rho\) stable at +0.37 to +0.53 in each seed alone, geometry \(\approx\) 0 in every seed, gradient alignment seed-unstable at −0.11 to −0.55). Two further results bound the claim: the initial two-axis grid's best predictor was delta-cosine (\(\rho\) = +0.60) --- an overlap artefact that the compatible-overlap control was added to expose, and did (collapse to +0.03); and the pre-registered internal prediction that confidence weighting would beat raw disagreement \emph{failed} (they are statistically indistinguishable as rank predictors), so the present evidence favours functional disagreement generally, not the DMI-specific refinement. The framework motivated the measurement and the controls; their success does not validate the specifically population-genetic mechanism. Whether the prediction improves a budget-matched operator choice, and whether it generalises to unfamiliar conflict structures and real task pairs, are the experiment's open front.
|
||||
|
||||
\begin{figure*}[p]\centering % fig6
|
||||
\includegraphics[width=\textwidth]{figs/fig6.pdf}
|
||||
\caption{The language-model tier. (A) Seed-replicated merging (0.5B, five seeds, fixed test sets; mean $\pm$ 95\% CI): merged specialists exceed the best single specialist overall, and only merged models are competent on every task family. (B) Hard, unsaturated tasks at 7B (single run): the weight-average dilutes a fragile specialist below the best single parent; routing among intact specialists preserves it. (C) The controlled predictive test (13 conditions $\times$ 3 seeds): pre-merge confidence-weighted functional conflict against merge penalty, coloured by grid axis --- penalty concentrates on the conflict axis. (D) Predictor comparison, $|$Spearman $\rho|$ against merge penalty over the full grid: functional measures carry signal, the tested weight-geometry baselines do not; paired differences between predictors are not individually significant.}\label{fig6}
|
||||
\end{figure*}
|
||||
|
||||
\textbf{Table 2.} Headline quantitative results with sample sizes, uncertainty, and outcome definitions (full per-experiment tables and falsifier status in SI Appendix and per-experiment documentation).
|
||||
|
||||
\medskip\noindent\begin{center}\footnotesize
|
||||
\begin{tabular}{p{0.230\textwidth} p{0.230\textwidth} p{0.230\textwidth} p{0.230\textwidth}}
|
||||
\hline
|
||||
Result & Setting / n & Outcome definition & Headline \\ \hline
|
||||
Closed-form validation & Analytic tier; standing tests & Simulated vs closed-form H-decay, immigration equilibrium, multi-teacher union & Agreement < 0.5\% \\[3pt]
|
||||
Grounding retention & Minimal model; 18+ replicates per point & Fraction of equilibrium diversity retained at grounding g (operational threshold) & g \(\approx\) 0.05 retained \(\geq\)95\% (tested setting); smooth in g \\[3pt]
|
||||
MNIST collapse \& rescue & Conv-VAE, 4 replicates; frozen oracle (98.5\% mode acc.) & Mode support / forward-KL over generations & Dry: 30\(\rightarrow\)1 modes; 10\% grounding: 30/30 held \\[3pt]
|
||||
Fisher--Muller in LLMs & 5 seeds (0.5B), fixed tests; single 7B run & Merged vs best-specialist accuracy (overall; worst family) & Ties 0.647±0.027 vs 0.592±0.009; 7B 0.87 vs 0.77 \\[3pt]
|
||||
Union vs blend (headroom) & 3 seeds (0.5B hard); single 7B-hard run & Paired per-seed ordering, routing vs weight-average & Routing > blend in 3/3 seeds; one catastrophic blend failure avoided \\[3pt]
|
||||
Speciation decomposition & MLPs, 3 replicates & LMC error barrier residual after permutation+rescaling alignment & Same-task 0.001; conflict 0.497 (naive 0.502) \\[3pt]
|
||||
Emergent isolation & MLPs 4 reps to 6.4\(\times\) base training; LLM 1\(\rightarrow\)12 epochs & Residual barrier; merged vs parent accuracy & 0.000 everywhere; merge rescues parents (\(\approx\)0.955 vs \(\approx\)0.50) \\[3pt]
|
||||
Predictive test & 13 conditions \(\times\) 3 seeds (0.5B) & Merge penalty vs oracle parent potential (pre-registered; ±: clustered 95\% CI) & Functional \(\rho\) +0.45/+0.46, CI excl. 0; LOCO \(\rho\) \(\approx\) 0.4; geometry n.s.; paired differences n.s. \\[3pt]
|
||||
\hline\end{tabular}\end{center}\medskip
|
||||
Headline quantitative results, with sample sizes, uncertainty, and outcome definitions, are collected in SI Appendix, Table S2.
|
||||
|
||||
\section*{Discussion}
|
||||
|
||||
|
|
|
|||
|
|
@ -22,6 +22,16 @@ OUT = HERE / "body.tex"
|
|||
# figure name -> (single publication PDF from make_figs.py, caption)
|
||||
FIGURES: dict[str, tuple[list[str], str]] = {
|
||||
"fig1": (["paper/pnas/figs/fig1.pdf"],
|
||||
"The experimental programme. Each population-genetic abstraction (Table 1) is tested at up "
|
||||
"to three tiers of model architecture, ordered left to right by increasing realism: an exact "
|
||||
"Wright--Fisher simulator over knowledge distributions (closed forms; bitwise-reproducible), "
|
||||
"trained neural networks measured against exact oracles (recurrent, feedforward, and "
|
||||
"variational-autoencoder generators on a synthetic mode universe, and a convolutional VAE on "
|
||||
"MNIST), and language models (LoRA specialists on Qwen bases at 0.5B and 7B, scored by an "
|
||||
"exact-match verifier). Rows are the paper's five questions; filled cells name the "
|
||||
"experiments run at each tier; dashed cells were not tested, and the composed society at "
|
||||
"language-model scale is the paper's stated gap."),
|
||||
"fig2": (["paper/pnas/figs/fig2.pdf"],
|
||||
"Grounding is immigration. (A) Stationary diversity against the grounding fraction in the "
|
||||
"minimal inheritance model: simulation (points, 95\\% CI) matches the exact immigration--drift "
|
||||
"equilibrium (dashed). The equilibrium is smooth in $g$; $g \\approx 0.05$ marks the "
|
||||
|
|
@ -30,7 +40,7 @@ FIGURES: dict[str, tuple[list[str], str]] = {
|
|||
"is zero). (B) The same signs on real images: samples from a convolutional VAE retrained each "
|
||||
"generation on its own output (rows: generations 0--15 of an ungrounded lineage) collapse "
|
||||
"toward a single blurred mode; 10\\% grounding holds all thirty modes (quantified in SI)."),
|
||||
"fig2": (["paper/pnas/figs/fig2.pdf"],
|
||||
"fig4": (["paper/pnas/figs/fig4.pdf"],
|
||||
"Recombination in the minimal model: blending inheritance and the Fisher--Muller effect. "
|
||||
"(A) Expected rare-capability survival in a child refit from $K$ uncorrelated parents: the "
|
||||
"output-mean (blending) stays at the single-parent level --- the first-order cancellation --- "
|
||||
|
|
@ -38,7 +48,7 @@ FIGURES: dict[str, tuple[list[str], str]] = {
|
|||
"with parent count. (B) Multi-locus recombination of decorrelated specialists produces "
|
||||
"offspring fitter than any parent, approaching the optimum as parents are added; the best "
|
||||
"single parent and the blended average plateau below (mean $\\pm$ 95\\% CI)."),
|
||||
"fig3": (["paper/pnas/figs/fig3.pdf"],
|
||||
"fig5": (["paper/pnas/figs/fig5.pdf"],
|
||||
"Rugged (epistatic) landscapes: risk, remedy, and population structure. (A) Outbreeding "
|
||||
"depression: the mean offspring of blindly recombined specialist parents falls below the best "
|
||||
"parent, more steeply the more rugged the landscape (NK ruggedness $K$) and the higher the "
|
||||
|
|
@ -47,7 +57,7 @@ FIGURES: dict[str, tuple[list[str], str]] = {
|
|||
"(C) Mating structure: the best champion arises at wide mate-pool breadth on smooth landscapes "
|
||||
"and at intermediate breadth on rugged ones. (D) Wide breadth monotonically erodes population "
|
||||
"diversity at every ruggedness (mean $\\pm$ 95\\% CI, 20 replicates)."),
|
||||
"fig4": (["paper/pnas/figs/fig4.pdf"],
|
||||
"fig6": (["paper/pnas/figs/fig6.pdf"],
|
||||
"The tested society: grounded evaluation, recombination, and diversity preservation make "
|
||||
"complementary contributions. A finite agent population on a rugged NK landscape; selection "
|
||||
"weights true fitness against conformity to the population consensus. (A) Best real fitness: "
|
||||
|
|
@ -55,7 +65,7 @@ FIGURES: dict[str, tuple[list[str], str]] = {
|
|||
"population onto a confident, unfit consensus; removing recombination or diversity "
|
||||
"preservation strands it lower. (B) Population diversity. (C) The self-consumption signature: "
|
||||
"conformity minus true fitness (mean $\\pm$ 95\\% CI, 12 replicates)."),
|
||||
"fig5": (["paper/pnas/figs/fig5.pdf"],
|
||||
"fig7": (["paper/pnas/figs/fig7.pdf"],
|
||||
"Model speciation at three tiers. (A) Analytic model: hybrid fitness tracks the parents while "
|
||||
"lineages are compatible, then falls to inviability; the denser the incompatibilities, the "
|
||||
"earlier the fall. (B) The isolation cliff: probability of hybrid inviability against "
|
||||
|
|
@ -68,7 +78,7 @@ FIGURES: dict[str, tuple[list[str], str]] = {
|
|||
"own convention while the merged model falls below both --- function-specific hybrid "
|
||||
"breakdown. (F) Divergence without conflict: over-training disjoint specialists from 1 to 12 "
|
||||
"epochs produces no isolation; the merged model tracks or exceeds the parents throughout."),
|
||||
"fig6": (["paper/pnas/figs/fig6.pdf"],
|
||||
"fig3": (["paper/pnas/figs/fig3.pdf"],
|
||||
"The language-model tier. (A) Seed-replicated merging (0.5B, five seeds, fixed test sets; mean "
|
||||
"$\\pm$ 95\\% CI): merged specialists exceed the best single specialist overall, and only "
|
||||
"merged models are competent on every task family. (B) Hard, unsaturated tasks at 7B (single "
|
||||
|
|
|
|||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
BIN
paper/pnas/figs/fig7.pdf
Normal file
BIN
paper/pnas/figs/fig7.pdf
Normal file
Binary file not shown.
|
|
@ -1,6 +1,6 @@
|
|||
# The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations
|
||||
|
||||
**Giorgio F. Gilestro** — Department of Life Sciences, Imperial College London. giorgio@gilest.ro
|
||||
**Giorgio F. Gilestro** — Department of Life Sciences, Imperial College London. giorgiogilest.ro
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -92,8 +92,11 @@ convention and by trial-and-error search. They are also, recognisably, machine l
|
|||
problem at a new scale: *continual learning*, the struggle to acquire new abilities without losing old
|
||||
ones (26, 27), transposed from a single network to a population whose members inherit from one
|
||||
another. Population genetics, I will argue, prices these decisions. Table 1 summarises the
|
||||
correspondences on which the argument runs; the sections that follow develop them from closed-form
|
||||
theory to experiments in trained networks and language models.
|
||||
correspondences on which the argument runs, and Fig. 1 maps the experimental programme built on
|
||||
them: the same abstractions tested at three tiers of model architecture — an exact simulator,
|
||||
trained neural networks, and language models — with the sections that follow climbing that ladder.
|
||||
|
||||
*(FIG:fig1)*
|
||||
|
||||
## The minimal model, and where its exactness ends
|
||||
|
||||
|
|
@ -120,7 +123,7 @@ refit) reproduces both. Throughout, a real learner is therefore treated as Wrigh
|
|||
estimator bias*, and the drift signs (rare-first loss; the grounding response)
|
||||
survived that bias in every architecture I tested, including a convolutional VAE retrained on its own
|
||||
generated digits, where the dry lineage collapses to a single blurred digit class while 10% grounding
|
||||
holds all thirty modes (Fig. 1). Retraining on a single parent is *asexual reproduction*, and sustained loss under it carries the defining consequence
|
||||
holds all thirty modes (Fig. 2). Retraining on a single parent is *asexual reproduction*, and sustained loss under it carries the defining consequence
|
||||
of *Muller's ratchet* (28), the mechanism by which lineages that never recombine decay irreversibly —
|
||||
the reason non-recombining genomes such as the Y chromosome have shed most of their ancestral genes.
|
||||
Once every copy of a rare capability is gone from all parents and sources, no recombination can
|
||||
|
|
@ -171,11 +174,11 @@ or stratified sampling changes that cost, and recombination can recover rare cap
|
|||
still retained across complementary parents* (next section). In trained networks the *sign* of the grounding response transfers everywhere I
|
||||
looked, with two deviations, both traced to the estimator bias above: sharp thresholds soften,
|
||||
and support-counting metrics decouple from truth (forward-KL is the operative collapse metric for a
|
||||
smoothing learner). On real images (Fig. 1B), dry self-training collapses a convolutional VAE to one
|
||||
smoothing learner). On real images (Fig. 2B), dry self-training collapses a convolutional VAE to one
|
||||
mode while ~10% grounding holds all thirty (the trained model needs roughly twice the exact-operator
|
||||
fraction, the measured price of the estimator bias).
|
||||
|
||||
*(FIG:fig1)*
|
||||
*(FIG:fig2)*
|
||||
|
||||
### Recombination: a conservation law, its operators, and offspring that exceed every parent
|
||||
|
||||
|
|
@ -206,7 +209,7 @@ union-preserving operators beat the weight-average where that average falls shor
|
|||
performance, and add nothing where it does not (easy-versus-hard contrasts at two scales; a
|
||||
quantitative form of the relationship is untested). On easy tasks a capable base's average is already at ceiling and refinements
|
||||
add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent
|
||||
and routing wins by a wide margin (Fig. 6A–B).
|
||||
and routing wins by a wide margin (Fig. 3A–B).
|
||||
|
||||
The generative payoff is the *Fisher–Muller effect* (35, 36), the classical account of why sex speeds
|
||||
adaptation: in an asexual population, beneficial variants arising in different individuals can only
|
||||
|
|
@ -215,16 +218,16 @@ producing a *genotype* (an individual's combination of variants, one at each *lo
|
|||
fitter than any parent.
|
||||
In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum, a
|
||||
genotype no parent held, while the best single parent and the blended average both plateau below
|
||||
(Fig. 2). In real language models the signature replicates under seed replication: merges of three
|
||||
(Fig. 4). In real language models the signature replicates under seed replication: merges of three
|
||||
LoRA (37) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper
|
||||
worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 6A).
|
||||
worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 3A).
|
||||
|
||||
Sex has risks and, for AI, an unfair advantage, both quantified on Kauffman's NK fitness landscapes
|
||||
(38), the standard model of *epistasis*, biology's term for interaction between genes: the fitness
|
||||
contribution of a variant depends on which variants occupy the other loci, much as a component's
|
||||
value in an ML system depends on the components around it. Each of the landscape's `N` sites
|
||||
interacts with `K` others (the model's eponymous parameters), and raising that interaction count
|
||||
tunes the landscape from smooth and additive to rugged and many-peaked (Fig. 3). When skills are entangled, blind recombination produces offspring *below* their parents,
|
||||
tunes the landscape from smooth and additive to rugged and many-peaked (Fig. 5). When skills are entangled, blind recombination produces offspring *below* their parents,
|
||||
worsening with ruggedness, and the optimal recombination rate shrinks as entanglement grows. Biology
|
||||
knows this failure as *outbreeding depression*, the reason conservation practice warns against
|
||||
crossing locally adapted populations: in the textbook case, an ibex herd in the Tatra Mountains
|
||||
|
|
@ -234,21 +237,23 @@ parents, choose complementary mates, and *screen many candidate offspring agains
|
|||
keeping one*. This directed sex converts the outbreeding catastrophe into a reliable gain in the
|
||||
model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language
|
||||
models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks, including
|
||||
one seed where the blend failed catastrophically and selection was immune (Fig. 6A). Finally,
|
||||
one seed where the blend failed catastrophically and selection was immune (Fig. 3A). Finally,
|
||||
population *structure* is itself a knob: sweeping the mate-pool breadth from monogamous (repeated
|
||||
local pairings) to promiscuous (*panmixia*: any model may merge with any other) against ruggedness,
|
||||
wide mixing maximises the population mean while
|
||||
monotonically destroying diversity, and the best *champion* shifts from wide breadth on smooth
|
||||
landscapes to intermediate breadth on rugged ones (Fig. 3C), the mating-system phenomenon known to
|
||||
landscapes to intermediate breadth on rugged ones (Fig. 5C), the mating-system phenomenon known to
|
||||
structured-population search, mapped onto merging populations.
|
||||
|
||||
*(FIG:fig2)*
|
||||
|
||||
*(FIG:fig3)*
|
||||
|
||||
*(FIG:fig4)*
|
||||
|
||||
*(FIG:fig5)*
|
||||
|
||||
### The society: grounding, recombination, and diversity make complementary contributions
|
||||
|
||||
Composing the operators (Fig. 4) requires one definitional distinction first. In the inheritance
|
||||
Composing the operators (Fig. 6) requires one definitional distinction first. In the inheritance
|
||||
model, grounding is *grounded inheritance*: external samples added to the reproduction process (the
|
||||
data channel). In the society model, grounding is *grounded evaluation*: selection weights true
|
||||
fitness against conformity to the population's own consensus, `g`·true-fitness + (1−g)·conformity,
|
||||
|
|
@ -264,7 +269,7 @@ make complementary contributions *under the tested conditions*; general joint ne
|
|||
established (alternative mutation, restart, archive, or selection schemes could alter the picture). At
|
||||
language-model scale this composed loop remains unbuilt; it is the paper's largest stated gap.
|
||||
|
||||
*(FIG:fig4)*
|
||||
*(FIG:fig6)*
|
||||
|
||||
### The limit of sex: model speciation
|
||||
|
||||
|
|
@ -273,7 +278,7 @@ separate species (*reproductive isolation*) through Bateson–Dobzhansky–Mulle
|
|||
changes harmless on their own genetic background but deleterious in combination — the mechanism behind
|
||||
the mule's sterility and the inviability of many between-species crosses, in which two genomes that
|
||||
each work perfectly cannot run in the same cell. A merged model is exactly the
|
||||
exposed hybrid. I built the analytic model (Fig. 5A): hybrid fitness tracks the parents while
|
||||
exposed hybrid. I built the analytic model (Fig. 7A): hybrid fitness tracks the parents while
|
||||
compatible, then peels off and crashes below the ancestor; the isolation cliff arrives earlier the
|
||||
denser the incompatibilities; and the incompatibility *count* snowballs quadratically with divergence
|
||||
(42). Note that a super-linear count does not by itself entail a sharp performance cliff without
|
||||
|
|
@ -284,7 +289,7 @@ In trained networks, the claim must survive a known alternative: merge barriers
|
|||
trained networks are famously *coordinate artefacts*, removable by re-aligning hidden units (43);
|
||||
richer symmetry groups remove more (44), with known failures beyond the shared-data regime (45). I therefore aligned under the composition of
|
||||
permutation matching and exact per-unit rescaling (the unit symmetry group of plain ReLU MLPs, as the
|
||||
search space) and decomposed the barrier (Fig. 5 C and D): two networks trained from different
|
||||
search space) and decomposed the barrier (Fig. 7 C and D): two networks trained from different
|
||||
initialisations on the *same* task have a barrier that this alignment removes essentially entirely
|
||||
(residual ≈ 0.001, the aligned merge performing at parent level): coordinate, not functional; two
|
||||
networks trained on *conflicting* label maps have a barrier the same alignment leaves largely
|
||||
|
|
@ -302,7 +307,7 @@ emergent (each lineage's changes harmless alone), so I let children diverge with
|
|||
signal anywhere*, using complementary class specialists and divergent input conventions, to 6.4× the base
|
||||
training. No isolation emerged (residual 0.000 throughout); instead the merge *rescued* the two
|
||||
catastrophically-forgetting specialists (parents ≈ 0.50, merge ≈ 0.955, a sustained Fisher–Muller
|
||||
rescue). The same double result appears at the language-model tier (Fig. 5 E and F): conflicting conventions
|
||||
rescue). The same double result appears at the language-model tier (Fig. 7 E and F): conflicting conventions
|
||||
produce *function-specific* hybrid breakdown (the merge scores below both parents on the conflicted
|
||||
function, while a budget-controlled design shows the disjoint skills merge unharmed), and over-training
|
||||
disjoint specialists 1→12 epochs (cf. the merging literature's expert-duration effect; 46) produces
|
||||
|
|
@ -311,11 +316,11 @@ tested, isolation had to be provoked by functional conflict; specialisation alon
|
|||
— a bound on the analogy that sharpens the design rule: what breaks merging is conflicting conventions
|
||||
on shared circuitry, not divergence per se.
|
||||
|
||||
*(FIG:fig5)*
|
||||
*(FIG:fig7)*
|
||||
|
||||
### A controlled predictive test: functional conflict, measured pre-merge, predicts merge damage
|
||||
|
||||
The framework's prediction-level claim was put to a designed test (Fig. 6C). Thirty-nine parent pairs
|
||||
The framework's prediction-level claim was put to a designed test (Fig. 3C). Thirty-nine parent pairs
|
||||
(13 conditions × 3 seeds; rows are not independent — parents share task-data seeds across conditions,
|
||||
so inference is condition-clustered, and because shared seeds also couple rows *across* conditions I
|
||||
report per-seed and leave-one-seed-out sensitivity alongside) span three axes decorrelated by construction: *conflict*
|
||||
|
|
@ -349,21 +354,7 @@ the specifically population-genetic mechanism. Whether the prediction improves a
|
|||
operator choice, and whether it generalises to unfamiliar conflict structures and real task pairs,
|
||||
are the experiment's open front.
|
||||
|
||||
*(FIG:fig6)*
|
||||
|
||||
**Table 2.** Headline quantitative results with sample sizes, uncertainty, and outcome definitions
|
||||
(full per-experiment tables and falsifier status in SI Appendix and per-experiment documentation).
|
||||
|
||||
| Result | Setting / n | Outcome definition | Headline |
|
||||
|---|---|---|---|
|
||||
| Closed-form validation | Analytic tier; standing tests | Simulated vs closed-form H-decay, immigration equilibrium, multi-teacher union | Agreement < 0.5% |
|
||||
| Grounding retention | Minimal model; 18+ replicates per point | Fraction of equilibrium diversity retained at grounding g (operational threshold) | g ≈ 0.05 retained ≥95% (tested setting); smooth in g |
|
||||
| MNIST collapse & rescue | Conv-VAE, 4 replicates; frozen oracle (98.5% mode acc.) | Mode support / forward-KL over generations | Dry: 30→1 modes; 10% grounding: 30/30 held |
|
||||
| Fisher–Muller in LLMs | 5 seeds (0.5B), fixed tests; single 7B run | Merged vs best-specialist accuracy (overall; worst family) | Ties 0.647±0.027 vs 0.592±0.009; 7B 0.87 vs 0.77 |
|
||||
| Union vs blend (headroom) | 3 seeds (0.5B hard); single 7B-hard run | Paired per-seed ordering, routing vs weight-average | Routing > blend in 3/3 seeds; one catastrophic blend failure avoided |
|
||||
| Speciation decomposition | MLPs, 3 replicates | LMC error barrier residual after permutation+rescaling alignment | Same-task 0.001; conflict 0.497 (naive 0.502) |
|
||||
| Emergent isolation | MLPs 4 reps to 6.4× base training; LLM 1→12 epochs | Residual barrier; merged vs parent accuracy | 0.000 everywhere; merge rescues parents (≈0.955 vs ≈0.50) |
|
||||
| Predictive test | 13 conditions × 3 seeds (0.5B) | Merge penalty vs oracle parent potential (pre-registered; ±: clustered 95% CI) | Functional ρ +0.45/+0.46, CI excl. 0; LOCO ρ ≈ 0.4; geometry n.s.; paired differences n.s. |
|
||||
Headline quantitative results, with sample sizes, uncertainty, and outcome definitions, are collected in SI Appendix, Table S2.
|
||||
|
||||
## Discussion
|
||||
|
||||
|
|
|
|||
Binary file not shown.
|
|
@ -1,7 +1,7 @@
|
|||
"""Publication figures for the PNAS draft — unified, lettered, codename-free.
|
||||
|
||||
Re-plots every panel directly from the committed results artifacts into six single-file figures
|
||||
(figs/fig1.pdf .. fig6.pdf): no experiment codenames, no suptitles, no per-panel headline titles
|
||||
Renders fig1 (the experimental-programme schematic) and re-plots every data panel directly from the
|
||||
committed results artifacts (figs/fig2.pdf .. fig7.pdf): no experiment codenames, no suptitles, no per-panel headline titles
|
||||
(interpretation lives in the captions), bold panel letters, one consistent style. The per-experiment
|
||||
figures under results/ remain the exploratory versions; these are the manuscript's.
|
||||
|
||||
|
|
@ -42,8 +42,93 @@ def save(fig, name):
|
|||
print("wrote", OUT / f"{name}.pdf")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- fig 1: grounding + MNIST
|
||||
# ---------------------------------------------------------------- fig 1: experimental programme
|
||||
def fig1():
|
||||
from matplotlib.patches import FancyBboxPatch
|
||||
|
||||
TIERS = [
|
||||
("Exact model", "Wright\u2013Fisher simulator (NumPy)", "closed forms \u00b7 bitwise-reproducible",
|
||||
"#4292c6", "#eaf2fa"),
|
||||
("Trained networks", "RNN \u00b7 MLP \u00b7 VAE on a synthetic oracle;\nconvolutional VAE on MNIST",
|
||||
"sign-level tests \u00b7 exact oracles", "#41ab5d", "#edf8ea"),
|
||||
("Language models", "LoRA specialists on Qwen 0.5B & 7B;\nexact-match verifier",
|
||||
"seed-replicated signs", "#e6550d", "#fdf0e6"),
|
||||
]
|
||||
ROWS = [
|
||||
("Grounding", "how much real data?",
|
||||
["immigration\u2013drift equilibrium:\n$g \\approx 0.05$ retains $\\geq$95% diversity;\nobservation floor $1-e^{-mp}$",
|
||||
"collapse & rescue in every\narchitecture; MNIST: dry 30$\\to$1 modes,\n10% grounding holds 30/30;\nestimator-bias learning kernel",
|
||||
None]),
|
||||
("Recombination", "blend or merge?",
|
||||
["blending conservation law\n(first-order cancellation);\nunion-operator gain; Fisher\u2013Muller",
|
||||
"merge rescues two forgetting\nspecialists ($\\approx$0.50 $\\to$ 0.955)",
|
||||
"merged specialists beat every parent\n(5 seeds at 0.5B; 7B); routing vs\naveraging: the headroom rule"]),
|
||||
("Entangled skills", "who merges with whom?",
|
||||
["NK landscapes: outbreeding\ndepression; directed sex restores\nthe gain; mate-pool breadth optimum",
|
||||
None,
|
||||
"bred-and-screened offspring beat\nthe blind blend in every seed\n(hard, unsaturated tasks)"]),
|
||||
("The composed society", "can the loop sustain itself?",
|
||||
["four-arm ablation: grounding, sex,\ndiversity each removed\n$\\to$ three distinct failures",
|
||||
None,
|
||||
"OPEN"]),
|
||||
("Speciation & prediction", "when does merging fail?",
|
||||
["BDM incompatibility model:\nisolation cliff; quadratic snowball",
|
||||
"barrier decomposition under\npermutation+rescaling; conflict\nsweep 0.97$\\to$0.03; emergent null",
|
||||
"convention conflict $\\to$ hybrid\nbreakdown; duration null; pre-merge\npredictive test (13 cond. $\\times$ 3 seeds)"]),
|
||||
]
|
||||
|
||||
fig, ax = plt.subplots(figsize=(11.4, 5.4))
|
||||
ax.set_axis_off()
|
||||
ax.set_xlim(0, 1)
|
||||
ax.set_ylim(0, 1)
|
||||
x0, gap = 0.16, 0.008
|
||||
cw = (1.0 - x0) / 3
|
||||
row_h, row_top = 0.152, 0.79
|
||||
|
||||
ax.annotate("", xy=(0.995, 0.975), xytext=(x0 + 0.02, 0.975),
|
||||
arrowprops=dict(arrowstyle="->", color="#555", lw=1.1))
|
||||
ax.text(x0 + (1 - x0) / 2, 0.988, "the same population-genetic abstractions (Table 1), increasing realism",
|
||||
ha="center", va="bottom", fontsize=8, style="italic", color="#333")
|
||||
|
||||
for j, (name, arch, guarantee, edge, face) in enumerate(TIERS):
|
||||
x = x0 + j * cw
|
||||
ax.add_patch(FancyBboxPatch((x + gap, 0.795), cw - 2 * gap, 0.16,
|
||||
boxstyle="round,pad=0.004", fc=face, ec=edge, lw=1.4))
|
||||
ax.text(x + cw / 2, 0.944, name, ha="center", va="top", fontsize=9, fontweight="bold", color=edge)
|
||||
ax.text(x + cw / 2, 0.902, arch, ha="center", va="top", fontsize=6.8, linespacing=1.3)
|
||||
ax.text(x + cw / 2, 0.803, guarantee, ha="center", va="bottom", fontsize=6.4,
|
||||
style="italic", color="#555")
|
||||
|
||||
for i, (label, question, cells) in enumerate(ROWS):
|
||||
y1 = row_top - i * row_h
|
||||
y0 = y1 - row_h + 2 * gap
|
||||
yc = (y0 + y1) / 2
|
||||
ax.text(0.0, yc + 0.012, label, ha="left", va="center", fontsize=8, fontweight="bold")
|
||||
ax.text(0.0, yc - 0.022, question, ha="left", va="center", fontsize=6.8, style="italic", color="#555")
|
||||
for j, cell in enumerate(cells):
|
||||
x = x0 + j * cw
|
||||
edge, face = TIERS[j][3], TIERS[j][4]
|
||||
if cell is None:
|
||||
ax.add_patch(FancyBboxPatch((x + gap, y0), cw - 2 * gap, y1 - y0,
|
||||
boxstyle="round,pad=0.004", fc="white", ec="#bbbbbb",
|
||||
lw=0.8, ls=(0, (3, 2))))
|
||||
ax.text(x + cw / 2, yc, "not tested at this tier", ha="center", va="center",
|
||||
fontsize=6.4, style="italic", color="#999")
|
||||
elif cell == "OPEN":
|
||||
ax.add_patch(FancyBboxPatch((x + gap, y0), cw - 2 * gap, y1 - y0,
|
||||
boxstyle="round,pad=0.004", fc="white", ec=edge,
|
||||
lw=0.8, ls=(0, (3, 2))))
|
||||
ax.text(x + cw / 2, yc, "open \u2014 the stated gap", ha="center", va="center",
|
||||
fontsize=6.6, style="italic", color=edge)
|
||||
else:
|
||||
ax.add_patch(FancyBboxPatch((x + gap, y0), cw - 2 * gap, y1 - y0,
|
||||
boxstyle="round,pad=0.004", fc=face, ec=edge, lw=0.9))
|
||||
ax.text(x + cw / 2, yc, cell, ha="center", va="center", fontsize=6.4, linespacing=1.35)
|
||||
save(fig, "fig1")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- fig 2: grounding + MNIST
|
||||
def fig2():
|
||||
from knowledge.analysis import critical_grounding, reduce_to_stationary
|
||||
from knowledge.metrics import heterozygosity
|
||||
from knowledge.truth import make_true_distribution
|
||||
|
|
@ -84,15 +169,21 @@ def fig1():
|
|||
ax = axes[1]
|
||||
from PIL import Image
|
||||
im = np.asarray(Image.open("results/mnist_collapse/mnist_montage.png"))
|
||||
crop = int(im.shape[0] * 0.085) # remove the baked-in title band
|
||||
ax.imshow(im[crop:], interpolation="bilinear")
|
||||
# Strip the baked-in title band and left label margin (raster text is unreadable at panel
|
||||
# size); measured on the committed montage: boxes span y >= 69, x >= 75, row centres below.
|
||||
top, left = 60, 68
|
||||
ax.imshow(im[top:, left:], interpolation="bilinear")
|
||||
for yc, g in zip((101.5, 191.5, 282.0, 372.5, 462.5), (0, 4, 8, 12, 15)):
|
||||
ax.text(-10, yc - top, str(g), ha="right", va="center", fontsize=8.5)
|
||||
ax.text(-0.055, 0.5, "generation", transform=ax.transAxes, rotation=90,
|
||||
ha="center", va="center", fontsize=8.5)
|
||||
ax.set_axis_off()
|
||||
letter(ax, "B", x=-0.02)
|
||||
save(fig, "fig1")
|
||||
save(fig, "fig2")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- fig 2: blending vs union + Fisher–Muller
|
||||
def fig2():
|
||||
# ---------------------------------------------------------------- fig 4: blending vs union + Fisher–Muller
|
||||
def fig4():
|
||||
fig, axes = plt.subplots(1, 2, figsize=(10.6, 3.5))
|
||||
|
||||
df, _ = load_bundle("results/E4")
|
||||
|
|
@ -122,11 +213,11 @@ def fig2():
|
|||
ax.set(xlabel="number of parents", ylabel="offspring capability")
|
||||
ax.legend()
|
||||
letter(ax, "B")
|
||||
save(fig, "fig2")
|
||||
save(fig, "fig4")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- fig 3: rugged landscapes
|
||||
def fig3():
|
||||
# ---------------------------------------------------------------- fig 5: rugged landscapes
|
||||
def fig5():
|
||||
fig, axes = plt.subplots(2, 2, figsize=(10.6, 6.8))
|
||||
|
||||
df9, _ = load_bundle("results/E9")
|
||||
|
|
@ -173,11 +264,11 @@ def fig3():
|
|||
ax.set(xlabel="mate-pool breadth (monogamous → panmictic)", ylabel=ylab)
|
||||
ax.legend(title="ruggedness")
|
||||
letter(ax, L)
|
||||
save(fig, "fig3")
|
||||
save(fig, "fig5")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- fig 4: the society
|
||||
def fig4():
|
||||
# ---------------------------------------------------------------- fig 6: the society
|
||||
def fig6():
|
||||
df, _ = load_bundle("results/E11")
|
||||
arms = [("full", "#2ca02c", "full system"),
|
||||
("no_sex", "#ff7f0e", "no recombination"),
|
||||
|
|
@ -201,11 +292,11 @@ def fig4():
|
|||
ax.legend()
|
||||
ax.set(xlabel="generation", ylabel=ylab)
|
||||
letter(ax, L)
|
||||
save(fig, "fig4")
|
||||
save(fig, "fig6")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- fig 5: speciation, three tiers
|
||||
def fig5():
|
||||
# ---------------------------------------------------------------- fig 7: speciation, three tiers
|
||||
def fig7():
|
||||
fig, axes = plt.subplots(2, 3, figsize=(11.4, 6.6))
|
||||
|
||||
bdm, _ = load_bundle("results/E12")
|
||||
|
|
@ -294,11 +385,11 @@ def fig5():
|
|||
ax.set(xlabel="specialist training (epochs)", ylabel="accuracy", ylim=(0, 1.02))
|
||||
ax.legend()
|
||||
letter(ax, "F")
|
||||
save(fig, "fig5")
|
||||
save(fig, "fig7")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- fig 6: the language-model tier
|
||||
def fig6():
|
||||
# ---------------------------------------------------------------- fig 3: the language-model tier
|
||||
def fig3():
|
||||
import pandas as pd
|
||||
from scipy.stats import spearmanr
|
||||
|
||||
|
|
@ -373,9 +464,9 @@ def fig6():
|
|||
ax.set_xticklabels([l for _, l in preds], fontsize=6.5)
|
||||
ax.set(ylabel="|Spearman ρ| vs merge penalty", ylim=(0, 0.8))
|
||||
letter(ax, "D")
|
||||
save(fig, "fig6")
|
||||
save(fig, "fig3")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
for f in (fig1, fig2, fig3, fig4, fig5, fig6):
|
||||
for f in (fig1, fig2, fig3, fig4, fig5, fig6, fig7):
|
||||
f()
|
||||
|
|
|
|||
|
|
@ -96,6 +96,22 @@ long-horizon over-specialisation erodes mergeability at LLM scale (cf. arXiv:260
|
|||
| Emergent speciation without conflict | **Not observed** (pre-registered) | Shared ancestry, compatible tasks, tested divergences | E13b: residual 0.000; merge rescues specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
|
||||
| Grounding + sex + diversity jointly necessary | Exact-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |
|
||||
|
||||
## SI Table S2: headline quantitative results
|
||||
|
||||
Headline quantitative results with sample sizes, uncertainty, and outcome definitions (full
|
||||
per-experiment tables and falsifier status in the per-experiment documentation).
|
||||
|
||||
| Result | Setting / n | Outcome definition | Headline |
|
||||
|---|---|---|---|
|
||||
| Closed-form validation | Analytic tier; standing tests | Simulated vs closed-form H-decay, immigration equilibrium, multi-teacher union | Agreement < 0.5% |
|
||||
| Grounding retention | Minimal model; 18+ replicates per point | Fraction of equilibrium diversity retained at grounding g (operational threshold) | g ≈ 0.05 retained ≥95% (tested setting); smooth in g |
|
||||
| MNIST collapse & rescue | Conv-VAE, 4 replicates; frozen oracle (98.5% mode acc.) | Mode support / forward-KL over generations | Dry: 30→1 modes; 10% grounding: 30/30 held |
|
||||
| Fisher–Muller in LLMs | 5 seeds (0.5B), fixed tests; single 7B run | Merged vs best-specialist accuracy (overall; worst family) | Ties 0.647±0.027 vs 0.592±0.009; 7B 0.87 vs 0.77 |
|
||||
| Union vs blend (headroom) | 3 seeds (0.5B hard); single 7B-hard run | Paired per-seed ordering, routing vs weight-average | Routing > blend in 3/3 seeds; one catastrophic blend failure avoided |
|
||||
| Speciation decomposition | MLPs, 3 replicates | LMC error barrier residual after permutation+rescaling alignment | Same-task 0.001; conflict 0.497 (naive 0.502) |
|
||||
| Emergent isolation | MLPs 4 reps to 6.4× base training; LLM 1→12 epochs | Residual barrier; merged vs parent accuracy | 0.000 everywhere; merge rescues parents (≈0.955 vs ≈0.50) |
|
||||
| Predictive test | 13 conditions × 3 seeds (0.5B) | Merge penalty vs oracle parent potential (pre-registered; ±: clustered 95% CI) | Functional ρ +0.45/+0.46, CI excl. 0; LOCO ρ ≈ 0.4; geometry n.s.; paired differences n.s. |
|
||||
|
||||
## SI Methods (per tier — full details in the per-experiment READMEs and configs)
|
||||
|
||||
**Analytic tier (E1–E14).** Wright–Fisher simulator over K-item distributions; closed-form validation
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue