Phase 4: PNAS research-article draft (main.md + composed figures + SI skeleton)
paper/pnas/main.md — the manuscript restructured as a research article (~5.6k words main text): significance statement, abstract, introduction (diagnosis conceded; the management thesis; the interpretation/ explanation/prediction ladder with the prediction rung stated as a bounded controlled test), the minimal model with its exactness boundary (learning kernel cited against ourselves), Table 1 dictionary with per-row support levels, a five-step results ladder (grounding floor; conservation law + operator boundaries + Fisher-Muller + directed sex + mating structure; the jointly-necessary society; speciation across three tiers with the emergent null; the controlled predictive test at second-review calibration), discussion (design rules, borrowed-vs-ours ledger, limits with the reviewer's generalisation-before-scale ordering, what biology gets back), brief methods, 30 references. build.py composes 6 figures by stacking committed vector PDFs (bespoke unified re-plots deferred to submission polish); builds clean under tectonic (15 pp incl. 6 full-page figures). si.md: SI skeleton (propositions, claims ledger, per-tier methods, statistics, figure list). Manifesto sections of v6 (institutions, timescales, re-minting) compressed into Discussion per the plan; v6 remains the long-form perspective document. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
parent
a40ace1821
commit
bf4b1c077c
21 changed files with 928 additions and 4 deletions
168
paper/pnas/body.tex
Normal file
168
paper/pnas/body.tex
Normal file
|
|
@ -0,0 +1,168 @@
|
|||
\section*{Significance statement}
|
||||
|
||||
Artificial intelligence is shifting from single, frozen models to populations of models that specialise, are retrained on each other's output, and are combined (``merged'') into new models. Trained on their own output, model lineages degenerate --- a process already recognised as the mathematics of genetic drift. This paper imports the other half of population genetics: the biology of sexual reproduction. It treats model merging as recombination, real data as immigration, and merge failure as reproductive isolation, and tests each correspondence in simulations, small neural networks, and language models. The framework yields design rules --- when to average models, when to keep them separate, how much real data suffices --- and a first controlled test showing that measured functional conflict, not weight distance, predicts when merging fails.
|
||||
|
||||
\section*{Abstract}
|
||||
|
||||
AI development increasingly resembles a population process: models are specialised, retrained on model output, and recombined by weight merging, with an openly evolutionary vocabulary but little use of evolutionary theory. Here we treat multigenerational model populations as systems whose inheritance, diversity, and compatibility must be managed, and transfer the quantitative apparatus of the evolution of sex. We take as settled that training on model output is genetic drift (model collapse). In a minimal inheritance model that is exactly Wright--Fisher --- and measurably Wright--Fisher-plus-bias in trained networks --- we derive and test the remedies: grounding as immigration, with a critical real-data fraction far below one but a per-capability floor that leaves the rarest knowledge unrescuable; recombination, where averaging parents' output distributions exactly cancels the benefit of multiple parents while union-preserving operators realise it; the Fisher--Muller effect, with merged language-model specialists exceeding every parent in replicated experiments; outbreeding depression on rugged task landscapes, converted into reliable gains by directed, offspring-screened recombination; and population structure, where the optimal mating breadth shrinks as skills entangle. Sex has a limit: we introduce model speciation --- merge failure as reproductive isolation --- and show in trained networks that a merge barrier surviving the full function-preserving symmetry group tracks functional conflict, that isolation did not emerge from compatible specialisation, and, in a controlled predictive test, that pre-merge functional disagreement predicts merge damage where weight-geometry baselines do not. We state precisely what is exact, what is measured, and what remains hypothesis.
|
||||
|
||||
\medskip\hrule\medskip
|
||||
|
||||
\section*{Introduction}
|
||||
|
||||
The unit of AI progress is quietly changing. Multi-agent systems arrange many models across \emph{space} --- specialists cooperating on a task. A newer axis is \emph{time}: populations of models that persist across generations, each new model built from older ones --- specialised by fine-tuning, trained on data earlier models generated, and, increasingly, produced by \textbf{model merging}, the direct combination of trained weights (1, 2). The engineering literature describes this openly in evolutionary vocabulary --- ``crossover,'' ``mutation,'' ``mate choice,'' populations of merging models that climb benchmarks (2--5) --- but as metaphor over search algorithms. The organising claim of this paper is that the vocabulary deserves its mathematics: \textbf{multigenerational model populations are systems whose inheritance, diversity, and compatibility must be managed --- not merely collections of models to optimise --- and the branch of biology that studies exactly this problem, the population genetics of the evolution of sex, transfers as a quantitative framework.}
|
||||
|
||||
One half of the transfer is settled and is not our contribution. Training each generation of a model on the previous generation's output degrades it --- \emph{model collapse}: rare capabilities vanish first and the lineage drifts toward its own most common behaviour (6). That this is the mathematics of \textbf{genetic drift} in a finite population is now established from several directions (7--9); a closed-form first-extinction law even places collapse onset at the Wright--Fisher first-extinction time (8). We cite this literature as the diagnosis and build on it.
|
||||
|
||||
Our contribution is on the remedy side, and we are explicit about what kind of contribution each claim is, distinguishing \textbf{interpretation} (an existing result understood in population-genetic terms), \textbf{explanation} (the transferred mechanism accounts for observations existing accounts leave open), and \textbf{prediction} (the framework forecasts an unmeasured outcome). The paper is strongest on the first; makes concrete progress on the second --- separating merge failures that are coordinate artefacts from those that are functional; and reports a first, bounded step on the third --- a controlled predictive test in which pre-merge functional-disagreement measures, chosen by the framework, predicted merge damage on a constructed task grid while weight-geometry baselines did not.
|
||||
|
||||
The correspondences we develop, summarised in Table 1: single-teacher retraining is \textbf{asexual reproduction}, and the irreversible arm of its decay corresponds to \textbf{Muller's ratchet} (10) --- once every copy of a rare capability is gone from all parents and sources, no recombination can rebuild it, which is precisely why remedies must act before fixation-by-loss. Injecting verified real data is \textbf{immigration} from a non-drifting source (11--13). Model merging is \textbf{recombination}, and its celebrated payoff --- a merged model exceeding every parent --- is the \textbf{Fisher--Muller effect} (14, 15). Merging entangled skills courts \textbf{outbreeding depression}; screening many candidate merges is a form of \textbf{directed sex} with no biological analogue; restricting who merges with whom is \textbf{population structure}. And merging's hard limit --- models too diverged in function to combine --- is \textbf{reproductive isolation}, for which the Bateson--Dobzhansky--Muller theory of incompatibilities (16, 17) supplies the structure. The nearest precursor to this programme reads sex as an algorithm for mixability in the theory of computation (18), pre-dating model merging; the model-merging literature itself has strong empirical operators (1, 19, 20) and emerging merge-success predictors (21, 22), to which our delta is mechanism: \emph{when and why} failure is coordinate versus functional, and what moves the boundary.
|
||||
|
||||
We support the framework at three tiers of evidence, in ascending realism and descending exactness: a \textbf{minimal analytic model} validated against closed forms to a fraction of a percent; \textbf{small trained networks} (MLPs, recurrent networks, an MNIST image generator) where the operators are measured in real weights; and \textbf{language models} (LoRA-specialised Qwen models, 0.5B locally and 7B on a compute cluster) where the claims are tested as signs under seed replication. Throughout, we report negative and tempering results with the same prominence as confirmations: they include the failure of an internal pre-registered prediction, a null on emergent speciation that bounds the analogy, and the sensitivity analyses that temper the predictive test.
|
||||
|
||||
\section*{The minimal model, and where its exactness ends}
|
||||
|
||||
Knowledge is modelled as a distribution \texttt{p\_t} over \texttt{K} discrete items --- capabilities, facts, modes of behaviour --- with a fixed true distribution \texttt{p*} whose rare tail carries the knowledge most at risk. One generation is: *draw \texttt{n} samples from the parent's distribution, optionally mix in \texttt{m} verified real samples (``grounding'', \texttt{g = m/(n+m)}), and refit the child*. In this minimal inheritance model the resampling step \textbf{is} the Wright--Fisher process --- the same equations, which we exploit as an engineering gate: our simulator reproduces the classical closed forms (heterozygosity decay \texttt{E[H\_t] = H\_0(1 − 1/n)\textasciicircum{}t}; the exact immigration--drift equilibrium; the closed-form multi-teacher union) to within 0.5\%, and these are standing tests in the codebase, not one-off checks.
|
||||
|
||||
The boundary of the exactness matters, and we measured it rather than assumed it. Real training adds approximation, optimisation noise, and inductive bias, and when trained networks are fit against the exact drift null they deviate in \emph{opposite, architecture-specific} directions: a smoothing recurrent network resists collapse (keeping spurious variants alive), while a sharpening image generator accelerates it. A one-parameter \textbf{learning kernel} (a smoothing knob and a sharpening knob on the refit) reproduces both. The honest statement, used throughout: a real learner is Wright--Fisher \emph{plus a signed, measurable estimator bias} --- and the drift signs (rare-first loss; the grounding response) survived that bias in every architecture we tested, including a convolutional VAE retrained on its own generated digits, where the dry lineage collapses to a single blurred digit class while 10\% grounding holds all thirty modes (Fig. 1).
|
||||
|
||||
\textbf{Table 1.} The dictionary. Each correspondence is stated with the level of support it currently has (exact = closed form in the minimal model; empirical = measured in trained systems; hypothesis = stated with a falsifier, untested or unconfirmed). The full claim-by-claim ledger with assumptions and known limits is SI Appendix, Table S1.
|
||||
|
||||
\medskip\noindent\begin{center}\footnotesize
|
||||
\begin{tabular}{p{0.307\textwidth} p{0.307\textwidth} p{0.307\textwidth}}
|
||||
\hline
|
||||
Population genetics & Model populations & Support \\ \hline
|
||||
Genetic drift in a finite population & Training on finite samples of model output & Exact (minimal model); signs in trained nets; diagnosis conceded to prior work \\[3pt]
|
||||
Immigration from a fixed source & Grounding with verified real data & Exact equilibrium; signs in RNN/MLP/VAE/MNIST \\[3pt]
|
||||
Muller's ratchet (asexual decay) & Irreversible arm of model collapse & Correspondence, scoped: applies to unrecoverable loss \\[3pt]
|
||||
Recombination / sexual reproduction & Model merging & Empirical at 0.5B--7B \\[3pt]
|
||||
Fisher--Muller effect & Merged specialists exceed every parent & Exact-model result; replicated in LLMs \\[3pt]
|
||||
Outbreeding depression under epistasis & Merging entangled skills harms offspring & Exact-model (NK landscapes); hypothesis at LLM scale \\[3pt]
|
||||
Mating systems / population structure & Who merges with whom (breadth of the parent pool) & Exact-model result; hypothesis for real populations \\[3pt]
|
||||
Reproductive isolation (BDM incompatibilities) & Merge failure from functional conflict & Empirical (MLP + LLM tiers, conflict-associated); emergent form not observed \\[3pt]
|
||||
Selection on a fitness function & Verifier-anchored selection (``reality that can say no'') & Exact-model result (jointly necessary with sex and diversity) \\[3pt]
|
||||
\hline\end{tabular}\end{center}\medskip
|
||||
|
||||
\section*{Results}
|
||||
|
||||
\subsection*{Grounding is immigration: cheap, with a floor}
|
||||
|
||||
In the minimal model, grounding from a fixed real source is immigration into a drifting population, and the equilibrium diversity has a closed form our simulator matches exactly. The engineering headline is the \emph{magnitude}: a critical grounding fraction \texttt{g* \(\approx\) 0.05} retains most diversity indefinitely --- real data is cheap insurance. But the same analysis yields a floor the field's average-loss framing misses: an individual capability of rarity \texttt{p} survives only if the \emph{absolute} real-data budget satisfies \texttt{m\(\cdot\)p \(\gtrsim\) 1}. Protecting the rarest knowledge is priced per item, at cost \texttt{\(\propto\) 1/p}, and no affordable grounding fraction rescues the deepest tail --- that requires recombination (next section). In trained networks the \emph{sign} of the grounding response transfers everywhere we looked, with two honest deviations, both traced to the estimator bias above: sharp thresholds soften, and support-counting metrics decouple from truth (forward-KL is the operative collapse metric for a smoothing learner). On real images (Fig. 1B), dry self-training collapses a convolutional VAE to one mode while \textasciitilde{}10\% grounding holds all thirty (the trained model needs roughly twice the exact-operator fraction --- the measured price of the estimator bias).
|
||||
|
||||
\begin{figure*}[p]\centering % fig1
|
||||
\includegraphics[width=\textwidth,height=0.63\textheight,keepaspectratio]{figs/fig1_E2.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.63\textheight,keepaspectratio]{figs/fig1_mnist_montage.pdf}\par\smallskip
|
||||
\caption{Collapse is drift; grounding is immigration. (A, top) The grounding phase response in the minimal model: a critical real-data fraction $g^*\!\approx\!0.05$ retains most diversity indefinitely, while tail survival obeys the per-item floor $m\,p \gtrsim 1$. (B, bottom) The same signs on real images: a convolutional VAE retrained each generation on its own output collapses to a single blurred mode (rows: generations), while $\sim$10\% grounding holds all thirty class$\times$style modes.}\label{fig1}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{Recombination: a conservation law, its operators, and offspring that exceed every parent}
|
||||
|
||||
Merging is where the evolution-of-sex apparatus pays for itself, beginning with a result about the obvious operator. \textbf{Averaging is blending inheritance, and it cancels the benefit of multiple parents:} when a child is refit to the \emph{mean of its parents' output distributions}, the expected mass on any rare item is conserved at the single-parent level --- in the rare-item regime the 1/K dilution of averaging exactly cancels the union gain of K parents, so adding parents cannot help. An operator that keeps, per item, its strongest source (which presupposes a verifier or oracle to say which) realises the union. That statement is exact for those operators in the minimal model. The practically important operators --- \textbf{weight averaging} (a nonlinear network's weight-mean does not compute its parents' output-mean) and \textbf{routing among intact specialists} (different storage and inference budgets from a single child) --- are its empirical cousins, and the measured bridge is a \textbf{headroom rule}: in language models, union-preserving operators beat the weight-average in proportion to how far that average is from the best attainable. On easy tasks a capable base's average is already at ceiling and refinements add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent and routing wins by a wide margin (Fig. 6A--B).
|
||||
|
||||
The generative payoff is the \textbf{Fisher--Muller effect}: recombination assembles, in one offspring, complementary variants that arose in different lineages, producing a genotype fitter than any parent. In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum --- a genotype no parent held --- while the best single parent and the blended average both plateau below (Fig. 2). In real language models the signature replicates under seed replication: merges of three LoRA specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 6A).
|
||||
|
||||
Sex has risks and, for AI, an unfair advantage --- both quantified on rugged (epistatic) NK landscapes (Fig. 3). When skills are entangled, blind recombination produces offspring \emph{below} their parents --- \textbf{outbreeding depression} --- worsening with ruggedness, and the optimal recombination rate shrinks as entanglement grows. But an engineered population can do what biology cannot: recombine unbounded parents, choose complementary mates, and \emph{screen many candidate offspring against a verifier before keeping one}. This \textbf{directed sex} converts the outbreeding catastrophe into a reliable gain in the model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks --- including one seed where the blend failed catastrophically and selection was immune (Fig. 6A). Finally, population \emph{structure} is itself a knob: sweeping the mate-pool breadth from monogamous (local) to promiscuous (panmictic) against ruggedness, wide mixing maximises the population mean while monotonically destroying diversity, and the best \emph{champion} shifts from wide breadth on smooth landscapes to intermediate breadth on rugged ones (Fig. 3C) --- the mating-system phenomenon known to structured-population search, mapped onto merging populations.
|
||||
|
||||
\begin{figure*}[p]\centering % fig2
|
||||
\includegraphics[width=\textwidth,height=0.63\textheight,keepaspectratio]{figs/fig2_E4.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.63\textheight,keepaspectratio]{figs/fig2_E8.pdf}\par\smallskip
|
||||
\caption{Recombination: the conservation law and the Fisher--Muller effect. (A, top) Refitting a child to the mean of its parents' output distributions conserves rare-item mass at single-parent level regardless of parent count (blending inheritance); a strongest-source (union) operator realises the multi-parent gain. (B, bottom) Multi-locus recombination of decorrelated specialists assembles a genotype fitter than any parent, climbing to the optimum as parents are added, while the best single parent and the blended average plateau below.}\label{fig2}
|
||||
\end{figure*}
|
||||
|
||||
\begin{figure*}[p]\centering % fig3
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig3_E9.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig3_E10.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig3_E14.pdf}\par\smallskip
|
||||
\caption{Rugged (epistatic) landscapes: risk, remedy, and structure. (A, top) Outbreeding depression: blind recombination of specialists drops offspring below their parents, worsening with ruggedness; the optimal recombination rate shrinks as skills entangle. (B, middle) Directed sex --- unbounded parents, chosen mates, verifier-screened offspring --- converts the catastrophe into a reliable gain at every ruggedness. (C, bottom) Mating structure: wide (promiscuous) mixing maximises the population mean but monotonically destroys diversity; the champion-optimal mate-pool breadth narrows as the landscape roughens.}\label{fig3}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{The society: grounding, sex, and diversity are jointly necessary}
|
||||
|
||||
Composing the operators closes the loop (Fig. 4). A finite population of agents evolves on a rugged NK landscape, with selection acting on a grounded score --- \texttt{g}\(\cdot\)true-fitness + (1−g)\(\cdot\)conformity to the population's own consensus, the analogue of training on the crowd's output. A four-arm ablation separates the failure modes: the \textbf{full} system (grounding + directed recombination + diversity-preserving selection) climbs to near the global optimum while keeping its specialists; remove \emph{grounding} and the population converges confidently on an unfit consensus (self-consumption); remove \emph{sex} and it strands on local optima; remove \emph{diversity} and it converges prematurely to a worse answer. Each removal fails \emph{differently} --- the operators are jointly necessary, which is the system-level claim the single-operator results build toward. At language-model scale this composed loop remains unbuilt; it is the paper's largest stated gap.
|
||||
|
||||
\begin{figure*}[p]\centering % fig4
|
||||
\includegraphics[width=\textwidth,height=0.98\textheight,keepaspectratio]{figs/fig4_E11.pdf}\par\smallskip
|
||||
\caption{The society: grounding, sex, and diversity are jointly necessary. A finite agent population on a rugged NK landscape under a grounded selection score. Four-arm ablation: the full system climbs to near the global optimum; removing grounding collapses the population onto a confident, unfit consensus (self-consumption); removing recombination strands it on local optima; removing diversity converges it prematurely. Each ablation fails differently.}\label{fig4}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{The limit of sex: model speciation}
|
||||
|
||||
Recombination presupposes compatible parents. In biology, lineages pushed far enough apart become separate species --- \textbf{reproductive isolation} --- through Bateson--Dobzhansky--Muller incompatibilities: changes harmless on their own background but deleterious in combination. A merged model is exactly the exposed hybrid. We built the analytic model (Fig. 5A): hybrid fitness tracks the parents while compatible, then peels off and crashes below the ancestor; the isolation cliff arrives earlier the denser the incompatibilities; and the incompatibility \emph{count} snowballs quadratically with divergence (17) --- noting that a super-linear count does not by itself entail a sharp performance cliff without the count-to-effect-size link, which the analytic model supplies under its assumptions and any neural test must establish separately.
|
||||
|
||||
In trained networks, the claim must survive a known alternative: merge barriers between independently trained networks are famously \emph{coordinate artefacts}, removable by re-aligning hidden units (23), and richer symmetry groups remove more (24). We therefore aligned modulo the \textbf{complete} function-preserving unit symmetry group of the architecture tested (permutation composed with per-unit positive rescaling, for plain ReLU MLPs) and decomposed the barrier (Fig. 5B): two networks trained from different initialisations on the \emph{same} task have a barrier that alignment removes essentially entirely (residual \(\approx\) 0.001, the aligned merge performing at parent level) --- coordinate, not functional; two networks trained on \emph{conflicting} label maps have a barrier the full group leaves intact (0.502 \(\rightarrow\) 0.497), with the merged model functionally dead --- and this cannot be an alignment failure, because the same aligner succeeded on the control. Sweeping conflict traces the cliff as hybrid fitness, 0.97 \(\rightarrow\) 0.03. Two scope notes: exact recovery of a permuted-and-rescaled copy validates a special case rather than global optimality, so the removable share is a lower bound and the residual an upper bound; and the conflict floor itself is information-theoretic --- no single model can satisfy contradictory conventions (SI Appendix, Proposition S2) --- with the framework's role being the \emph{structure around it}: which divergences generate conflict, and what moves the cliff.
|
||||
|
||||
The sharpest honesty comes from the pre-registered \textbf{emergent test}: true BDM incompatibilities are emergent (each lineage's changes harmless alone), so we let children diverge with \emph{no conflicting signal anywhere} --- complementary class specialists, and divergent input conventions --- to 6.4\(\times\) the base training. \textbf{No isolation emerged} (residual 0.000 throughout); instead the merge \emph{rescued} the two catastrophically-forgetting specialists (parents \(\approx\) 0.50, merge \(\approx\) 0.955 --- a sustained Fisher--Muller rescue). The same double result appears at the language-model tier (Fig. 5C): conflicting conventions produce \textbf{function-specific} hybrid breakdown (the merge scores below both parents on the conflicted function, while a budget-controlled design shows the disjoint skills merge unharmed), and over-training disjoint specialists 1\(\rightarrow\)12 epochs produces no isolation at all --- the merge improves. Across every tier tested, \textbf{isolation had to be provoked by functional conflict; specialisation alone did not speciate} --- a bound on the analogy that sharpens the design rule: what breaks merging is conflicting conventions on shared circuitry, not divergence per se.
|
||||
|
||||
\begin{figure*}[p]\centering % fig5
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig5_E12.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig5_speciation_real.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig5_llm_speciation.pdf}\par\smallskip
|
||||
\caption{Model speciation across three tiers. (A, top) Analytic: hybrid fitness traces compatible $\rightarrow$ outbreeding depression $\rightarrow$ inviability; the cliff arrives earlier the denser the incompatibilities; incompatibility count snowballs with divergence. (B, middle) Trained MLPs: the merge barrier decomposed under the complete unit symmetry group --- same-task/different-init barriers are coordinate artefacts (removed by alignment); conflicting-task barriers survive in full, with hybrid fitness falling 0.97 $\rightarrow$ 0.03; divergence without conflict produced no isolation, the merge instead rescuing the forgetting specialists. (C, bottom) Language models: conflicting conventions produce function-specific hybrid breakdown; over-training disjoint specialists produces none --- at every tier tested, isolation had to be provoked by functional conflict.}\label{fig5}
|
||||
\end{figure*}
|
||||
|
||||
\subsection*{A controlled predictive test: functional conflict, measured pre-merge, predicts merge damage}
|
||||
|
||||
The framework's prediction-level claim was put to a designed test (Fig. 6C). Thirty-nine parent pairs (13 conditions \(\times\) 3 seeds; rows are not independent --- parents share task-data seeds across conditions --- so all inference is condition-clustered) span three axes decorrelated by construction: \emph{conflict} (contradictory conventions on shared prompts, private budgets fixed), \emph{compatible overlap} (the same shared prompts under the same convention --- overlap and volume without conflict), and \emph{duration} (weight divergence with zero conflict). Before merging, six predictors are computed: \textbf{confidence-weighted functional conflict} (bilateral confident disagreement on probes drawn blind to where conflict lives --- a proposed proxy for merge-relevant interactions, motivated by the observation that raw disagreement counts harmless complementation, one parent merely ignorant, as conflict), raw disagreement, gradient alignment at the shared base (21), LoRA-delta cosine and distance, and a cross-task performance baseline. The pre-registered outcome is the merge penalty against oracle parent potential (the hybrid-load analogue), also reported against best- and mean-parent references because the predictor ordering is sensitive to that choice.
|
||||
|
||||
The supported conclusion, stated conditionally: \textbf{across this controlled grid, pre-merge functional disagreement predicted merge penalties (clustered bootstrap CIs excluding zero; held-out leave-one-condition-out \(\rho\) \(\approx\) 0.35--0.40), whereas LoRA-delta cosine and L2 showed no statistically detectable association; gradient alignment carried intermediate signal.} Head-to-head predictor differences are not individually significant at this sample size, and only these baselines were tested. Two further results earn their place by tempering: the initial two-axis grid's best predictor was delta-cosine (\(\rho\) = +0.60) --- an overlap artefact that the compatible-overlap control was added to expose, and did (collapse to +0.03); and the pre-registered internal prediction that confidence weighting would beat raw disagreement \textbf{failed} (they are statistically indistinguishable as rank predictors), so the present evidence favours functional disagreement generally, not the DMI-specific refinement. The framework motivated the measurement and the controls; their success does not validate the specifically population-genetic mechanism. Whether the prediction improves a budget-matched operator choice, and whether it generalises to unfamiliar conflict structures and real task pairs, are the experiment's open front.
|
||||
|
||||
\begin{figure*}[p]\centering % fig6
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig6_llm_seeds.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig6_llm_moe.pdf}\par\smallskip
|
||||
\includegraphics[width=\textwidth,height=0.42\textheight,keepaspectratio]{figs/fig6_llm_epistasis.pdf}\par\smallskip
|
||||
\caption{The language-model tier. (A, top) Seed-replicated recombination claims (fixed test sets, training seed varied, 95\% CI): merges beat every specialist; union-preserving routing and directed offspring selection beat the blend in every seed on headroom tasks, including one catastrophic blend failure they avoided. (B, middle) The headroom rule at 7B on hard (unsaturated) tasks: the weight-average dilutes a fragile specialist below the best single parent; routing preserves it. (C, bottom) The controlled predictive test: across a task grid with conflict, compatible-overlap, and duration axes decorrelated by construction, pre-merge functional disagreement predicts merge penalty (held-out $\rho \approx 0.4$) while weight-geometry baselines show no detectable association; paired predictor differences are not individually significant.}\label{fig6}
|
||||
\end{figure*}
|
||||
|
||||
\section*{Discussion}
|
||||
|
||||
\textbf{Design rules.} Read as engineering, the results compress into rules an operator of a model population can apply. \emph{Ground every generation} in verified reality --- a few percent retains most diversity --- but price the rarest capabilities individually (\texttt{m\(\cdot\)p \(\gtrsim\) 1}) and use recombination, not grounding, to reach the deep tail. \emph{Merge, don't blend, when there is headroom}: keep specialists intact and route, or breed-and-screen candidate merges, whenever the naive average is far from ceiling; plain averaging is adequate only where a strong base has already composed the skills. \emph{Match the operator to entanglement}: merge freely when skills are additive; sparingly, with offspring selection, when they entangle; and expect the champion-optimal mating breadth to narrow as landscapes roughen. \emph{Preserve diversity as a first-class objective}, because selection can only preserve variety that exists, and the society result shows grounding, recombination, and diversity are jointly necessary. \emph{Before merging, measure functional conflict} --- cheap, pre-merge, and in our controlled setting predictive where weight distance was not; and expect specialisation alone to be merge-safe, with conflicting conventions on shared circuitry as the thing to detect and avoid.
|
||||
|
||||
\textbf{What is borrowed and what is ours.} The diagnosis --- collapse as drift --- is prior art (6--9), as are the empirical facts that merges can beat parents, that decorrelated parents merge better, and that naive averaging loses to interference-aware or routed merges (1, 19, 20), that model populations can climb (2--5), and that merge success admits ML-native predictors (21, 22). Ours is the framework-level synthesis --- inheritance, diversity, and compatibility as managed quantities --- together with: the conservation law for blending inheritance and its operator boundaries; the per-item grounding floor; the jointly-necessary society; model speciation as a named, tested question, with the coordinate-versus-functional decomposition under a complete symmetry group and the emergent null that bounds it; and the controlled predictive test with its controls. We claim the framework generated these measurements and experiments; we do not claim their outcomes validate a uniquely population-genetic mechanism, and one refinement it proposed was not supported.
|
||||
|
||||
\textbf{Limits and open problems.} The demonstrations are deliberately small: exact where small is a virtue, sign-level and seed-replicated at the language-model tier, on constructed task families with a trivially separable router and one model lineage (Qwen, 0.5B--7B). The composed society has not been built at language-model scale. The predictive test's next bars, in order of value: generalisation to \emph{unfamiliar} conflict structures and real task pairs; a demonstrably better \emph{budget-matched} merging decision; then scale replication. Beyond engineering, the framework's hardest open problem is the fitness function itself: selection optimises what is measured, and for knowledge systems the persuasive and the true compete --- grounding against a reality that can refuse is the only anchor we trust, and institutionalising that anchor (verification, replication, and challenge among models) is the society-level problem we pose but do not solve. What biology receives in return is a new model system: populations of learners where every genotype, environment, and mating decision is observable and manipulable --- where the evolution of sex can be studied with interventions (unbounded parents, offspring preview, directed mating) that no living system permits.
|
||||
|
||||
\section*{Materials and Methods}
|
||||
|
||||
\textbf{Analytic tier.} Pure NumPy/SciPy Wright--Fisher simulator over \texttt{K}-item distributions (knowledge as \texttt{p\_t}; Zipf-tailed truth \texttt{p*}; drift--grounding--refit generations), extended with a learning kernel (smoothing/sharpening refit), multi-locus genotypes on additive and Kauffman NK landscapes, n-parent crossover, and finite-population society loops. All parameters live in per-experiment YAML configs; every run derives all randomness from one master seed (\texttt{SeedSequence.spawn}) and is bitwise reproducible; scientific-validation tests assert the closed forms (heterozygosity decay, immigration equilibrium, closed-form union) to <0.5\% and run in CI with 151 further correctness tests.
|
||||
|
||||
\textbf{Neural tier.} Trained-network experiments realise the same abstractions with an exact oracle: histogram/RNN/MLP/VAE generators on a synthetic mode universe (the histogram model reduces the harness exactly to the analytic tier --- the bridge gate), and a convolutional VAE on MNIST with a frozen CNN oracle (98.5\% mode accuracy; confusion matrix recorded as the measurement floor). Speciation experiments fork no-BatchNorm MLPs (784--512--512--10) from a shared base, weight-average, and measure linear-mode-connectivity error barriers before and after alignment; alignment composes deterministic Git Re-Basin permutation matching with exact per-unit scale canonicalisation (the complete unit symmetry group for this class), gated by exact recovery of a permuted-and-rescaled copy.
|
||||
|
||||
\textbf{Language-model tier.} LoRA specialists (rank 16) on procedurally generated task families with an exact-match verifier, on frozen Qwen2.5-Instruct bases (0.5B on one 16 GB GPU; 7B on one L40S). Operators: weight-space merges (soup/TIES via adapter arithmetic), per-input routing, and Dirichlet-sampled offspring populations screened on held-out validation splits. Multi-seed protocols fix the test sets and vary the training seed. The predictive test computes all predictors pre-merge (generation confidence from token log-probabilities; base-model gradient cosines; exact r-space LoRA-delta geometry) and evaluates merges on held-out tests; robust statistics (condition-clustered bootstrap, paired predictor contrasts, leave-one-condition-out prediction, multi-reference outcomes) are produced by a committed script. Statistical, per-seed reproducibility is documented for GPU tiers.
|
||||
|
||||
\textbf{Data and code availability.} All code, configs, seeds, results artifacts (with content hashes), figures, and a one-command reproduction script will be deposited openly (repository + archived DOI) on publication; every figure in this paper regenerates from committed artifacts without re-simulation.
|
||||
|
||||
\section*{References}
|
||||
|
||||
\begin{enumerate}
|
||||
\item Yadav P, Tam D, Choshen L, Raffel C, Bansal M (2023) TIES-Merging: resolving interference when merging models. \emph{NeurIPS}. arXiv:2306.01708.
|
||||
\item Akiba T, Shing M, Tang Y, Sun Q, Ha D (2025) Evolutionary optimization of model merging recipes. \emph{Nat Mach Intell} 7:195--204.
|
||||
\item GENOME: Nature-inspired population-based evolution of large language models (2025). arXiv:2503.01155.
|
||||
\item Sakana AI (2025) Competition and attraction improve model fusion (M2N2). \emph{GECCO}. arXiv:2508.16204.
|
||||
\item Subramaniam V, Du Y, Tenenbaum JB, Torralba A, Li S, Mordatch I (2025) Multiagent finetuning: self-improvement with diverse reasoning chains. arXiv:2501.05707.
|
||||
\item Shumailov I, et al. (2024) AI models collapse when trained on recursively generated data. \emph{Nature} 631:755--759.
|
||||
\item Riis S (2026) Drift and selection in LLM text ecosystems. arXiv:2604.08554.
|
||||
\item Benati M, Londei A, Lanzieri D, Loreto V (2025) First-extinction law for resampling processes. arXiv:2509.20101.
|
||||
\item Yoon Y, Hu D, Weissburg I, Qin Y, Jeong H (2025) Model collapse in the self-consuming chain of diffusion finetuning: a quantitative trait modeling perspective. \emph{ICLR}. arXiv:2407.17493.
|
||||
\item Muller HJ (1964) The relation of recombination to mutational advance. \emph{Mutat Res} 1:2--9.
|
||||
\item Gerstgrasser M, et al. (2024) Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. arXiv:2404.01413.
|
||||
\item Yi B, Liu Q, Cheng Y, Xu H (2025) Escaping model collapse via synthetic data verification. arXiv:2510.16657.
|
||||
\item Wright S (1931) Evolution in Mendelian populations. \emph{Genetics} 16:97--159.
|
||||
\item Fisher RA (1930) \emph{The Genetical Theory of Natural Selection} (Clarendon, Oxford).
|
||||
\item Muller HJ (1932) Some genetic aspects of sex. \emph{Am Nat} 66:118--138.
|
||||
\item Orr HA (1995) The population genetics of speciation: the evolution of hybrid incompatibilities. \emph{Genetics} 139:1805--1813.
|
||||
\item Orr HA, Turelli M (2001) The evolution of postzygotic isolation: accumulating Dobzhansky--Muller incompatibilities. \emph{Evolution} 55:1085--1094.
|
||||
\item Livnat A, Papadimitriou C (2016) Sex as an algorithm: the theory of evolution under the lens of computation. \emph{Commun ACM} 59(11):84--93.
|
||||
\item Yu L, Yu B, Yu H, Huang F, Li Y (2023) Language models are super Mario: absorbing abilities from homologous models (DARE). arXiv:2311.03099.
|
||||
\item Wortsman M, et al. (2022) Model soups: averaging weights of multiple fine-tuned models. \emph{ICML}. arXiv:2203.05482.
|
||||
\item Zhou L, Zhao B, Yu R, Rodolà E (2026) Demystifying mergeability: interpretable properties to predict model merging success. arXiv:2601.22285.
|
||||
\item Cao Y, Ran D, Guo Y, Wu M, Chen S, et al. (2026) An empirical study and theoretical explanation on task-level model-merging collapse. arXiv:2603.09463.
|
||||
\item Ainsworth S, Hayase J, Srinivasa S (2022) Git Re-Basin: merging models modulo permutation symmetries. arXiv:2209.04836.
|
||||
\item Li T, Shen Z (2026) Scaling linear mode connectivity and merging to billion-parameter pretrained transformers. arXiv:2606.23607.
|
||||
\item Kauffman SA, Levin S (1987) Towards a general theory of adaptive walks on rugged landscapes. \emph{J Theor Biol} 128:11--45.
|
||||
\item Lehman J, Stanley KO (2011) Abandoning objectives: evolution through the search for novelty alone. \emph{Evol Comput} 19:189--223.
|
||||
\item Pari J, Jelassi S, Agrawal P (2024) Collective model intelligence requires compatible specialization. arXiv:2411.02207.
|
||||
\item Hu EJ, et al. (2021) LoRA: low-rank adaptation of large language models. arXiv:2106.09685.
|
||||
\item Sharma E, Roy DM, Dziugaite GK (2024) The non-local model merging problem: permutation symmetries and variance collapse. arXiv:2410.12766.
|
||||
\item Kozodoi N, Afolabi Z, Butler J (2026) Are we merging the right models? Impact of expert training duration on model merging for LLMs. arXiv:2607.11997.
|
||||
\end{enumerate}
|
||||
|
||||
192
paper/pnas/build.py
Normal file
192
paper/pnas/build.py
Normal file
|
|
@ -0,0 +1,192 @@
|
|||
r"""Build the PNAS-draft PDF from main.md (Markdown stays the source of truth).
|
||||
|
||||
Adapted from paper/arxiv/md2tex.py (same Markdown subset + pipe tables), with one addition: standalone
|
||||
`*(FIG:name)*` markers compose multi-panel figures by stacking existing per-experiment vector PDFs
|
||||
(LaTeX-level consolidation; bespoke unified figures are a submission-time polish, tracked in the work
|
||||
order). Captions define the panel letters positionally (A = top, ...) because the sub-figures carry
|
||||
their own internal panel labels.
|
||||
|
||||
Usage: python paper/pnas/build.py && (cd paper/pnas && tectonic main.tex)
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
HERE = Path(__file__).resolve().parent
|
||||
SRC = HERE / "main.md"
|
||||
OUT = HERE / "body.tex"
|
||||
|
||||
# figure name -> (list of source PDFs (stacked top->bottom), caption)
|
||||
FIGURES: dict[str, tuple[list[str], str]] = {
|
||||
"fig1": (["results/E2/E2.pdf", "results/mnist_collapse/mnist_montage.pdf"],
|
||||
"Collapse is drift; grounding is immigration. (A, top) The grounding phase response in the "
|
||||
"minimal model: a critical real-data fraction $g^*\\!\\approx\\!0.05$ retains most diversity "
|
||||
"indefinitely, while tail survival obeys the per-item floor $m\\,p \\gtrsim 1$. (B, bottom) "
|
||||
"The same signs on real images: a convolutional VAE retrained each generation on its own "
|
||||
"output collapses to a single blurred mode (rows: generations), while $\\sim$10\\% grounding "
|
||||
"holds all thirty class$\\times$style modes."),
|
||||
"fig2": (["results/E4/E4.pdf", "results/E8/E8.pdf"],
|
||||
"Recombination: the conservation law and the Fisher--Muller effect. (A, top) Refitting a child "
|
||||
"to the mean of its parents' output distributions conserves rare-item mass at single-parent "
|
||||
"level regardless of parent count (blending inheritance); a strongest-source (union) operator "
|
||||
"realises the multi-parent gain. (B, bottom) Multi-locus recombination of decorrelated "
|
||||
"specialists assembles a genotype fitter than any parent, climbing to the optimum as parents "
|
||||
"are added, while the best single parent and the blended average plateau below."),
|
||||
"fig3": (["results/E9/E9.pdf", "results/E10/E10.pdf", "results/E14/E14.pdf"],
|
||||
"Rugged (epistatic) landscapes: risk, remedy, and structure. (A, top) Outbreeding depression: "
|
||||
"blind recombination of specialists drops offspring below their parents, worsening with "
|
||||
"ruggedness; the optimal recombination rate shrinks as skills entangle. (B, middle) Directed "
|
||||
"sex --- unbounded parents, chosen mates, verifier-screened offspring --- converts the "
|
||||
"catastrophe into a reliable gain at every ruggedness. (C, bottom) Mating structure: wide "
|
||||
"(promiscuous) mixing maximises the population mean but monotonically destroys diversity; the "
|
||||
"champion-optimal mate-pool breadth narrows as the landscape roughens."),
|
||||
"fig4": (["results/E11/E11.pdf"],
|
||||
"The society: grounding, sex, and diversity are jointly necessary. A finite agent population "
|
||||
"on a rugged NK landscape under a grounded selection score. Four-arm ablation: the full system "
|
||||
"climbs to near the global optimum; removing grounding collapses the population onto a "
|
||||
"confident, unfit consensus (self-consumption); removing recombination strands it on local "
|
||||
"optima; removing diversity converges it prematurely. Each ablation fails differently."),
|
||||
"fig5": (["results/E12/E12.pdf", "results/speciation_real/speciation_real.pdf",
|
||||
"results/llm_speciation/llm_speciation.pdf"],
|
||||
"Model speciation across three tiers. (A, top) Analytic: hybrid fitness traces compatible "
|
||||
"$\\rightarrow$ outbreeding depression $\\rightarrow$ inviability; the cliff arrives earlier "
|
||||
"the denser the incompatibilities; incompatibility count snowballs with divergence. (B, "
|
||||
"middle) Trained MLPs: the merge barrier decomposed under the complete unit symmetry group --- "
|
||||
"same-task/different-init barriers are coordinate artefacts (removed by alignment); "
|
||||
"conflicting-task barriers survive in full, with hybrid fitness falling 0.97 $\\rightarrow$ "
|
||||
"0.03; divergence without conflict produced no isolation, the merge instead rescuing the "
|
||||
"forgetting specialists. (C, bottom) Language models: conflicting conventions produce "
|
||||
"function-specific hybrid breakdown; over-training disjoint specialists produces none --- at "
|
||||
"every tier tested, isolation had to be provoked by functional conflict."),
|
||||
"fig6": (["results/llm_merge_seeds/llm_seeds.pdf", "results/llm_moe_hard_hpc/llm_moe.pdf",
|
||||
"results/llm_epistasis/llm_epistasis.pdf"],
|
||||
"The language-model tier. (A, top) Seed-replicated recombination claims (fixed test sets, "
|
||||
"training seed varied, 95\\% CI): merges beat every specialist; union-preserving routing and "
|
||||
"directed offspring selection beat the blend in every seed on headroom tasks, including one "
|
||||
"catastrophic blend failure they avoided. (B, middle) The headroom rule at 7B on hard "
|
||||
"(unsaturated) tasks: the weight-average dilutes a fragile specialist below the best single "
|
||||
"parent; routing preserves it. (C, bottom) The controlled predictive test: across a task grid "
|
||||
"with conflict, compatible-overlap, and duration axes decorrelated by construction, pre-merge "
|
||||
"functional disagreement predicts merge penalty (held-out $\\rho \\approx 0.4$) while "
|
||||
"weight-geometry baselines show no detectable association; paired predictor differences are "
|
||||
"not individually significant."),
|
||||
}
|
||||
|
||||
UNICODE = {"—": "---", "–": "--", "→": r"\(\rightarrow\)", "≈": r"\(\approx\)", "≥": r"\(\geq\)",
|
||||
"≳": r"\(\gtrsim\)", "×": r"\(\times\)", "·": r"\(\cdot\)", "μ": r"\(\mu\)",
|
||||
"ρ": r"\(\rho\)", "≤": r"\(\leq\)", "≪": r"\(\ll\)", "∝": r"\(\propto\)"}
|
||||
SPECIALS = {"&": r"\&", "%": r"\%", "#": r"\#", "_": r"\_", "$": r"\$",
|
||||
"~": r"\textasciitilde{}", "^": r"\textasciicircum{}"}
|
||||
|
||||
|
||||
def esc(s: str) -> str:
|
||||
s = s.replace("\\", r"\textbackslash{}")
|
||||
for k, v in SPECIALS.items():
|
||||
s = s.replace(k, v)
|
||||
for k, v in UNICODE.items():
|
||||
s = s.replace(k, v)
|
||||
return s
|
||||
|
||||
|
||||
def inline(s: str) -> str:
|
||||
parts = re.split(r"(`[^`]*`)", s)
|
||||
out = []
|
||||
for p in parts:
|
||||
if p.startswith("`") and p.endswith("`") and len(p) >= 2:
|
||||
out.append(r"\texttt{" + esc(p[1:-1]) + "}")
|
||||
else:
|
||||
p = esc(p)
|
||||
p = re.sub(r"\[([^\]]+)\]\((https?://[^)]+)\)", r"\\href{\2}{\1}", p)
|
||||
p = re.sub(r"\*\*([^*]+)\*\*", r"\\textbf{\1}", p)
|
||||
p = re.sub(r"\*([^*]+)\*", r"\\emph{\1}", p)
|
||||
p = re.sub(r'"([^"]+)"', r"``\1''", p)
|
||||
out.append(p)
|
||||
return "".join(out)
|
||||
|
||||
|
||||
def figure_env(name: str) -> str:
|
||||
pdfs, caption = FIGURES[name]
|
||||
(HERE / "figs").mkdir(exist_ok=True)
|
||||
lines = [f"\\begin{{figure*}}[p]\\centering % {name}"]
|
||||
for src in pdfs:
|
||||
dst = HERE / "figs" / (name + "_" + Path(src).name)
|
||||
shutil.copyfile(ROOT / src, dst)
|
||||
frac = min(0.98, 3.0 / len(pdfs) * 0.42)
|
||||
lines.append(f"\\includegraphics[width=\\textwidth,height={frac:.2f}\\textheight,"
|
||||
f"keepaspectratio]{{figs/{dst.name}}}\\par\\smallskip")
|
||||
lines.append(f"\\caption{{{caption}}}\\label{{{name}}}")
|
||||
lines.append("\\end{figure*}")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def convert(text: str) -> str:
|
||||
lines = text.split("\n")
|
||||
i = 0
|
||||
while i < len(lines) and lines[i].strip() != "---":
|
||||
i += 1
|
||||
i += 1
|
||||
|
||||
blocks: list[list[str]] = []
|
||||
cur: list[str] = []
|
||||
for line in lines[i:]:
|
||||
if line.strip() == "":
|
||||
if cur:
|
||||
blocks.append(cur); cur = []
|
||||
else:
|
||||
cur.append(line)
|
||||
if cur:
|
||||
blocks.append(cur)
|
||||
|
||||
def emit_table(block, out):
|
||||
rows = [[c.strip() for c in line.strip().strip("|").split("|")] for line in block]
|
||||
header, body = rows[0], rows[2:]
|
||||
n = len(header)
|
||||
widths = " ".join([f"p{{{0.92 / n:.3f}\\textwidth}}"] * n)
|
||||
out += ["\\medskip\\noindent\\begin{center}\\footnotesize",
|
||||
f"\\begin{{tabular}}{{{widths}}}", "\\hline",
|
||||
" & ".join(inline(c) for c in header) + " \\\\ \\hline"]
|
||||
for r in body:
|
||||
r = (r + [""] * n)[:n]
|
||||
out.append(" & ".join(inline(c) for c in r) + " \\\\[3pt]")
|
||||
out += ["\\hline\\end{tabular}\\end{center}\\medskip", ""]
|
||||
|
||||
out: list[str] = []
|
||||
for block in blocks:
|
||||
first = block[0].strip()
|
||||
m = re.match(r"^\*?\(FIG:(\w+)\)\*?$", first)
|
||||
if m:
|
||||
out.append(figure_env(m.group(1))); out.append("")
|
||||
elif first.startswith("|") and len(block) >= 2 and set(block[1].strip()) <= set("|-: "):
|
||||
emit_table(block, out)
|
||||
elif first == "---" and len(block) == 1:
|
||||
out.append("\\medskip\\hrule\\medskip"); out.append("")
|
||||
elif first.startswith("## "):
|
||||
out.append(f"\\section*{{{inline(first[3:])}}}"); out.append("")
|
||||
elif first.startswith("### "):
|
||||
out.append(f"\\subsection*{{{inline(first[4:])}}}"); out.append("")
|
||||
elif re.match(r"^(- |\d+\. )", first):
|
||||
env = "itemize" if first.startswith("- ") else "enumerate"
|
||||
out.append(f"\\begin{{{env}}}")
|
||||
items: list[str] = []
|
||||
for l in block:
|
||||
s = l.strip()
|
||||
if re.match(r"^(- |\d+\. )", s):
|
||||
items.append(re.sub(r"^(- |\d+\. )", "", s))
|
||||
else:
|
||||
items[-1] += " " + s
|
||||
for it in items:
|
||||
out.append("\\item " + inline(it.strip()))
|
||||
out.append(f"\\end{{{env}}}"); out.append("")
|
||||
else:
|
||||
joined = re.sub(r"\s{2,}", " ", " ".join(l.strip() for l in block)).strip()
|
||||
out.append(inline(joined)); out.append("")
|
||||
return "\n".join(out) + "\n"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
OUT.write_text(convert(SRC.read_text()))
|
||||
print(f"wrote {OUT}")
|
||||
BIN
paper/pnas/figs/fig1_E2.pdf
Normal file
BIN
paper/pnas/figs/fig1_E2.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig1_mnist_montage.pdf
Normal file
BIN
paper/pnas/figs/fig1_mnist_montage.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig2_E4.pdf
Normal file
BIN
paper/pnas/figs/fig2_E4.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig2_E8.pdf
Normal file
BIN
paper/pnas/figs/fig2_E8.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig3_E10.pdf
Normal file
BIN
paper/pnas/figs/fig3_E10.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig3_E14.pdf
Normal file
BIN
paper/pnas/figs/fig3_E14.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig3_E9.pdf
Normal file
BIN
paper/pnas/figs/fig3_E9.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig4_E11.pdf
Normal file
BIN
paper/pnas/figs/fig4_E11.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig5_E12.pdf
Normal file
BIN
paper/pnas/figs/fig5_E12.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig5_llm_speciation.pdf
Normal file
BIN
paper/pnas/figs/fig5_llm_speciation.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig5_speciation_real.pdf
Normal file
BIN
paper/pnas/figs/fig5_speciation_real.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig6_llm_epistasis.pdf
Normal file
BIN
paper/pnas/figs/fig6_llm_epistasis.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig6_llm_moe.pdf
Normal file
BIN
paper/pnas/figs/fig6_llm_moe.pdf
Normal file
Binary file not shown.
BIN
paper/pnas/figs/fig6_llm_seeds.pdf
Normal file
BIN
paper/pnas/figs/fig6_llm_seeds.pdf
Normal file
Binary file not shown.
396
paper/pnas/main.md
Normal file
396
paper/pnas/main.md
Normal file
|
|
@ -0,0 +1,396 @@
|
|||
# The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations
|
||||
|
||||
**Giorgio F. Gilestro** — Department of Life Sciences, Imperial College London. giorgio@gilest.ro
|
||||
|
||||
---
|
||||
|
||||
## Significance statement
|
||||
|
||||
Artificial intelligence is shifting from single, frozen models to populations of models that
|
||||
specialise, are retrained on each other's output, and are combined ("merged") into new models. Trained
|
||||
on their own output, model lineages degenerate — a process already recognised as the mathematics of
|
||||
genetic drift. This paper imports the other half of population genetics: the biology of sexual
|
||||
reproduction. It treats model merging as recombination, real data as immigration, and merge failure as
|
||||
reproductive isolation, and tests each correspondence in simulations, small neural networks, and
|
||||
language models. The framework yields design rules — when to average models, when to keep them
|
||||
separate, how much real data suffices — and a first controlled test showing that measured functional
|
||||
conflict, not weight distance, predicts when merging fails.
|
||||
|
||||
## Abstract
|
||||
|
||||
AI development increasingly resembles a population process: models are specialised, retrained on model
|
||||
output, and recombined by weight merging, with an openly evolutionary vocabulary but little use of
|
||||
evolutionary theory. Here we treat multigenerational model populations as systems whose inheritance,
|
||||
diversity, and compatibility must be managed, and transfer the quantitative apparatus of the evolution
|
||||
of sex. We take as settled that training on model output is genetic drift (model collapse). In a
|
||||
minimal inheritance model that is exactly Wright–Fisher — and measurably Wright–Fisher-plus-bias in
|
||||
trained networks — we derive and test the remedies: grounding as immigration, with a critical
|
||||
real-data fraction far below one but a per-capability floor that leaves the rarest knowledge
|
||||
unrescuable; recombination, where averaging parents' output distributions exactly cancels the benefit
|
||||
of multiple parents while union-preserving operators realise it; the Fisher–Muller effect, with merged
|
||||
language-model specialists exceeding every parent in replicated experiments; outbreeding depression on
|
||||
rugged task landscapes, converted into reliable gains by directed, offspring-screened recombination;
|
||||
and population structure, where the optimal mating breadth shrinks as skills entangle. Sex has a
|
||||
limit: we introduce model speciation — merge failure as reproductive isolation — and show in trained
|
||||
networks that a merge barrier surviving the full function-preserving symmetry group tracks functional
|
||||
conflict, that isolation did not emerge from compatible specialisation, and, in a controlled
|
||||
predictive test, that pre-merge functional disagreement predicts merge damage where weight-geometry
|
||||
baselines do not. We state precisely what is exact, what is measured, and what remains hypothesis.
|
||||
|
||||
---
|
||||
|
||||
## Introduction
|
||||
|
||||
The unit of AI progress is quietly changing. Multi-agent systems arrange many models across *space* —
|
||||
specialists cooperating on a task. A newer axis is *time*: populations of models that persist across
|
||||
generations, each new model built from older ones — specialised by fine-tuning, trained on data earlier
|
||||
models generated, and, increasingly, produced by **model merging**, the direct combination of trained
|
||||
weights (1, 2). The engineering literature describes this openly in evolutionary vocabulary —
|
||||
"crossover," "mutation," "mate choice," populations of merging models that climb benchmarks (2–5) —
|
||||
but as metaphor over search algorithms. The organising claim of this paper is that the vocabulary
|
||||
deserves its mathematics: **multigenerational model populations are systems whose inheritance,
|
||||
diversity, and compatibility must be managed — not merely collections of models to optimise — and the
|
||||
branch of biology that studies exactly this problem, the population genetics of the evolution of sex,
|
||||
transfers as a quantitative framework.**
|
||||
|
||||
One half of the transfer is settled and is not our contribution. Training each generation of a model
|
||||
on the previous generation's output degrades it — *model collapse*: rare capabilities vanish first and
|
||||
the lineage drifts toward its own most common behaviour (6). That this is the mathematics of **genetic
|
||||
drift** in a finite population is now established from several directions (7–9); a closed-form
|
||||
first-extinction law even places collapse onset at the Wright–Fisher first-extinction time (8). We cite
|
||||
this literature as the diagnosis and build on it.
|
||||
|
||||
Our contribution is on the remedy side, and we are explicit about what kind of contribution each claim
|
||||
is, distinguishing **interpretation** (an existing result understood in population-genetic terms),
|
||||
**explanation** (the transferred mechanism accounts for observations existing accounts leave open),
|
||||
and **prediction** (the framework forecasts an unmeasured outcome). The paper is strongest on the
|
||||
first; makes concrete progress on the second — separating merge failures that are coordinate artefacts
|
||||
from those that are functional; and reports a first, bounded step on the third — a controlled
|
||||
predictive test in which pre-merge functional-disagreement measures, chosen by the framework,
|
||||
predicted merge damage on a constructed task grid while weight-geometry baselines did not.
|
||||
|
||||
The correspondences we develop, summarised in Table 1: single-teacher retraining is **asexual
|
||||
reproduction**, and the irreversible arm of its decay corresponds to **Muller's ratchet** (10) — once
|
||||
every copy of a rare capability is gone from all parents and sources, no recombination can rebuild it,
|
||||
which is precisely why remedies must act before fixation-by-loss. Injecting verified real data is
|
||||
**immigration** from a non-drifting source (11–13). Model merging is **recombination**, and its
|
||||
celebrated payoff — a merged model exceeding every parent — is the **Fisher–Muller effect** (14, 15).
|
||||
Merging entangled skills courts **outbreeding depression**; screening many candidate merges is a form
|
||||
of **directed sex** with no biological analogue; restricting who merges with whom is **population
|
||||
structure**. And merging's hard limit — models too diverged in function to combine — is **reproductive
|
||||
isolation**, for which the Bateson–Dobzhansky–Muller theory of incompatibilities (16, 17) supplies the
|
||||
structure. The nearest precursor to this programme reads sex as an algorithm for mixability in the
|
||||
theory of computation (18), pre-dating model merging; the model-merging literature itself has strong
|
||||
empirical operators (1, 19, 20) and emerging merge-success predictors (21, 22), to which our delta is
|
||||
mechanism: *when and why* failure is coordinate versus functional, and what moves the boundary.
|
||||
|
||||
We support the framework at three tiers of evidence, in ascending realism and descending exactness: a
|
||||
**minimal analytic model** validated against closed forms to a fraction of a percent; **small trained
|
||||
networks** (MLPs, recurrent networks, an MNIST image generator) where the operators are measured in
|
||||
real weights; and **language models** (LoRA-specialised Qwen models, 0.5B locally and 7B on a compute
|
||||
cluster) where the claims are tested as signs under seed replication. Throughout, we report negative
|
||||
and tempering results with the same prominence as confirmations: they include the failure of an
|
||||
internal pre-registered prediction, a null on emergent speciation that bounds the analogy, and the
|
||||
sensitivity analyses that temper the predictive test.
|
||||
|
||||
## The minimal model, and where its exactness ends
|
||||
|
||||
Knowledge is modelled as a distribution `p_t` over `K` discrete items — capabilities, facts, modes of
|
||||
behaviour — with a fixed true distribution `p*` whose rare tail carries the knowledge most at risk.
|
||||
One generation is: *draw `n` samples from the parent's distribution, optionally mix in `m` verified
|
||||
real samples ("grounding", `g = m/(n+m)`), and refit the child*. In this minimal inheritance model the
|
||||
resampling step **is** the Wright–Fisher process — the same equations, which we exploit as an
|
||||
engineering gate: our simulator reproduces the classical closed forms (heterozygosity decay
|
||||
`E[H_t] = H_0(1 − 1/n)^t`; the exact immigration–drift equilibrium; the closed-form multi-teacher
|
||||
union) to within 0.5%, and these are standing tests in the codebase, not one-off checks.
|
||||
|
||||
The boundary of the exactness matters, and we measured it rather than assumed it. Real training adds
|
||||
approximation, optimisation noise, and inductive bias, and when trained networks are fit against the
|
||||
exact drift null they deviate in *opposite, architecture-specific* directions: a smoothing recurrent
|
||||
network resists collapse (keeping spurious variants alive), while a sharpening image generator
|
||||
accelerates it. A one-parameter **learning kernel** (a smoothing knob and a sharpening knob on the
|
||||
refit) reproduces both. The honest statement, used throughout: a real learner is Wright–Fisher *plus a
|
||||
signed, measurable estimator bias* — and the drift signs (rare-first loss; the grounding response)
|
||||
survived that bias in every architecture we tested, including a convolutional VAE retrained on its own
|
||||
generated digits, where the dry lineage collapses to a single blurred digit class while 10% grounding
|
||||
holds all thirty modes (Fig. 1).
|
||||
|
||||
**Table 1.** The dictionary. Each correspondence is stated with the level of support it currently has
|
||||
(exact = closed form in the minimal model; empirical = measured in trained systems; hypothesis =
|
||||
stated with a falsifier, untested or unconfirmed). The full claim-by-claim ledger with assumptions and
|
||||
known limits is SI Appendix, Table S1.
|
||||
|
||||
| Population genetics | Model populations | Support |
|
||||
|---|---|---|
|
||||
| Genetic drift in a finite population | Training on finite samples of model output | Exact (minimal model); signs in trained nets; diagnosis conceded to prior work |
|
||||
| Immigration from a fixed source | Grounding with verified real data | Exact equilibrium; signs in RNN/MLP/VAE/MNIST |
|
||||
| Muller's ratchet (asexual decay) | Irreversible arm of model collapse | Correspondence, scoped: applies to unrecoverable loss |
|
||||
| Recombination / sexual reproduction | Model merging | Empirical at 0.5B–7B |
|
||||
| Fisher–Muller effect | Merged specialists exceed every parent | Exact-model result; replicated in LLMs |
|
||||
| Outbreeding depression under epistasis | Merging entangled skills harms offspring | Exact-model (NK landscapes); hypothesis at LLM scale |
|
||||
| Mating systems / population structure | Who merges with whom (breadth of the parent pool) | Exact-model result; hypothesis for real populations |
|
||||
| Reproductive isolation (BDM incompatibilities) | Merge failure from functional conflict | Empirical (MLP + LLM tiers, conflict-associated); emergent form not observed |
|
||||
| Selection on a fitness function | Verifier-anchored selection ("reality that can say no") | Exact-model result (jointly necessary with sex and diversity) |
|
||||
|
||||
## Results
|
||||
|
||||
### Grounding is immigration: cheap, with a floor
|
||||
|
||||
In the minimal model, grounding from a fixed real source is immigration into a drifting population,
|
||||
and the equilibrium diversity has a closed form our simulator matches exactly. The engineering
|
||||
headline is the *magnitude*: a critical grounding fraction `g* ≈ 0.05` retains most diversity
|
||||
indefinitely — real data is cheap insurance. But the same analysis yields a floor the field's
|
||||
average-loss framing misses: an individual capability of rarity `p` survives only if the *absolute*
|
||||
real-data budget satisfies `m·p ≳ 1`. Protecting the rarest knowledge is priced per item, at cost
|
||||
`∝ 1/p`, and no affordable grounding fraction rescues the deepest tail — that requires recombination
|
||||
(next section). In trained networks the *sign* of the grounding response transfers everywhere we
|
||||
looked, with two honest deviations, both traced to the estimator bias above: sharp thresholds soften,
|
||||
and support-counting metrics decouple from truth (forward-KL is the operative collapse metric for a
|
||||
smoothing learner). On real images (Fig. 1B), dry self-training collapses a convolutional VAE to one
|
||||
mode while ~10% grounding holds all thirty (the trained model needs roughly twice the exact-operator
|
||||
fraction — the measured price of the estimator bias).
|
||||
|
||||
*(FIG:fig1)*
|
||||
|
||||
### Recombination: a conservation law, its operators, and offspring that exceed every parent
|
||||
|
||||
Merging is where the evolution-of-sex apparatus pays for itself, beginning with a result about the
|
||||
obvious operator. **Averaging is blending inheritance, and it cancels the benefit of multiple
|
||||
parents:** when a child is refit to the *mean of its parents' output distributions*, the expected mass
|
||||
on any rare item is conserved at the single-parent level — in the rare-item regime the 1/K dilution of
|
||||
averaging exactly cancels the union gain of K parents, so adding parents cannot help. An operator that
|
||||
keeps, per item, its strongest source (which presupposes a verifier or oracle to say which) realises
|
||||
the union. That statement is exact for those operators in the minimal model. The practically important
|
||||
operators — **weight averaging** (a nonlinear network's weight-mean does not compute its parents'
|
||||
output-mean) and **routing among intact specialists** (different storage and inference budgets from a
|
||||
single child) — are its empirical cousins, and the measured bridge is a **headroom rule**: in language
|
||||
models, union-preserving operators beat the weight-average in proportion to how far that average is
|
||||
from the best attainable. On easy tasks a capable base's average is already at ceiling and refinements
|
||||
add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent
|
||||
and routing wins by a wide margin (Fig. 6A–B).
|
||||
|
||||
The generative payoff is the **Fisher–Muller effect**: recombination assembles, in one offspring,
|
||||
complementary variants that arose in different lineages, producing a genotype fitter than any parent.
|
||||
In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum — a
|
||||
genotype no parent held — while the best single parent and the blended average both plateau below
|
||||
(Fig. 2). In real language models the signature replicates under seed replication: merges of three
|
||||
LoRA specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper
|
||||
worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 6A).
|
||||
|
||||
Sex has risks and, for AI, an unfair advantage — both quantified on rugged (epistatic) NK landscapes
|
||||
(Fig. 3). When skills are entangled, blind recombination produces offspring *below* their parents —
|
||||
**outbreeding depression** — worsening with ruggedness, and the optimal recombination rate shrinks as
|
||||
entanglement grows. But an engineered population can do what biology cannot: recombine unbounded
|
||||
parents, choose complementary mates, and *screen many candidate offspring against a verifier before
|
||||
keeping one*. This **directed sex** converts the outbreeding catastrophe into a reliable gain in the
|
||||
model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language
|
||||
models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks — including
|
||||
one seed where the blend failed catastrophically and selection was immune (Fig. 6A). Finally,
|
||||
population *structure* is itself a knob: sweeping the mate-pool breadth from monogamous (local) to
|
||||
promiscuous (panmictic) against ruggedness, wide mixing maximises the population mean while
|
||||
monotonically destroying diversity, and the best *champion* shifts from wide breadth on smooth
|
||||
landscapes to intermediate breadth on rugged ones (Fig. 3C) — the mating-system phenomenon known to
|
||||
structured-population search, mapped onto merging populations.
|
||||
|
||||
*(FIG:fig2)*
|
||||
|
||||
*(FIG:fig3)*
|
||||
|
||||
### The society: grounding, sex, and diversity are jointly necessary
|
||||
|
||||
Composing the operators closes the loop (Fig. 4). A finite population of agents evolves on a rugged NK
|
||||
landscape, with selection acting on a grounded score — `g`·true-fitness + (1−g)·conformity to the
|
||||
population's own consensus, the analogue of training on the crowd's output. A four-arm ablation
|
||||
separates the failure modes: the **full** system (grounding + directed recombination +
|
||||
diversity-preserving selection) climbs to near the global optimum while keeping its specialists;
|
||||
remove *grounding* and the population converges confidently on an unfit consensus (self-consumption);
|
||||
remove *sex* and it strands on local optima; remove *diversity* and it converges prematurely to a
|
||||
worse answer. Each removal fails *differently* — the operators are jointly necessary, which is the
|
||||
system-level claim the single-operator results build toward. At language-model scale this composed
|
||||
loop remains unbuilt; it is the paper's largest stated gap.
|
||||
|
||||
*(FIG:fig4)*
|
||||
|
||||
### The limit of sex: model speciation
|
||||
|
||||
Recombination presupposes compatible parents. In biology, lineages pushed far enough apart become
|
||||
separate species — **reproductive isolation** — through Bateson–Dobzhansky–Muller incompatibilities:
|
||||
changes harmless on their own background but deleterious in combination. A merged model is exactly the
|
||||
exposed hybrid. We built the analytic model (Fig. 5A): hybrid fitness tracks the parents while
|
||||
compatible, then peels off and crashes below the ancestor; the isolation cliff arrives earlier the
|
||||
denser the incompatibilities; and the incompatibility *count* snowballs quadratically with divergence
|
||||
(17) — noting that a super-linear count does not by itself entail a sharp performance cliff without
|
||||
the count-to-effect-size link, which the analytic model supplies under its assumptions and any neural
|
||||
test must establish separately.
|
||||
|
||||
In trained networks, the claim must survive a known alternative: merge barriers between independently
|
||||
trained networks are famously *coordinate artefacts*, removable by re-aligning hidden units (23), and
|
||||
richer symmetry groups remove more (24). We therefore aligned modulo the **complete**
|
||||
function-preserving unit symmetry group of the architecture tested (permutation composed with per-unit
|
||||
positive rescaling, for plain ReLU MLPs) and decomposed the barrier (Fig. 5B): two networks trained
|
||||
from different initialisations on the *same* task have a barrier that alignment removes essentially
|
||||
entirely (residual ≈ 0.001, the aligned merge performing at parent level) — coordinate, not
|
||||
functional; two networks trained on *conflicting* label maps have a barrier the full group leaves
|
||||
intact (0.502 → 0.497), with the merged model functionally dead — and this cannot be an alignment
|
||||
failure, because the same aligner succeeded on the control. Sweeping conflict traces the cliff as
|
||||
hybrid fitness, 0.97 → 0.03. Two scope notes: exact recovery of a permuted-and-rescaled copy validates
|
||||
a special case rather than global optimality, so the removable share is a lower bound and the residual
|
||||
an upper bound; and the conflict floor itself is information-theoretic — no single model can satisfy
|
||||
contradictory conventions (SI Appendix, Proposition S2) — with the framework's role being the
|
||||
*structure around it*: which divergences generate conflict, and what moves the cliff.
|
||||
|
||||
The sharpest honesty comes from the pre-registered **emergent test**: true BDM incompatibilities are
|
||||
emergent (each lineage's changes harmless alone), so we let children diverge with *no conflicting
|
||||
signal anywhere* — complementary class specialists, and divergent input conventions — to 6.4× the base
|
||||
training. **No isolation emerged** (residual 0.000 throughout); instead the merge *rescued* the two
|
||||
catastrophically-forgetting specialists (parents ≈ 0.50, merge ≈ 0.955 — a sustained Fisher–Muller
|
||||
rescue). The same double result appears at the language-model tier (Fig. 5C): conflicting conventions
|
||||
produce **function-specific** hybrid breakdown (the merge scores below both parents on the conflicted
|
||||
function, while a budget-controlled design shows the disjoint skills merge unharmed), and over-training
|
||||
disjoint specialists 1→12 epochs produces no isolation at all — the merge improves. Across every tier
|
||||
tested, **isolation had to be provoked by functional conflict; specialisation alone did not speciate**
|
||||
— a bound on the analogy that sharpens the design rule: what breaks merging is conflicting conventions
|
||||
on shared circuitry, not divergence per se.
|
||||
|
||||
*(FIG:fig5)*
|
||||
|
||||
### A controlled predictive test: functional conflict, measured pre-merge, predicts merge damage
|
||||
|
||||
The framework's prediction-level claim was put to a designed test (Fig. 6C). Thirty-nine parent pairs
|
||||
(13 conditions × 3 seeds; rows are not independent — parents share task-data seeds across conditions —
|
||||
so all inference is condition-clustered) span three axes decorrelated by construction: *conflict*
|
||||
(contradictory conventions on shared prompts, private budgets fixed), *compatible overlap* (the same
|
||||
shared prompts under the same convention — overlap and volume without conflict), and *duration* (weight
|
||||
divergence with zero conflict). Before merging, six predictors are computed: **confidence-weighted
|
||||
functional conflict** (bilateral confident disagreement on probes drawn blind to where conflict lives —
|
||||
a proposed proxy for merge-relevant interactions, motivated by the observation that raw disagreement
|
||||
counts harmless complementation, one parent merely ignorant, as conflict), raw disagreement, gradient
|
||||
alignment at the shared base (21), LoRA-delta cosine and distance, and a cross-task performance
|
||||
baseline. The pre-registered outcome is the merge penalty against oracle parent potential (the
|
||||
hybrid-load analogue), also reported against best- and mean-parent references because the predictor
|
||||
ordering is sensitive to that choice.
|
||||
|
||||
The supported conclusion, stated conditionally: **across this controlled grid, pre-merge functional
|
||||
disagreement predicted merge penalties (clustered bootstrap CIs excluding zero; held-out
|
||||
leave-one-condition-out ρ ≈ 0.35–0.40), whereas LoRA-delta cosine and L2 showed no statistically
|
||||
detectable association; gradient alignment carried intermediate signal.** Head-to-head predictor
|
||||
differences are not individually significant at this sample size, and only these baselines were
|
||||
tested. Two further results earn their place by tempering: the initial two-axis grid's best predictor
|
||||
was delta-cosine (ρ = +0.60) — an overlap artefact that the compatible-overlap control was added to
|
||||
expose, and did (collapse to +0.03); and the pre-registered internal prediction that confidence
|
||||
weighting would beat raw disagreement **failed** (they are statistically indistinguishable as rank
|
||||
predictors), so the present evidence favours functional disagreement generally, not the DMI-specific
|
||||
refinement. The framework motivated the measurement and the controls; their success does not validate
|
||||
the specifically population-genetic mechanism. Whether the prediction improves a budget-matched
|
||||
operator choice, and whether it generalises to unfamiliar conflict structures and real task pairs,
|
||||
are the experiment's open front.
|
||||
|
||||
*(FIG:fig6)*
|
||||
|
||||
## Discussion
|
||||
|
||||
**Design rules.** Read as engineering, the results compress into rules an operator of a model
|
||||
population can apply. *Ground every generation* in verified reality — a few percent retains most
|
||||
diversity — but price the rarest capabilities individually (`m·p ≳ 1`) and use recombination, not
|
||||
grounding, to reach the deep tail. *Merge, don't blend, when there is headroom*: keep specialists
|
||||
intact and route, or breed-and-screen candidate merges, whenever the naive average is far from
|
||||
ceiling; plain averaging is adequate only where a strong base has already composed the skills. *Match
|
||||
the operator to entanglement*: merge freely when skills are additive; sparingly, with offspring
|
||||
selection, when they entangle; and expect the champion-optimal mating breadth to narrow as landscapes
|
||||
roughen. *Preserve diversity as a first-class objective*, because selection can only preserve variety
|
||||
that exists, and the society result shows grounding, recombination, and diversity are jointly
|
||||
necessary. *Before merging, measure functional conflict* — cheap, pre-merge, and in our controlled
|
||||
setting predictive where weight distance was not; and expect specialisation alone to be merge-safe,
|
||||
with conflicting conventions on shared circuitry as the thing to detect and avoid.
|
||||
|
||||
**What is borrowed and what is ours.** The diagnosis — collapse as drift — is prior art (6–9), as are
|
||||
the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
|
||||
naive averaging loses to interference-aware or routed merges (1, 19, 20), that model populations can
|
||||
climb (2–5), and that merge success admits ML-native predictors (21, 22). Ours is the framework-level
|
||||
synthesis — inheritance, diversity, and compatibility as managed quantities — together with: the
|
||||
conservation law for blending inheritance and its operator boundaries; the per-item grounding floor;
|
||||
the jointly-necessary society; model speciation as a named, tested question, with the
|
||||
coordinate-versus-functional decomposition under a complete symmetry group and the emergent null that
|
||||
bounds it; and the controlled predictive test with its controls. We claim the framework generated
|
||||
these measurements and experiments; we do not claim their outcomes validate a uniquely
|
||||
population-genetic mechanism, and one refinement it proposed was not supported.
|
||||
|
||||
**Limits and open problems.** The demonstrations are deliberately small: exact where small is a virtue,
|
||||
sign-level and seed-replicated at the language-model tier, on constructed task families with a
|
||||
trivially separable router and one model lineage (Qwen, 0.5B–7B). The composed society has not been
|
||||
built at language-model scale. The predictive test's next bars, in order of value: generalisation to
|
||||
*unfamiliar* conflict structures and real task pairs; a demonstrably better *budget-matched* merging
|
||||
decision; then scale replication. Beyond engineering, the framework's hardest open problem is the
|
||||
fitness function itself: selection optimises what is measured, and for knowledge systems the
|
||||
persuasive and the true compete — grounding against a reality that can refuse is the only anchor we
|
||||
trust, and institutionalising that anchor (verification, replication, and challenge among models) is
|
||||
the society-level problem we pose but do not solve. What biology receives in return is a new model
|
||||
system: populations of learners where every genotype, environment, and mating decision is observable
|
||||
and manipulable — where the evolution of sex can be studied with interventions (unbounded parents,
|
||||
offspring preview, directed mating) that no living system permits.
|
||||
|
||||
## Materials and Methods
|
||||
|
||||
**Analytic tier.** Pure NumPy/SciPy Wright–Fisher simulator over `K`-item distributions (knowledge as
|
||||
`p_t`; Zipf-tailed truth `p*`; drift–grounding–refit generations), extended with a learning kernel
|
||||
(smoothing/sharpening refit), multi-locus genotypes on additive and Kauffman NK landscapes, n-parent
|
||||
crossover, and finite-population society loops. All parameters live in per-experiment YAML configs;
|
||||
every run derives all randomness from one master seed (`SeedSequence.spawn`) and is bitwise
|
||||
reproducible; scientific-validation tests assert the closed forms (heterozygosity decay, immigration
|
||||
equilibrium, closed-form union) to <0.5% and run in CI with 151 further correctness tests.
|
||||
|
||||
**Neural tier.** Trained-network experiments realise the same abstractions with an exact oracle:
|
||||
histogram/RNN/MLP/VAE generators on a synthetic mode universe (the histogram model reduces the harness
|
||||
exactly to the analytic tier — the bridge gate), and a convolutional VAE on MNIST with a frozen CNN
|
||||
oracle (98.5% mode accuracy; confusion matrix recorded as the measurement floor). Speciation
|
||||
experiments fork no-BatchNorm MLPs (784–512–512–10) from a shared base, weight-average, and measure
|
||||
linear-mode-connectivity error barriers before and after alignment; alignment composes deterministic
|
||||
Git Re-Basin permutation matching with exact per-unit scale canonicalisation (the complete unit
|
||||
symmetry group for this class), gated by exact recovery of a permuted-and-rescaled copy.
|
||||
|
||||
**Language-model tier.** LoRA specialists (rank 16) on procedurally generated task families with an
|
||||
exact-match verifier, on frozen Qwen2.5-Instruct bases (0.5B on one 16 GB GPU; 7B on one L40S).
|
||||
Operators: weight-space merges (soup/TIES via adapter arithmetic), per-input routing, and
|
||||
Dirichlet-sampled offspring populations screened on held-out validation splits. Multi-seed protocols
|
||||
fix the test sets and vary the training seed. The predictive test computes all predictors pre-merge
|
||||
(generation confidence from token log-probabilities; base-model gradient cosines; exact r-space
|
||||
LoRA-delta geometry) and evaluates merges on held-out tests; robust statistics (condition-clustered
|
||||
bootstrap, paired predictor contrasts, leave-one-condition-out prediction, multi-reference outcomes)
|
||||
are produced by a committed script. Statistical, per-seed reproducibility is documented for GPU tiers.
|
||||
|
||||
**Data and code availability.** All code, configs, seeds, results artifacts (with content hashes),
|
||||
figures, and a one-command reproduction script will be deposited openly (repository + archived DOI) on
|
||||
publication; every figure in this paper regenerates from committed artifacts without re-simulation.
|
||||
|
||||
## References
|
||||
|
||||
1. Yadav P, Tam D, Choshen L, Raffel C, Bansal M (2023) TIES-Merging: resolving interference when merging models. *NeurIPS*. arXiv:2306.01708.
|
||||
2. Akiba T, Shing M, Tang Y, Sun Q, Ha D (2025) Evolutionary optimization of model merging recipes. *Nat Mach Intell* 7:195–204.
|
||||
3. GENOME: Nature-inspired population-based evolution of large language models (2025). arXiv:2503.01155.
|
||||
4. Sakana AI (2025) Competition and attraction improve model fusion (M2N2). *GECCO*. arXiv:2508.16204.
|
||||
5. Subramaniam V, Du Y, Tenenbaum JB, Torralba A, Li S, Mordatch I (2025) Multiagent finetuning: self-improvement with diverse reasoning chains. arXiv:2501.05707.
|
||||
6. Shumailov I, et al. (2024) AI models collapse when trained on recursively generated data. *Nature* 631:755–759.
|
||||
7. Riis S (2026) Drift and selection in LLM text ecosystems. arXiv:2604.08554.
|
||||
8. Benati M, Londei A, Lanzieri D, Loreto V (2025) First-extinction law for resampling processes. arXiv:2509.20101.
|
||||
9. Yoon Y, Hu D, Weissburg I, Qin Y, Jeong H (2025) Model collapse in the self-consuming chain of diffusion finetuning: a quantitative trait modeling perspective. *ICLR*. arXiv:2407.17493.
|
||||
10. Muller HJ (1964) The relation of recombination to mutational advance. *Mutat Res* 1:2–9.
|
||||
11. Gerstgrasser M, et al. (2024) Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. arXiv:2404.01413.
|
||||
12. Yi B, Liu Q, Cheng Y, Xu H (2025) Escaping model collapse via synthetic data verification. arXiv:2510.16657.
|
||||
13. Wright S (1931) Evolution in Mendelian populations. *Genetics* 16:97–159.
|
||||
14. Fisher RA (1930) *The Genetical Theory of Natural Selection* (Clarendon, Oxford).
|
||||
15. Muller HJ (1932) Some genetic aspects of sex. *Am Nat* 66:118–138.
|
||||
16. Orr HA (1995) The population genetics of speciation: the evolution of hybrid incompatibilities. *Genetics* 139:1805–1813.
|
||||
17. Orr HA, Turelli M (2001) The evolution of postzygotic isolation: accumulating Dobzhansky–Muller incompatibilities. *Evolution* 55:1085–1094.
|
||||
18. Livnat A, Papadimitriou C (2016) Sex as an algorithm: the theory of evolution under the lens of computation. *Commun ACM* 59(11):84–93.
|
||||
19. Yu L, Yu B, Yu H, Huang F, Li Y (2023) Language models are super Mario: absorbing abilities from homologous models (DARE). arXiv:2311.03099.
|
||||
20. Wortsman M, et al. (2022) Model soups: averaging weights of multiple fine-tuned models. *ICML*. arXiv:2203.05482.
|
||||
21. Zhou L, Zhao B, Yu R, Rodolà E (2026) Demystifying mergeability: interpretable properties to predict model merging success. arXiv:2601.22285.
|
||||
22. Cao Y, Ran D, Guo Y, Wu M, Chen S, et al. (2026) An empirical study and theoretical explanation on task-level model-merging collapse. arXiv:2603.09463.
|
||||
23. Ainsworth S, Hayase J, Srinivasa S (2022) Git Re-Basin: merging models modulo permutation symmetries. arXiv:2209.04836.
|
||||
24. Li T, Shen Z (2026) Scaling linear mode connectivity and merging to billion-parameter pretrained transformers. arXiv:2606.23607.
|
||||
25. Kauffman SA, Levin S (1987) Towards a general theory of adaptive walks on rugged landscapes. *J Theor Biol* 128:11–45.
|
||||
26. Lehman J, Stanley KO (2011) Abandoning objectives: evolution through the search for novelty alone. *Evol Comput* 19:189–223.
|
||||
27. Pari J, Jelassi S, Agrawal P (2024) Collective model intelligence requires compatible specialization. arXiv:2411.02207.
|
||||
28. Hu EJ, et al. (2021) LoRA: low-rank adaptation of large language models. arXiv:2106.09685.
|
||||
29. Sharma E, Roy DM, Dziugaite GK (2024) The non-local model merging problem: permutation symmetries and variance collapse. arXiv:2410.12766.
|
||||
30. Kozodoi N, Afolabi Z, Butler J (2026) Are we merging the right models? Impact of expert training duration on model merging for LLMs. arXiv:2607.11997.
|
||||
BIN
paper/pnas/main.pdf
Normal file
BIN
paper/pnas/main.pdf
Normal file
Binary file not shown.
27
paper/pnas/main.tex
Normal file
27
paper/pnas/main.tex
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
% PNAS draft — readable single-column build for review/iteration (tectonic/XeLaTeX). Content is
|
||||
% generated from main.md by build.py; the pnas.cls reflow happens at submission (Phase 5).
|
||||
\ifdefined\XeTeXversion\else\ifdefined\pdfoutput\pdfoutput=1\fi\fi
|
||||
\documentclass[11pt]{article}
|
||||
|
||||
\usepackage[a4paper, margin=1.0in]{geometry}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{amsmath, amssymb}
|
||||
\usepackage[hidelinks]{hyperref}
|
||||
\usepackage{microtype}
|
||||
|
||||
\setlength{\parskip}{0.35em}
|
||||
|
||||
\title{\textbf{The evolution of sex for artificial intelligence}\\[0.5em]
|
||||
\large A population-genetic framework for multigenerational model populations}
|
||||
\author{Giorgio F.\ Gilestro\\[0.2em]
|
||||
\normalsize Department of Life Sciences, Imperial College London\\
|
||||
\normalsize \href{mailto:giorgio@gilest.ro}{giorgio@gilest.ro} \,\(\cdot\)\,
|
||||
\href{https://lab.gilest.ro}{lab.gilest.ro}}
|
||||
\date{PNAS draft --- generated from \texttt{paper/pnas/main.md}}
|
||||
|
||||
\begin{document}
|
||||
\maketitle
|
||||
|
||||
\input{body}
|
||||
|
||||
\end{document}
|
||||
135
paper/pnas/si.md
Normal file
135
paper/pnas/si.md
Normal file
|
|
@ -0,0 +1,135 @@
|
|||
# SI Appendix — The evolution of sex for artificial intelligence
|
||||
|
||||
*Skeleton assembled at Phase 4; finalised at submission. Every numbered experiment has a committed
|
||||
config (`configs/`), artifact triple (`results/<name>/results.parquet` + resolved config + manifest
|
||||
with content hashes and git commit), a README with its legend and falsifier status, and a figure that
|
||||
regenerates from the parquet alone. `reproduce.sh` re-runs everything from the master seeds.*
|
||||
|
||||
## SI Text S1–S2: formal statements
|
||||
|
||||
## S1. The incompatibility floor: what no alignment can remove (E13c)
|
||||
|
||||
**Setting.** Models A and B are trained on the same input distribution; their target label functions
|
||||
`f_A` and `f_B` agree except on a conflict set `S` of probability mass `μ(S)` (in E13's conflict
|
||||
condition, the cyclically-relabelled classes; `μ(S) ≈ conflict_frac` up to class balance). A
|
||||
*function-preserving transformation* `T` (any composition of hidden-unit permutations and, for ReLU
|
||||
networks, positive per-unit rescalings — the full unit symmetry group of a plain ReLU MLP) satisfies
|
||||
`T(B)(x) = B(x)` for all `x` by construction.
|
||||
|
||||
**Proposition 1 (endpoint invariance — with the term "chord" defined precisely).** Here "chord"
|
||||
means the α-linear interpolation **of the endpoint loss values**, `(1−α)·L(A) + α·L(B)` — the
|
||||
baseline in the barrier definition, a function of the endpoints only — NOT the weight-space
|
||||
interpolation path. For every function-preserving `T`, the endpoint functions, hence the endpoint
|
||||
losses and this chord, are identical for `(A, T(B))` and `(A, B)`. The **interpolation path itself is
|
||||
generally NOT invariant** — losses along `(1−α)·A + α·T(B)` change with `T`, which is precisely why
|
||||
alignment can lower a barrier. *(Immediate from the definition of function-preserving.)* Scope
|
||||
caveat: our aligner provably recovers a permuted-and-rescaled copy exactly — an important special
|
||||
case — but this does not establish global optimality of the alignment over the symmetry group for
|
||||
independently trained networks; the decomposition's "removable" share is therefore a lower bound, and
|
||||
the "residual" an upper bound, on their true values.
|
||||
|
||||
**Proposition 2 (no merged model can serve both parents).** Let `h` be *any* single classifier (in
|
||||
particular, any interpolated/merged model, under any alignment). On every `x ∈ S`, `f_A(x) ≠ f_B(x)`,
|
||||
so `h(x)` disagrees with at least one of them. Hence
|
||||
|
||||
`ε_A(h) + ε_B(h) ≥ μ(S)`, and therefore `max(ε_A(h), ε_B(h)) ≥ μ(S)/2`,
|
||||
|
||||
where `ε_P(h)` is `h`'s error against parent `P`'s labels. A hybrid of two models whose conventions
|
||||
conflict on mass `μ(S)` errs at rate at least `μ(S)/2` against at least one parent — **hybrid
|
||||
disadvantage with an information-theoretic floor, independent of the alignment group, the
|
||||
architecture, and the merging operator.** This is reproductive isolation in the fitness sense: past a
|
||||
given functional conflict, *no* recombination operator produces an offspring loyal to both lineages.
|
||||
|
||||
**What remains empirical, and why the experiment is designed as it is.** Propositions 1–2 do *not*
|
||||
bound the single-task path barrier (the loss along the interpolation between A and `T(B)` evaluated
|
||||
on one parent's task): in principle a path could dip toward one parent's function. Whether it does is
|
||||
exactly what E13 measures — and the measured answer is that it does not: the conflict-condition
|
||||
barrier is unchanged by permutation alignment (`residual`) *and* by alignment modulo the full
|
||||
permutation × positive-rescaling group (`residual_scale`), while the same aligner removes ~all of the
|
||||
independent-init barrier (the positive control). Richer-symmetry results for transformers
|
||||
(arXiv:2606.23607; neuron-identifiability approaches to linear mode connectivity, 2026) strengthen
|
||||
the *removable* side of the decomposition and are therefore complementary: the more barrier a larger
|
||||
group can remove for *compatible* models, the sharper the meaning of the residual that survives for
|
||||
*incompatible* ones — and Proposition 2 caps what any of them could ever achieve on the conflict set.
|
||||
|
||||
**Terminology note for the paper.** "Residual (after alignment)" = the estimated functional
|
||||
incompatibility; for ReLU MLPs we align modulo the full unit symmetry group, so the estimate is not
|
||||
confounded by missed symmetries of that architecture class.
|
||||
|
||||
## S2. Emergent vs imposed incompatibility (E13b framing)
|
||||
|
||||
The conflict condition *imposes* contradiction (the two label maps disagree on `S`), which pins
|
||||
`μ(S) > 0` and activates Proposition 2. A true Bateson–Dobzhansky–Muller incompatibility is
|
||||
*emergent*: each lineage's substitutions are harmless on their own background (`μ(S) = 0` — the
|
||||
training signals never contradict), and incompatibility, if any, arises only in the *combination*.
|
||||
The `disjoint` (complementary class specialists) and `augment` (divergent input conventions)
|
||||
conditions realise this: any residual barrier they develop cannot be attributed to label conflict and
|
||||
is the emergent-speciation signal proper. Pre-registered readings: residual grows with divergence →
|
||||
model speciation is emergent in real weights (E12's trajectory realised); residual stays at the
|
||||
`shared`-control level → within this regime, trained networks are *more* merge-compatible than the
|
||||
biological analogy predicts — an honest bound on the analogy, and itself a design-relevant result
|
||||
(merging is safe absent functional conflict).
|
||||
|
||||
**Outcome (2026-08-11 run, 4 reps, t_div ≤ 3200): the second reading.** Residual 0.000 at every
|
||||
divergence in both emergent conditions, and the merge *rescues* the forgetting `disjoint` specialists
|
||||
(parents → 0.535/0.474 on the full task; merged ≈ 0.955 throughout — a sustained Fisher–Muller rescue
|
||||
at zero barrier). Isolation in real weights required functional conflict in this regime; whether
|
||||
long-horizon over-specialisation erodes mergeability at LLM scale (cf. arXiv:2607.11997) is the
|
||||
`llm_speciation` question (Phase 3).
|
||||
|
||||
|
||||
## SI Table S1: the claims ledger (status / assumptions / evidence / limits)
|
||||
|
||||
| Claim | Status | Key assumptions | Evidence | Known limits |
|
||||
|---|---|---|---|---|
|
||||
| Collapse = Wright–Fisher drift (minimal model) | Exact (diagnosis conceded to prior work) | Knowledge = categorical distribution; refit = resample | Closed forms reproduced to <0.5% | Real learners add a signed, architecture-specific estimator bias (measured) |
|
||||
| Grounding = immigration; critical real-data fraction ≪ 1 | Exact + empirical sign | Fresh samples from a fixed, non-drifting truth | Exact `H_eq`; `g*≈0.048`; sign holds in RNN/MLP/VAE and on MNIST | Deepest tail unrescuable at feasible budgets (`m ∼ 1/p`); sharp threshold softens in trained nets |
|
||||
| "Merge, don't average" conservation | Exact **for the output-mean operator** | Rare-item regime; an oracle/verifier identifies the strongest source | E4 closed form + simulation; neural reproduction | Weight-averaging and routing are empirical cousins, not instances; budgets differ; bridge = the headroom rule |
|
||||
| Offspring exceed every parent (Fisher–Muller) | Interpretation + empirical | Complementary (decorrelated) parents; verifiable fitness | E8 analytic; 7B LoRA merge beats every specialist on every family | LLM tier: 3 lexically-distinct families; multi-seed replication in progress |
|
||||
| Outbreeding depression on rugged landscapes; operator design rule | Exact-model result; hypothesis at LLM scale | NK epistasis stands in for skill entanglement | E9–E10; directed selection rescues | Not yet mapped onto a real task-entanglement measure |
|
||||
| Optimal mate-pool breadth shrinks with ruggedness | Exact-model result; hypothesis for merging populations | Ring population, local selection | E14 | Phenomenon known to island-model evolutionary computation; our contribution is the mapping and the diversity/mean decomposition |
|
||||
| Merge failure decomposes into coordinate artefact + functional residual | Empirical (MLP tier; LLM tier in progress) | Alignment enumerates the architecture's unit symmetries | Full-symmetry residual ≈ 0 (compatible) vs ≈ naive (conflict); cliff in hybrid fitness | Scoped to aligned linear interpolation; conflict floor is information-theoretic, not genetic |
|
||||
| Epistasis (not divergence) sets the cliff; snowball onset | Exact-model result; **hypothesis** at the neural tier | BDM incompatibility structure | E12 | Snowball count ≠ performance cliff without the effect-size link; neural test outstanding |
|
||||
| Pre-merge functional disagreement predicts merge penalty | Empirical, within a controlled grid (0.5B, 13 conditions × 3 seeds) | Constructed conflict/overlap/duration axes; oracle-potential outcome (pre-registered; ordering sensitive to reference) | Clustered CIs exclude 0; held-out LOCO ρ≈0.4; selected geometry baselines ≈ 0 | Head-to-head predictor differences not individually significant; only selected baselines; generalisation to real task pairs open |
|
||||
| Confidence weighting improves rank prediction over raw disagreement | **Not supported** (pre-registered internal prediction) | — | Paired Δ\|ρ\| ≈ −0.02, CI [−0.13, +0.06] | Weighting does double the conflict-vs-compat level contrast |
|
||||
| The predictor improves budget-matched operator choice | **Open** | — | Soup-vs-route gap readout noise-dominated at 0.5B | The practical payoff; untested |
|
||||
| Emergent speciation without conflict | **Not observed** (pre-registered) | Shared ancestry, compatible tasks, tested divergences | E13b: residual 0.000; merge rescues specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
|
||||
| Grounding + sex + diversity jointly necessary | Exact-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |
|
||||
|
||||
## SI Methods (per tier — full details in the per-experiment READMEs and configs)
|
||||
|
||||
**Analytic tier (E1–E14).** Wright–Fisher simulator over K-item distributions; closed-form validation
|
||||
suite (`tests/test_scientific_validation.py`, <0.5% tolerances); learning kernel; multi-locus
|
||||
genotypes, NK landscapes, n-parent crossover (E7–E11); BDM speciation model (E12); mating structure
|
||||
(E14). Bitwise reproducible from master seeds.
|
||||
|
||||
**Neural tier.** Histogram bridge (exact reduction to the analytic tier — the harness gate);
|
||||
RNN/MLP/VAE collapse+grounding on a synthetic mode universe with an exact oracle; conv-VAE on MNIST
|
||||
with a frozen CNN oracle (98.5% mode accuracy; 30x30 confusion matrix recorded as the measurement
|
||||
floor); E13 speciation: no-BatchNorm MLPs, weight-average merges, LMC error barriers before/after
|
||||
alignment under the complete unit symmetry group (deterministic Re-Basin permutation matching composed
|
||||
with exact scale canonicalisation; sanity gate recovers a permuted-and-rescaled copy exactly);
|
||||
pre-registered emergent conditions (disjoint classes; shifted-view conventions).
|
||||
|
||||
**Language-model tier.** LoRA rank-16 specialists on procedural task families with an exact-match
|
||||
verifier; Qwen2.5-Instruct 0.5B/7B; operators: soup/TIES adapter arithmetic, per-input routing,
|
||||
Dirichlet offspring populations screened on held-out validation; multi-seed protocol (fixed tests,
|
||||
varied training seed); speciation knobs (conflicting conventions on ambiguous prompts; duration);
|
||||
the controlled predictive test (six pre-merge predictors; three decorrelated axes; robust statistics
|
||||
via `figures/stats_llm_epistasis.py`: condition-clustered bootstrap, paired contrasts,
|
||||
leave-one-condition-out prediction, three outcome references). Statistical (per-seed)
|
||||
reproducibility documented for GPU tiers.
|
||||
|
||||
## SI Statistics
|
||||
|
||||
Output of `figures/stats_llm_epistasis.py` (clustered CIs, paired predictor contrasts, LOCO held-out
|
||||
prediction, outcome-reference sensitivity, within/between-axis decomposition) — reproduced verbatim at
|
||||
submission. Chronology of the predictive test (prospective / adaptive / post-hoc) as disclosed in
|
||||
`results/llm_epistasis/README.md`.
|
||||
|
||||
## SI Figures
|
||||
|
||||
One per experiment, regenerated from committed artifacts: E1–E14, bridge/collapse/grounding/
|
||||
architectures/recombination, kernel (sharpen/smooth), mnist_collapse (+ montage), speciation_real
|
||||
(decomposition/cliff/emergent), llm_merge(_hpc/_seeds), llm_moe(_hpc/_hard_hpc/_hard_seeds),
|
||||
llm_directed(_hpc/_hard_hpc/_hard_seeds), llm_speciation(_add), llm_epistasis(_compat).
|
||||
Loading…
Add table
Add a link
Reference in a new issue