third review round: mathematical corrections + operator separation + headline calibration

The five priority fixes, in the PNAS draft and propagated to the
long-form document and results documentation:

1. The averaging proposition now proves what it claims: a FIRST-ORDER
   cancellation of the multi-parent retention gain under output-mean
   inheritance in the rare-item regime (n·p/K << 1), with the convexity
   boundary stated (averaging's variance reduction can reduce extinction
   outside that regime — the reviewer's argument) and the union
   operator's renormalisation + oracle requirement explicit. "Adding
   parents cannot help" deleted everywhere.
2. Grounding: g*~=0.05 restated as an operational threshold (equilibrium
   smooth in g — no phase transition); m·p floor restated as
   1−exp(−m·p) per-batch observation probability with
   retention/occupancy/reintroduction distinguished; the deep-tail rule
   de-categoricalised (stratified sampling; recombination recovers only
   what parents retain).
3. Grounded INHERITANCE (data channel) separated from grounded
   EVALUATION (fitness channel) in the society section; retitled to
   "complementary contributions"; general joint necessity disclaimed.
   Table 1 + v6 ledger updated.
4. Alignment contradiction removed everywhere ("cannot be an alignment
   failure" -> the reviewer's formulation); abstract says "remaining
   after permutation-and-rescaling alignment"; group = search space,
   control recovery != global optimality; "specialisation is merge-safe"
   -> "do not treat divergence/specialisation alone as evidence of
   incompatibility".
5. Significance headline matched to the bounded evidence; seed-
   dependence sensitivity added (per-seed rho stable +0.37..+0.53 for
   functional measures, ~0 for geometry, gradient alignment
   seed-UNSTABLE −0.11..−0.55 — reported as its own caveat; LOSO ranges
   in stats script).

Presentation: review-process meta-language stripped; "exact" reserved
for closed forms ("analytic model" labels); headroom rule qualitative;
directed-sex phrasing per review; ratchet = consequence-level
correspondence; compact results table (Table 2) added. Response letter:
paper/response-to-review-3.md. Both PDFs rebuilt; 151 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
Giorgio Gilestro 2026-09-06 19:29:09 +01:00
parent bf4b1c077c
commit 6b5591c92f
9 changed files with 322 additions and 138 deletions

View file

@ -13,8 +13,9 @@ genetic drift. This paper imports the other half of population genetics: the bio
reproduction. It treats model merging as recombination, real data as immigration, and merge failure as
reproductive isolation, and tests each correspondence in simulations, small neural networks, and
language models. The framework yields design rules — when to average models, when to keep them
separate, how much real data suffices — and a first controlled test showing that measured functional
conflict, not weight distance, predicts when merging fails.
separate, how much real data suffices — and a controlled small-model test in which pre-merge
functional disagreement predicted merge damage, motivating further comparison with weight-space
measures.
## Abstract
@ -24,18 +25,20 @@ evolutionary theory. Here we treat multigenerational model populations as system
diversity, and compatibility must be managed, and transfer the quantitative apparatus of the evolution
of sex. We take as settled that training on model output is genetic drift (model collapse). In a
minimal inheritance model that is exactly WrightFisher — and measurably WrightFisher-plus-bias in
trained networks — we derive and test the remedies: grounding as immigration, with a critical
real-data fraction far below one but a per-capability floor that leaves the rarest knowledge
unrescuable; recombination, where averaging parents' output distributions exactly cancels the benefit
of multiple parents while union-preserving operators realise it; the FisherMuller effect, with merged
language-model specialists exceeding every parent in replicated experiments; outbreeding depression on
rugged task landscapes, converted into reliable gains by directed, offspring-screened recombination;
and population structure, where the optimal mating breadth shrinks as skills entangle. Sex has a
limit: we introduce model speciation — merge failure as reproductive isolation — and show in trained
networks that a merge barrier surviving the full function-preserving symmetry group tracks functional
conflict, that isolation did not emerge from compatible specialisation, and, in a controlled
predictive test, that pre-merge functional disagreement predicts merge damage where weight-geometry
baselines do not. We state precisely what is exact, what is measured, and what remains hypothesis.
trained networks — we derive and test the remedies: grounding as immigration, where a real-data
fraction far below one retained most equilibrium diversity in the tested settings, with a
per-capability observation floor that makes the rarest knowledge expensive under unstratified
sampling; recombination, where refitting to the mean of parents' output distributions cancels the
multi-parent gain to first order in the rare-item regime while union-preserving operators realise it;
the FisherMuller effect, with merged language-model specialists exceeding every parent in replicated
experiments; outbreeding depression on rugged task landscapes, converted into reliable gains by
directed, offspring-screened recombination; and population structure, where the optimal mating breadth
shrinks as skills entangle. Sex has a limit: we introduce model speciation — merge failure as
reproductive isolation — and show in trained networks that a merge barrier remaining after
permutation-and-rescaling alignment tracks functional conflict, that isolation did not emerge from
compatible specialisation, and, in a controlled predictive test, that pre-merge functional
disagreement predicted merge damage while the tested weight-geometry baselines showed no detectable
association. We state precisely what is exact, what is measured, and what remains hypothesis.
---
@ -67,17 +70,20 @@ and **prediction** (the framework forecasts an unmeasured outcome). The paper is
first; makes concrete progress on the second — separating merge failures that are coordinate artefacts
from those that are functional; and reports a first, bounded step on the third — a controlled
predictive test in which pre-merge functional-disagreement measures, chosen by the framework,
predicted merge damage on a constructed task grid while weight-geometry baselines did not.
predicted merge damage on a constructed task grid while the tested weight-geometry baselines showed
no detectable association.
The correspondences we develop, summarised in Table 1: single-teacher retraining is **asexual
reproduction**, and the irreversible arm of its decay corresponds to **Muller's ratchet** (10) — once
every copy of a rare capability is gone from all parents and sources, no recombination can rebuild it,
which is precisely why remedies must act before fixation-by-loss. Injecting verified real data is
reproduction**, and the irreversible arm of its decay shares the defining consequence of **Muller's
ratchet** (10) — once every copy of a rare capability is gone from all parents and sources, no
recombination can rebuild it, which is why remedies must act before fixation-by-loss (a consequence-
level correspondence: the minimal model lacks the ratchet's recurrent deleterious-mutation mechanism,
so irreversible loss alone does not identify that specific mechanism). Injecting verified real data is
**immigration** from a non-drifting source (1113). Model merging is **recombination**, and its
celebrated payoff — a merged model exceeding every parent — is the **FisherMuller effect** (14, 15).
Merging entangled skills courts **outbreeding depression**; screening many candidate merges is a form
of **directed sex** with no biological analogue; restricting who merges with whom is **population
structure**. And merging's hard limit — models too diverged in function to combine — is **reproductive
Merging entangled skills courts **outbreeding depression**; screening many candidate merges is
engineered recombination with unusually flexible parent choice and pre-deployment screening (we use
the shorthand **directed sex**); restricting who merges with whom is **population structure**. And merging's hard limit — models too diverged in function to combine — is **reproductive
isolation**, for which the BatesonDobzhanskyMuller theory of incompatibilities (16, 17) supplies the
structure. The nearest precursor to this programme reads sex as an algorithm for mixability in the
theory of computation (18), pre-dating model merging; the model-merging literature itself has strong
@ -88,10 +94,9 @@ We support the framework at three tiers of evidence, in ascending realism and de
**minimal analytic model** validated against closed forms to a fraction of a percent; **small trained
networks** (MLPs, recurrent networks, an MNIST image generator) where the operators are measured in
real weights; and **language models** (LoRA-specialised Qwen models, 0.5B locally and 7B on a compute
cluster) where the claims are tested as signs under seed replication. Throughout, we report negative
and tempering results with the same prominence as confirmations: they include the failure of an
internal pre-registered prediction, a null on emergent speciation that bounds the analogy, and the
sensitivity analyses that temper the predictive test.
cluster) where the claims are tested as signs under seed replication. Negative results are reported with the same prominence as confirmations; they include the failure of
an internal pre-registered prediction, a null on emergent speciation that bounds the analogy, and the
sensitivity analyses on the predictive test.
## The minimal model, and where its exactness ends
@ -109,8 +114,8 @@ approximation, optimisation noise, and inductive bias, and when trained networks
exact drift null they deviate in *opposite, architecture-specific* directions: a smoothing recurrent
network resists collapse (keeping spurious variants alive), while a sharpening image generator
accelerates it. A one-parameter **learning kernel** (a smoothing knob and a sharpening knob on the
refit) reproduces both. The honest statement, used throughout: a real learner is WrightFisher *plus a
signed, measurable estimator bias* — and the drift signs (rare-first loss; the grounding response)
refit) reproduces both. Throughout, a real learner is therefore treated as WrightFisher *plus a signed, measurable
estimator bias* — and the drift signs (rare-first loss; the grounding response)
survived that bias in every architecture we tested, including a convolutional VAE retrained on its own
generated digits, where the dry lineage collapses to a single blurred digit class while 10% grounding
holds all thirty modes (Fig. 1).
@ -126,25 +131,33 @@ known limits is SI Appendix, Table S1.
| Immigration from a fixed source | Grounding with verified real data | Exact equilibrium; signs in RNN/MLP/VAE/MNIST |
| Muller's ratchet (asexual decay) | Irreversible arm of model collapse | Correspondence, scoped: applies to unrecoverable loss |
| Recombination / sexual reproduction | Model merging | Empirical at 0.5B7B |
| FisherMuller effect | Merged specialists exceed every parent | Exact-model result; replicated in LLMs |
| Outbreeding depression under epistasis | Merging entangled skills harms offspring | Exact-model (NK landscapes); hypothesis at LLM scale |
| Mating systems / population structure | Who merges with whom (breadth of the parent pool) | Exact-model result; hypothesis for real populations |
| FisherMuller effect | Merged specialists exceed every parent | Analytic model; replicated in LLMs |
| Outbreeding depression under epistasis | Merging entangled skills harms offspring | Analytic model (NK landscapes); hypothesis at LLM scale |
| Mating systems / population structure | Who merges with whom (breadth of the parent pool) | Analytic model; hypothesis for real populations |
| Reproductive isolation (BDM incompatibilities) | Merge failure from functional conflict | Empirical (MLP + LLM tiers, conflict-associated); emergent form not observed |
| Selection on a fitness function | Verifier-anchored selection ("reality that can say no") | Exact-model result (jointly necessary with sex and diversity) |
| Selection on a fitness function | Verifier-anchored selection ("reality that can say no") | Analytic model (complementary with recombination and diversity in the tested society) |
## Results
### Grounding is immigration: cheap, with a floor
In the minimal model, grounding from a fixed real source is immigration into a drifting population,
and the equilibrium diversity has a closed form our simulator matches exactly. The engineering
headline is the *magnitude*: a critical grounding fraction `g* ≈ 0.05` retains most diversity
indefinitely — real data is cheap insurance. But the same analysis yields a floor the field's
average-loss framing misses: an individual capability of rarity `p` survives only if the *absolute*
real-data budget satisfies `m·p ≳ 1`. Protecting the rarest knowledge is priced per item, at cost
`∝ 1/p`, and no affordable grounding fraction rescues the deepest tail — that requires recombination
(next section). In trained networks the *sign* of the grounding response transfers everywhere we
looked, with two honest deviations, both traced to the estimator bias above: sharp thresholds soften,
and the equilibrium diversity has a closed form our simulator matches exactly. That equilibrium is
*smooth* in the grounding fraction — there is no phase transition in aggregate diversity — so the
practical number is an operational threshold, and we define it as such: under the tested population
size and Zipf source distribution, `g ≈ 0.05` retained most (≥95%) of equilibrium diversity
indefinitely, with the required fraction depending on sample size, source distribution, and the
chosen retention target (dependencies in SI). The engineering point survives the definition: verified
real data is cheap insurance at fractions far below one. But the same analysis yields a floor the field's
average-loss framing misses: under unstratified sampling from the source, a capability of rarity `p`
appears in a real-data batch of size `m` with probability `1 e^{m·p}`, so `m·p ≈ 1` marks roughly a
63% chance of one example per batch — a soft observation floor, with higher confidence priced
accordingly, and with distinct consequences for continuous retention, stationary occupancy, and
reintroduction after loss (immigration can restore an absent item; SI separates these). Protecting the
rarest knowledge under unstratified grounding is therefore priced per item at cost `∝ 1/p`; targeted
or stratified sampling changes that cost, and recombination can recover rare capabilities *that are
still retained across complementary parents* (next section). In trained networks the *sign* of the grounding response transfers everywhere we
looked, with two deviations, both traced to the estimator bias above: sharp thresholds soften,
and support-counting metrics decouple from truth (forward-KL is the operative collapse metric for a
smoothing learner). On real images (Fig. 1B), dry self-training collapses a convolutional VAE to one
mode while ~10% grounding holds all thirty (the trained model needs roughly twice the exact-operator
@ -155,17 +168,25 @@ fraction — the measured price of the estimator bias).
### Recombination: a conservation law, its operators, and offspring that exceed every parent
Merging is where the evolution-of-sex apparatus pays for itself, beginning with a result about the
obvious operator. **Averaging is blending inheritance, and it cancels the benefit of multiple
parents:** when a child is refit to the *mean of its parents' output distributions*, the expected mass
on any rare item is conserved at the single-parent level — in the rare-item regime the 1/K dilution of
averaging exactly cancels the union gain of K parents, so adding parents cannot help. An operator that
keeps, per item, its strongest source (which presupposes a verifier or oracle to say which) realises
the union. That statement is exact for those operators in the minimal model. The practically important
obvious operator, stated with its assumptions. **Proposition (blending inheritance, rare-item
regime).** Let K parents independently retain a rare item (mass `p` when retained), and let the child
draw `n` samples either from one parent chosen at random or from the *mean of the parents' output
distributions*. Expected item mass is identical under the two schemes; and in the rare-item regime
`n·p/K ≪ 1`, where per-item survival is first-order in sampled mass, expected *survival* is also
identical — the 1/K dilution of averaging cancels the K-parent union gain to first order, so in this
regime adding parents through the output-mean does not increase expected tail retention. Two
boundaries: outside that regime, survival is a convex function of mixed mass, so the variance
reduction from averaging can *reduce* extinction relative to a randomly chosen single parent — the
cancellation is a first-order result about rare items, not a universal impossibility; and the
contrasting union operator (keep each item's strongest source, then renormalise — which itself
redistributes mass, and presupposes a verifier or oracle to identify the strongest source) increases
expected retention with K in all regimes in the minimal model. The practically important
operators — **weight averaging** (a nonlinear network's weight-mean does not compute its parents'
output-mean) and **routing among intact specialists** (different storage and inference budgets from a
single child) — are its empirical cousins, and the measured bridge is a **headroom rule**: in language
models, union-preserving operators beat the weight-average in proportion to how far that average is
from the best attainable. On easy tasks a capable base's average is already at ceiling and refinements
single child) — are its empirical cousins, and the measured bridge is a **headroom rule**, stated qualitatively: in language models,
union-preserving operators beat the weight-average where that average falls short of attainable
performance, and add nothing where it does not (easy-versus-hard contrasts at two scales; a
quantitative form of the relationship is untested). On easy tasks a capable base's average is already at ceiling and refinements
add nothing; on hard tasks the average dilutes a fragile specialist below even the best single parent
and routing wins by a wide margin (Fig. 6AB).
@ -196,18 +217,23 @@ structured-population search, mapped onto merging populations.
*(FIG:fig3)*
### The society: grounding, sex, and diversity are jointly necessary
### The society: grounding, recombination, and diversity make complementary contributions
Composing the operators closes the loop (Fig. 4). A finite population of agents evolves on a rugged NK
landscape, with selection acting on a grounded score — `g`·true-fitness + (1g)·conformity to the
population's own consensus, the analogue of training on the crowd's output. A four-arm ablation
separates the failure modes: the **full** system (grounding + directed recombination +
diversity-preserving selection) climbs to near the global optimum while keeping its specialists;
remove *grounding* and the population converges confidently on an unfit consensus (self-consumption);
remove *sex* and it strands on local optima; remove *diversity* and it converges prematurely to a
worse answer. Each removal fails *differently* — the operators are jointly necessary, which is the
system-level claim the single-operator results build toward. At language-model scale this composed
loop remains unbuilt; it is the paper's largest stated gap.
Composing the operators (Fig. 4) requires one definitional distinction first. In the inheritance
model, grounding is **grounded inheritance**: external samples added to the reproduction process (the
data channel). In the society model, grounding is **grounded evaluation**: selection weights true
fitness against conformity to the population's own consensus — `g`·true-fitness + (1g)·conformity —
the analogue of scoring models by the crowd's approval (the fitness channel). These are related design
ideas — both couple the lineage to a non-drifting external signal — but they are different operators,
and we name them separately. In the tested society (a finite agent population on a rugged NK
landscape), a four-arm ablation separates the failure modes: the full system (grounded evaluation +
directed recombination + diversity-preserving selection) climbs to near the global optimum while
keeping its specialists; removing grounded evaluation converges the population confidently on an unfit
consensus (self-consumption); removing recombination strands it on local optima; removing diversity
converges it prematurely to a worse answer. Each removal fails differently — the three implementations
make complementary contributions *under the tested conditions*; general joint necessity is not
established (alternative mutation, restart, archive, or selection schemes could alter the picture). At
language-model scale this composed loop remains unbuilt; it is the paper's largest stated gap.
*(FIG:fig4)*
@ -225,21 +251,22 @@ test must establish separately.
In trained networks, the claim must survive a known alternative: merge barriers between independently
trained networks are famously *coordinate artefacts*, removable by re-aligning hidden units (23), and
richer symmetry groups remove more (24). We therefore aligned modulo the **complete**
function-preserving unit symmetry group of the architecture tested (permutation composed with per-unit
positive rescaling, for plain ReLU MLPs) and decomposed the barrier (Fig. 5B): two networks trained
from different initialisations on the *same* task have a barrier that alignment removes essentially
entirely (residual ≈ 0.001, the aligned merge performing at parent level) — coordinate, not
functional; two networks trained on *conflicting* label maps have a barrier the full group leaves
intact (0.502 → 0.497), with the merged model functionally dead — and this cannot be an alignment
failure, because the same aligner succeeded on the control. Sweeping conflict traces the cliff as
hybrid fitness, 0.97 → 0.03. Two scope notes: exact recovery of a permuted-and-rescaled copy validates
a special case rather than global optimality, so the removable share is a lower bound and the residual
an upper bound; and the conflict floor itself is information-theoretic — no single model can satisfy
contradictory conventions (SI Appendix, Proposition S2) — with the framework's role being the
*structure around it*: which divergences generate conflict, and what moves the cliff.
richer symmetry groups remove more (24). We therefore aligned under the composition of
permutation matching and exact per-unit rescaling (the unit symmetry group of plain ReLU MLPs, as the
search space) and decomposed the barrier (Fig. 5B): two networks trained from different
initialisations on the *same* task have a barrier that this alignment removes essentially entirely
(residual ≈ 0.001, the aligned merge performing at parent level) — coordinate, not functional; two
networks trained on *conflicting* label maps have a barrier the same alignment leaves largely
unchanged (0.502 → 0.497), with the merged model functionally dead. The tested alignment removes the
same-task barrier but leaves the conflict-associated barrier intact — supporting a functional-conflict
interpretation without proving optimal alignment: exact recovery of a permuted-and-rescaled copy
validates a special case, so the removable share is a lower bound and the residual an upper bound.
Sweeping conflict traces the cliff as hybrid fitness, 0.97 → 0.03. The conflict floor itself is
information-theoretic — no single model can satisfy contradictory conventions (SI Appendix,
Proposition S2) — with the framework's role being the *structure around it*: which divergences
generate conflict, and what moves the cliff.
The sharpest honesty comes from the pre-registered **emergent test**: true BDM incompatibilities are
The strongest constraint comes from the pre-registered **emergent test**: true BDM incompatibilities are
emergent (each lineage's changes harmless alone), so we let children diverge with *no conflicting
signal anywhere* — complementary class specialists, and divergent input conventions — to 6.4× the base
training. **No isolation emerged** (residual 0.000 throughout); instead the merge *rescued* the two
@ -257,8 +284,9 @@ on shared circuitry, not divergence per se.
### A controlled predictive test: functional conflict, measured pre-merge, predicts merge damage
The framework's prediction-level claim was put to a designed test (Fig. 6C). Thirty-nine parent pairs
(13 conditions × 3 seeds; rows are not independent — parents share task-data seeds across conditions —
so all inference is condition-clustered) span three axes decorrelated by construction: *conflict*
(13 conditions × 3 seeds; rows are not independent — parents share task-data seeds across conditions,
so inference is condition-clustered, and because shared seeds also couple rows *across* conditions we
report per-seed and leave-one-seed-out sensitivity alongside) span three axes decorrelated by construction: *conflict*
(contradictory conventions on shared prompts, private budgets fixed), *compatible overlap* (the same
shared prompts under the same convention — overlap and volume without conflict), and *duration* (weight
divergence with zero conflict). Before merging, six predictors are computed: **confidence-weighted
@ -274,8 +302,11 @@ The supported conclusion, stated conditionally: **across this controlled grid, p
disagreement predicted merge penalties (clustered bootstrap CIs excluding zero; held-out
leave-one-condition-out ρ ≈ 0.350.40), whereas LoRA-delta cosine and L2 showed no statistically
detectable association; gradient alignment carried intermediate signal.** Head-to-head predictor
differences are not individually significant at this sample size, and only these baselines were
tested. Two further results earn their place by tempering: the initial two-axis grid's best predictor
differences are not individually significant at this sample size; only these baselines were tested;
and with three seeds, uncertainty about seed generalisation remains substantial — though the seed
sensitivity favours the functional measures (per-seed ρ stable at +0.37 to +0.53 in each seed alone,
geometry ≈ 0 in every seed, gradient alignment seed-unstable at 0.11 to 0.55). Two further results
bound the claim: the initial two-axis grid's best predictor
was delta-cosine (ρ = +0.60) — an overlap artefact that the compatible-overlap control was added to
expose, and did (collapse to +0.03); and the pre-registered internal prediction that confidence
weighting would beat raw disagreement **failed** (they are statistically indistinguishable as rank
@ -287,21 +318,37 @@ are the experiment's open front.
*(FIG:fig6)*
**Table 2.** Headline quantitative results with sample sizes, uncertainty, and outcome definitions
(full per-experiment tables and falsifier status in SI Appendix and per-experiment documentation).
| Result | Setting / n | Outcome definition | Headline |
|---|---|---|---|
| Closed-form validation | Analytic tier; standing tests | Simulated vs closed-form H-decay, immigration equilibrium, multi-teacher union | Agreement < 0.5% |
| Grounding retention | Minimal model; 18+ replicates per point | Fraction of equilibrium diversity retained at grounding g (operational threshold) | g ≈ 0.05 retained ≥95% (tested setting); smooth in g |
| MNIST collapse & rescue | Conv-VAE, 4 replicates; frozen oracle (98.5% mode acc.) | Mode support / forward-KL over generations | Dry: 30→1 modes; 10% grounding: 30/30 held |
| FisherMuller in LLMs | 5 seeds (0.5B), fixed tests; single 7B run | Merged vs best-specialist accuracy (overall; worst family) | Ties 0.647±0.027 vs 0.592±0.009; 7B 0.87 vs 0.77 |
| Union vs blend (headroom) | 3 seeds (0.5B hard); single 7B-hard run | Paired per-seed ordering, routing vs weight-average | Routing > blend in 3/3 seeds; one catastrophic blend failure avoided |
| Speciation decomposition | MLPs, 3 replicates | LMC error barrier residual after permutation+rescaling alignment | Same-task 0.001; conflict 0.497 (naive 0.502) |
| Emergent isolation | MLPs 4 reps to 6.4× base training; LLM 1→12 epochs | Residual barrier; merged vs parent accuracy | 0.000 everywhere; merge rescues parents (≈0.955 vs ≈0.50) |
| Predictive test | 13 conditions × 3 seeds (0.5B) | Merge penalty vs oracle parent potential (pre-registered; ±: clustered 95% CI) | Functional ρ +0.45/+0.46, CI excl. 0; LOCO ρ ≈ 0.4; geometry n.s.; paired differences n.s. |
## Discussion
**Design rules.** Read as engineering, the results compress into rules an operator of a model
population can apply. *Ground every generation* in verified reality — a few percent retains most
diversity — but price the rarest capabilities individually (`m·p ≳ 1`) and use recombination, not
grounding, to reach the deep tail. *Merge, don't blend, when there is headroom*: keep specialists
population can apply. *Ground every generation* in verified reality — a few percent retained most diversity in our tested
settings — but price the rarest capabilities individually (observation probability `1 e^{m·p}` per
batch under unstratified sampling), consider targeted sampling for the deep tail, and use
recombination to recover rare capabilities still retained across complementary parents. *Merge, don't blend, when there is headroom*: keep specialists
intact and route, or breed-and-screen candidate merges, whenever the naive average is far from
ceiling; plain averaging is adequate only where a strong base has already composed the skills. *Match
the operator to entanglement*: merge freely when skills are additive; sparingly, with offspring
selection, when they entangle; and expect the champion-optimal mating breadth to narrow as landscapes
roughen. *Preserve diversity as a first-class objective*, because selection can only preserve variety
that exists, and the society result shows grounding, recombination, and diversity are jointly
necessary. *Before merging, measure functional conflict* — cheap, pre-merge, and in our controlled
setting predictive where weight distance was not; and expect specialisation alone to be merge-safe,
with conflicting conventions on shared circuitry as the thing to detect and avoid.
roughen. *Preserve diversity as a first-class objective*, because selection can only preserve variety that
exists, and in the tested society its removal produced a distinct failure mode. *Before merging,
measure functional conflict* — cheap, pre-merge, and in our controlled setting predictive where the
tested weight-distance baselines were not; and *do not treat divergence or specialisation alone as
evidence of incompatibility* — in every regime we tested, what broke merging was conflicting
conventions on shared circuitry, which is the thing to detect.
**What is borrowed and what is ours.** The diagnosis — collapse as drift — is prior art (69), as are
the empirical facts that merges can beat parents, that decorrelated parents merge better, and that
@ -309,9 +356,9 @@ naive averaging loses to interference-aware or routed merges (1, 19, 20), that m
climb (25), and that merge success admits ML-native predictors (21, 22). Ours is the framework-level
synthesis — inheritance, diversity, and compatibility as managed quantities — together with: the
conservation law for blending inheritance and its operator boundaries; the per-item grounding floor;
the jointly-necessary society; model speciation as a named, tested question, with the
coordinate-versus-functional decomposition under a complete symmetry group and the emergent null that
bounds it; and the controlled predictive test with its controls. We claim the framework generated
the society ablation with its complementary failure modes; model speciation as a named, tested question, with the
coordinate-versus-functional decomposition under permutation-and-rescaling alignment and the emergent
null that bounds it; and the controlled predictive test with its controls. We claim the framework generated
these measurements and experiments; we do not claim their outcomes validate a uniquely
population-genetic mechanism, and one refinement it proposed was not supported.
@ -345,8 +392,9 @@ exactly to the analytic tier — the bridge gate), and a convolutional VAE on MN
oracle (98.5% mode accuracy; confusion matrix recorded as the measurement floor). Speciation
experiments fork no-BatchNorm MLPs (78451251210) from a shared base, weight-average, and measure
linear-mode-connectivity error barriers before and after alignment; alignment composes deterministic
Git Re-Basin permutation matching with exact per-unit scale canonicalisation (the complete unit
symmetry group for this class), gated by exact recovery of a permuted-and-rescaled copy.
Git Re-Basin permutation matching with exact per-unit scale canonicalisation (the unit symmetry
group of this class, as the alignment search space; control recovery does not establish global
optimality), gated by exact recovery of a permuted-and-rescaled copy.
**Language-model tier.** LoRA specialists (rank 16) on procedurally generated task families with an
exact-match verifier, on frozen Qwen2.5-Instruct bases (0.5B on one 16 GB GPU; 7B on one L40S).