third review round: mathematical corrections + operator separation + headline calibration
The five priority fixes, in the PNAS draft and propagated to the
long-form document and results documentation:
1. The averaging proposition now proves what it claims: a FIRST-ORDER
cancellation of the multi-parent retention gain under output-mean
inheritance in the rare-item regime (n·p/K << 1), with the convexity
boundary stated (averaging's variance reduction can reduce extinction
outside that regime — the reviewer's argument) and the union
operator's renormalisation + oracle requirement explicit. "Adding
parents cannot help" deleted everywhere.
2. Grounding: g*~=0.05 restated as an operational threshold (equilibrium
smooth in g — no phase transition); m·p floor restated as
1−exp(−m·p) per-batch observation probability with
retention/occupancy/reintroduction distinguished; the deep-tail rule
de-categoricalised (stratified sampling; recombination recovers only
what parents retain).
3. Grounded INHERITANCE (data channel) separated from grounded
EVALUATION (fitness channel) in the society section; retitled to
"complementary contributions"; general joint necessity disclaimed.
Table 1 + v6 ledger updated.
4. Alignment contradiction removed everywhere ("cannot be an alignment
failure" -> the reviewer's formulation); abstract says "remaining
after permutation-and-rescaling alignment"; group = search space,
control recovery != global optimality; "specialisation is merge-safe"
-> "do not treat divergence/specialisation alone as evidence of
incompatibility".
5. Significance headline matched to the bounded evidence; seed-
dependence sensitivity added (per-seed rho stable +0.37..+0.53 for
functional measures, ~0 for geometry, gradient alignment
seed-UNSTABLE −0.11..−0.55 — reported as its own caveat; LOSO ranges
in stats script).
Presentation: review-process meta-language stripped; "exact" reserved
for closed forms ("analytic model" labels); headroom rule qualitative;
directed-sex phrasing per review; ratchet = consequence-level
correspondence; compact results table (Table 2) added. Response letter:
paper/response-to-review-3.md. Both PDFs rebuilt; 151 tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
parent
bf4b1c077c
commit
6b5591c92f
9 changed files with 322 additions and 138 deletions
|
|
@ -81,9 +81,11 @@ question the framework is built to test.
|
|||
|
||||
From the geneticist's apparatus we extract falsifiable, load-bearing claims (each stated with its
|
||||
operator and scope in the text): (i) **"merge, don't average"** — a conservation result: refitting a
|
||||
child to the *mean of its parents' output distributions* conserves rare-capability mass at the
|
||||
single-parent level, so adding parents cannot help, while union-preserving operators realise the gain
|
||||
— exact in the minimal model, with its weight-space image verified as the headroom rule below; (ii)
|
||||
child to the *mean of its parents' output distributions* conserves expected rare-capability mass at
|
||||
the single-parent level, cancelling the multi-parent gain *to first order in the rare-item regime*
|
||||
(outside it, variance reduction from averaging can help — the result is a first-order cancellation,
|
||||
not a universal impossibility), while union-preserving operators realise the gain in all regimes —
|
||||
derived in the minimal model, with its weight-space image the headroom rule below; (ii)
|
||||
**offspring can exceed every parent** (Fisher–Muller), the real argument for sex in model societies;
|
||||
(iii) on **rugged, epistatic** task landscapes, blind recombination causes **outbreeding depression**,
|
||||
yielding a design rule — *merge freely when skills are additive, sparingly and with selection when
|
||||
|
|
@ -100,8 +102,8 @@ pre-registered and found: absent conflicting training signals, divergently-speci
|
|||
shared ancestry developed *no* isolation at any divergence tested, the merge instead *rescuing* the
|
||||
forgetting specialists. Isolation must be provoked by conflict; specialisation alone did not speciate.
|
||||
AI also has an advantage biology lacks: **directed sex** — unbounded parents, chosen mates,
|
||||
and offspring screened before they are kept — which converts recombination from a gamble into a
|
||||
reliable engine and has no biological analogue.
|
||||
and offspring screened before they are kept — engineered recombination with a flexibility of parent
|
||||
choice and pre-deployment screening that natural mating systems do not approach.
|
||||
|
||||
We support the argument with **minimal, reproducible models** — a closed-form-exact account of drift
|
||||
and grounding, the same effects in small trained networks and an MNIST image generator, a real-weight
|
||||
|
|
@ -322,9 +324,11 @@ the fancier operator.
|
|||
Four different operators travel under these words, and the conservation result belongs to exactly one
|
||||
of them. What is *derived* is this: when a pupil's knowledge is refit to the **mean of the parents'
|
||||
output distributions**, the expected mass on any rare item is conserved at the single-parent level —
|
||||
in the rare-item regime the 1/K dilution of averaging exactly cancels the union gain of having K
|
||||
parents — so adding parents cannot help; whereas an operator that keeps, per item, its **strongest
|
||||
source** realises the union. That statement is exact in the minimal model, and it presupposes an
|
||||
in the rare-item regime (`n·p/K ≪ 1`) the 1/K dilution of averaging cancels the union gain of having
|
||||
K parents to first order — outside that regime, survival is convex in mixed mass and averaging's
|
||||
variance reduction can help, so this is a first-order cancellation, not a universal impossibility;
|
||||
whereas an operator that keeps, per item, its **strongest source** (and renormalises, which itself
|
||||
redistributes mass) realises the union in all regimes. That statement is exact in the minimal model, and it presupposes an
|
||||
oracle (or verifier) able to say which source is strongest. The two operators the LLM prototype
|
||||
tests — **weight averaging** (a nonlinear network's weight-mean does not compute the mean of its
|
||||
parents' outputs) and **routing among intact specialists** (which keeps K models' storage and an input
|
||||
|
|
@ -436,8 +440,9 @@ on the same task* have a real naive barrier that alignment removes almost entire
|
|||
and the aligned merge performs at parent level) — same species, different basis, the canonical Re-Basin
|
||||
result, which also proves the aligner works. Two nets that learned *conflicting* label maps have a
|
||||
large barrier of which the full symmetry group removes **essentially nothing** (0.502 → 0.497) —
|
||||
genuine reproductive isolation, not a missed symmetry, and it cannot be dismissed as a failure to align
|
||||
because the very same aligner erased the same-task barrier. It also carries a floor no future alignment
|
||||
a conflict-associated barrier the tested alignment leaves largely unchanged — supporting a
|
||||
functional-conflict interpretation without proving optimal alignment (control recovery validates a
|
||||
special case; the removable share is a lower bound, the residual an upper bound). It also carries a floor no future alignment
|
||||
method can breach: models loyal to label maps that conflict on a fraction *μ* of inputs cannot both be
|
||||
served by *any* single merged model, which must err at rate ≥ *μ*/2 against at least one parent
|
||||
(SI proposition). Sweeping the fraction of conflicting classes traces the **isolation cliff in real
|
||||
|
|
@ -754,8 +759,9 @@ shape.)
|
|||
offspring, beats the naive average — but *only when the task leaves headroom*. On easy tasks a strong
|
||||
model's plain average is already at the ceiling and the refinements add nothing; on hard tasks the
|
||||
average dilutes a specialist below even the best single parent, and the union-preserving operators win
|
||||
clearly. The practical rule is exact: these tricks pay off in proportion to how far the naive average
|
||||
is from the best attainable. This is a prototype (three task families, one seed), so we read it as
|
||||
clearly. The practical rule, stated qualitatively: these tricks pay off where the naive average falls
|
||||
short of attainable performance, and add nothing where it does not (a quantitative form is untested).
|
||||
This is a prototype (three task families, one seed), so we read it as
|
||||
signs, not magnitudes; the *whole grounded society* on a language model remains the open step.
|
||||
- *The whole society, and why every part is needed.* In a population evolving on a "reality" landscape,
|
||||
the full system — grounding + sexual recombination + preserved diversity — climbs to the top while
|
||||
|
|
@ -786,7 +792,7 @@ falsifier, not yet established):
|
|||
| Confidence weighting improves rank prediction over raw disagreement | **Not supported** (pre-registered internal prediction) | — | Paired Δ\|ρ\| ≈ −0.02, CI [−0.13, +0.06] | Weighting does double the conflict-vs-compat level contrast |
|
||||
| The predictor improves budget-matched operator choice | **Open** | — | Soup-vs-route gap readout noise-dominated at 0.5B | The practical payoff; untested |
|
||||
| Emergent speciation without conflict | **Not observed** (pre-registered) | Shared ancestry, compatible tasks, tested divergences | E13b: residual 0.000; merge rescues specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
|
||||
| Grounding + sex + diversity jointly necessary | Exact-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |
|
||||
| Grounding + sex + diversity complementary (each ablation fails distinctly) | Analytic-model result; hypothesis at LLM scale | Conformity stands in for self-consumption; general joint necessity not established | E11 four-arm ablation | The full grounded LLM society is unbuilt; alternative schemes untested |
|
||||
|
||||
**What is borrowed, and what is ours.** We are deliberate about the ledger, because the surrounding
|
||||
literature is crowded and a reader deserves to know exactly where the line falls. **Conceded as prior
|
||||
|
|
@ -811,8 +817,9 @@ from a thing you must run a search to discover into a thing the landscape's rugg
|
|||
the operator-choice design rule that follows (average / union-route / directed-select); **grounding as
|
||||
migration–drift balance**, giving a critical real-data fraction and a phase boundary a closed
|
||||
self-consuming loop cannot have; **directed sex** as the distinctly-AI advantage (unbounded parents,
|
||||
offspring preview, mate choice); and the **integrated society** whose four operators are shown *jointly
|
||||
necessary*. The value-add over the machine-learning-native merge theory is that ours predicts *which
|
||||
offspring preview, mate choice); and the **integrated society** whose operators make
|
||||
*complementary, distinctly-failing contributions* in the tested model (general joint necessity is not
|
||||
established). The value-add over the machine-learning-native merge theory is that ours predicts *which
|
||||
operator to use and when it will backfire*, not merely how fast quality decays. And it opens — and
|
||||
begins to occupy — a question nobody has framed: **model speciation**, the population-genetics of
|
||||
*reproductive isolation* (Bateson–Dobzhansky–Muller incompatibilities) as the account of *when two
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue