second review round: tempered claims, robust statistics, corrected technical statements

Analyses (figures/stats_llm_epistasis.py, committed + reproducible):
condition-clustered bootstrap CIs (functional measures exclude zero:
dis_raw [+0.04,+0.69], conf-weighted [+0.02,+0.68]; gradient alignment
[-0.59,-0.06]; geometry straddles zero), PAIRED predictor contrasts (not
individually significant — stated), leave-one-condition-out held-out
prediction (functional replicates, geometry ~0, performance baseline
unstable), three outcome references (ordering sensitive to reference —
reported, with the mechanism), between/within-axis decomposition
(within-conflict identification impossible by design; the compat axis
identifies), and seed-level paired reliability (routing/directed beat
soup 3/3 seeds incl. one catastrophic soup failure; CI-width fragility
claim withdrawn).

Renames and corrections: "decisive experiment" -> "controlled predictive
test"; "operational epistasis" -> "confidence-weighted functional
conflict (proposed proxy)"; "functional by construction" -> "controls a
major source of coordinate mismatch / conflict-associated" (module,
configs, READMEs, figures); SI proposition's "chord" defined precisely
(endpoint-loss interpolation, invariant) vs the path (not invariant) +
no-global-optimality caveat (removable = lower bound, residual = upper);
snowball count != performance cliff distinction added; claims table
gains four rows (grid finding / weighting NOT supported / functional-vs-
all-geometry not established / operator choice open); §1 ladder states
the prediction rung as a bounded small-model result.

paper/response-to-review-2.md: point-by-point, opening with the
bookkeeping correction (E13b/c were in the reviewed draft — revised
interpretation, not new results). READMEs rewritten around the four
analyses with the chronology (prospective/adaptive/post-hoc) disclosed.
151 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
This commit is contained in:
Giorgio Gilestro 2026-09-06 17:55:46 +01:00
parent 1ae950cb7c
commit a40ace1821
17 changed files with 418 additions and 95 deletions

View file

@ -164,10 +164,15 @@ different things are easily conflated: **interpretation** (an existing result is
in these terms — e.g., merged offspring beating their parents as FisherMuller), **explanation** (the
transferred mechanism accounts for observations existing accounts leave open — e.g., which merge
failures are coordinate artefacts and which are functional), and **prediction** (the framework
forecasts an unmeasured outcome and improves a design decision — e.g., an epistasis measure taken
*before* merging that beats geometry-based predictors of merge success). This paper is strongest on
the first, makes concrete progress on the second, and states the third as its open, decisive test —
proposed here with pre-registered falsifiers, not claimed as done. The organising shift we argue for
forecasts an unmeasured outcome and improves a design decision). This paper is strongest on the
first, makes concrete progress on the second, and reports a first, bounded step on the third: a
**controlled predictive test** at small scale in which pre-merge *functional-disagreement* measures —
chosen by the framework — showed a detectable, held-out-robust association with merge damage on a
constructed task grid, while the selected weight-geometry baselines did not. We are precise about
that result's boundary where it is reported: it is a small-model demonstration on a constructed grid;
the proposed epistasis-specific refinement did not outperform plain disagreement; predictor
differences are not individually significant head-to-head; and whether the prediction improves a
budget-matched operator choice remains open. The organising shift we argue for
is prior to any single mechanism: **treat multigenerational model populations as systems whose
inheritance, diversity, and compatibility must be managed — not merely as collections of models to
optimise.**
@ -470,9 +475,15 @@ cannot satisfy two contradictory answer conventions — is information-theoretic
genetics; what the genetic frame adds is *structure around it*: which divergences generate such
conflicts, the prediction that epistasis rather than distance sets the cliff's position, and the
snowball's super-linear onset — the latter two verified so far only in the analytic model, and
therefore carried as **hypotheses at the neural tier, not results**. Second, our alignment removes the
symmetries we enumerate for this architecture class; richer transformation families for other
architectures could reapportion removable vs residual, though not below the conflict floor. Third,
therefore carried as **hypotheses at the neural tier, not results**. (On the snowball, one more
distinction: super-linear growth in the *number* of incompatibilities does not by itself entail a
sharp *performance* cliff — that needs the link from incompatibility count through effect sizes to
measured performance, which the analytic model supplies under its assumptions and any neural test
must establish separately.) Second, our alignment removes the symmetries we enumerate for this
architecture class, and exactly recovering a permuted-and-rescaled copy validates a special case
rather than proving global optimality for independently trained networks — so the removable share is
a lower bound and the residual an upper bound; richer transformation families for other architectures
could reapportion the split, though not below the conflict floor. Third,
"unmergeable" here means by aligned linear interpolation of weights — a barrier to that operator does
not preclude every conceivable recombination method (routing, for one, sidesteps it by not blending).
Emergent DobzhanskyMuller incompatibilities in real weights remain the flagship *hypothesis* of this
@ -770,7 +781,10 @@ falsifier, not yet established):
| Outbreeding depression on rugged landscapes; operator design rule | Exact-model result; hypothesis at LLM scale | NK epistasis stands in for skill entanglement | E9E10; directed selection rescues | Not yet mapped onto a real task-entanglement measure |
| Optimal mate-pool breadth shrinks with ruggedness | Exact-model result; hypothesis for merging populations | Ring population, local selection | E14 | Phenomenon known to island-model evolutionary computation; our contribution is the mapping and the diversity/mean decomposition |
| Merge failure decomposes into coordinate artefact + functional residual | Empirical (MLP tier; LLM tier in progress) | Alignment enumerates the architecture's unit symmetries | Full-symmetry residual ≈ 0 (compatible) vs ≈ naive (conflict); cliff in hybrid fitness | Scoped to aligned linear interpolation; conflict floor is information-theoretic, not genetic |
| Epistasis (not divergence) sets the cliff; snowball onset | Exact-model result; **hypothesis** at the neural tier | BDM incompatibility structure | E12 | The decisive pre-merge prediction test is proposed, not run |
| Epistasis (not divergence) sets the cliff; snowball onset | Exact-model result; **hypothesis** at the neural tier | BDM incompatibility structure | E12 | Snowball count ≠ performance cliff without the effect-size link; neural test outstanding |
| Pre-merge functional disagreement predicts merge penalty | Empirical, within a controlled grid (0.5B, 13 conditions × 3 seeds) | Constructed conflict/overlap/duration axes; oracle-potential outcome (pre-registered; ordering sensitive to reference) | Clustered CIs exclude 0; held-out LOCO ρ≈0.4; selected geometry baselines ≈ 0 | Head-to-head predictor differences not individually significant; only selected baselines; generalisation to real task pairs open |
| Confidence weighting improves rank prediction over raw disagreement | **Not supported** (pre-registered internal prediction) | — | Paired Δ\|ρ\| ≈ 0.02, CI [0.13, +0.06] | Weighting does double the conflict-vs-compat level contrast |
| The predictor improves budget-matched operator choice | **Open** | — | Soup-vs-route gap readout noise-dominated at 0.5B | The practical payoff; untested |
| Emergent speciation without conflict | **Not observed** (pre-registered) | Shared ancestry, compatible tasks, tested divergences | E13b: residual 0.000; merge rescues specialists | Bounds the hypothesis; longer horizons/distribution shift/capacity pressure untested |
| Grounding + sex + diversity jointly necessary | Exact-model result; hypothesis at LLM scale | Conformity stands in for self-consumption | E11 four-arm ablation, each arm failing distinctly | The full grounded LLM society is unbuilt |