Analyses (figures/stats_llm_epistasis.py, committed + reproducible): condition-clustered bootstrap CIs (functional measures exclude zero: dis_raw [+0.04,+0.69], conf-weighted [+0.02,+0.68]; gradient alignment [-0.59,-0.06]; geometry straddles zero), PAIRED predictor contrasts (not individually significant — stated), leave-one-condition-out held-out prediction (functional replicates, geometry ~0, performance baseline unstable), three outcome references (ordering sensitive to reference — reported, with the mechanism), between/within-axis decomposition (within-conflict identification impossible by design; the compat axis identifies), and seed-level paired reliability (routing/directed beat soup 3/3 seeds incl. one catastrophic soup failure; CI-width fragility claim withdrawn). Renames and corrections: "decisive experiment" -> "controlled predictive test"; "operational epistasis" -> "confidence-weighted functional conflict (proposed proxy)"; "functional by construction" -> "controls a major source of coordinate mismatch / conflict-associated" (module, configs, READMEs, figures); SI proposition's "chord" defined precisely (endpoint-loss interpolation, invariant) vs the path (not invariant) + no-global-optimality caveat (removable = lower bound, residual = upper); snowball count != performance cliff distinction added; claims table gains four rows (grid finding / weighting NOT supported / functional-vs- all-geometry not established / operator choice open); §1 ladder states the prediction rung as a bounded small-model result. paper/response-to-review-2.md: point-by-point, opening with the bookkeeping correction (E13b/c were in the reviewed draft — revised interpretation, not new results). READMEs rewritten around the four analyses with the chronology (prospective/adaptive/post-hoc) disclosed. 151 tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
146 lines
10 KiB
Markdown
146 lines
10 KiB
Markdown
# Response to the second review
|
||
|
||
*This response accompanies a further revision. Every number below is reproduced by a committed script
|
||
(`figures/stats_llm_epistasis.py`) from committed artifacts; the revised documents are
|
||
`results/llm_epistasis/README.md` (rewritten around your four analyses), the manuscript, and
|
||
`paper/si-notes.md`.*
|
||
|
||
---
|
||
|
||
## 0. A correction first (your §7)
|
||
|
||
You are right, and we apologise for the bookkeeping error: the draft you reviewed **already
|
||
contained** the full-symmetry alignment experiment, the conflict residual, and the
|
||
compatible-specialisation null. Our previous letter's "new since the review" conflated three things
|
||
that we now state separately: **new results** (the LLM-tier speciation runs, the multi-seed
|
||
replication, and the controlled predictive test with its control axis), **new controls and analyses**
|
||
(the compatible-overlap axis; the robust statistics in this letter), and **revised interpretation**
|
||
(everything about E13b/c, which was experimental content you had already seen and whose *presentation*
|
||
we changed). The experimental delta attributable to the review process is the first and second
|
||
categories only.
|
||
|
||
## 1. Your two-conclusion distinction — adopted as the evidential boundary
|
||
|
||
We accept the boundary exactly as you drew it:
|
||
|
||
- **Demonstrated:** a small-model, controlled predictive test in which pre-merge functional
|
||
disagreement predicted merge penalties where the selected weight-space measures did not.
|
||
- **Not demonstrated:** that an epistasis-specific predictor adds value beyond ordinary functional
|
||
disagreement, or that the prediction improves operator choice.
|
||
|
||
The experiment is now titled a **"controlled predictive test"** everywhere ("decisive experiment
|
||
delivered" is gone), and the manuscript's §1 ladder describes the prediction rung in your conditional
|
||
formulation, with its four boundary clauses stated in place: constructed grid, small scale, refinement
|
||
not superior, operator choice open. The measure itself is renamed **"confidence-weighted functional
|
||
conflict — a proposed proxy for merge-relevant interactions"**; we accept that bilateral confident
|
||
contradiction measures incompatible endpoint behaviour, not non-additive interaction in the strict
|
||
biological sense, and the paper no longer calls any measured quantity "epistasis." Your distinction —
|
||
*the framework motivated the measurement and controls* vs *their success validates the specifically
|
||
population-genetic mechanism* — is adopted verbatim; we claim the former.
|
||
|
||
## 2. The four analyses — run
|
||
|
||
**(1) Direct predictor comparison.** Condition-clustered bootstrap (13 clusters, B = 4000), 95% CIs
|
||
for each predictor's ρ against the pre-registered primary outcome:
|
||
|
||
| predictor | ρ | clustered 95% CI |
|
||
|---|---|---|
|
||
| raw functional disagreement | +0.460 | [+0.04, +0.69] |
|
||
| confidence-weighted functional conflict | +0.446 | [+0.02, +0.68] |
|
||
| gradient alignment | −0.347 | [−0.59, −0.06] |
|
||
| delta L2 | +0.165 | [−0.27, +0.58] |
|
||
| delta cosine | +0.030 | [−0.46, +0.51] |
|
||
| cross-family accuracy | −0.005 | [−0.29, +0.31] |
|
||
|
||
**Paired contrasts are not individually significant** (e.g. |ρ(dis_raw)| − |ρ(delta_cos)| = +0.23,
|
||
CI [−0.23, +0.59]). Held-out prediction (leave-one-condition-out linear fits): functional measures
|
||
replicate (dis_raw ρ = +0.396, p = 0.013; conf-weighted +0.352, p = 0.028); geometry ≈ 0; the
|
||
performance baseline is unstable out-of-sample (−0.435). So the supported statement — now the
|
||
conclusion in the README, the figure title, and the manuscript — is yours: *across this controlled
|
||
grid, functional disagreement showed a detectable, held-out-robust association with merge penalty;
|
||
LoRA-delta cosine and L2 showed no statistically detectable association; gradient alignment carried
|
||
intermediate signal (its CI excludes zero), so this is not a clean functional-versus-all-geometric
|
||
divide; head-to-head predictor differences are not individually significant; only these baselines were
|
||
tested.* "Weight divergence does not predict merge failure" has been removed as over-broad.
|
||
|
||
**(2) Sample structure.** 39 rows = **13 conditions × 3 seeds**. Parents are retrained per
|
||
condition × seed, but children share task-data seeds across conditions within a seed (e.g. the
|
||
duration-3 parent and the conflict-0 parent are trained on essentially the same data), so rows are not
|
||
independent — hence clustering by condition in all uncertainty estimates, disclosed in the README.
|
||
|
||
**(3) Between- vs within-axis.** You were right that the pooled correlation is substantially axis
|
||
discrimination (mean penalty: conflict 0.061 vs compat 0.005 / duration 0.011). Within the conflict
|
||
axis (n = 15): functional measures +0.59/+0.62 — but delta-L2 is +0.71 there, because *within that
|
||
axis* conflict fraction, added-data volume, and delta growth are collinear: **within-axis
|
||
identification is impossible by design**, and we now say so; the identification comes from the control
|
||
axes, where the same volumes and L2 ranges occur at ~zero penalty. Scatterplots are coloured by axis
|
||
in the figure. Predicting penalties for a held-out *conflict mechanism* (not just held-out conditions
|
||
of the same mechanism) is listed as the next test — we agree it is the more valuable one.
|
||
|
||
**(4) Outcome references.** Reported under all three: oracle parent potential (pre-registered
|
||
primary), best parent, and mean parent. The ordering is **sensitive to the reference** — under the
|
||
best-parent reference, delta-L2 correlates comparably to the functional measures (+0.48 vs +0.34).
|
||
Our reading, stated in the README rather than hidden: that reference inherits parent-strength trends
|
||
that track training volume, which geometry also tracks, coupling predictor and outcome through the
|
||
reference rather than through merge damage. On structural coupling between the disagreement predictor
|
||
and the oracle-potential outcome: probe and test sets are disjoint by construction, but we agree
|
||
definitional coupling through parental complementarity cannot be fully excluded, which is one more
|
||
reason all three references are now on the table.
|
||
|
||
## 3. Hypothesis labels
|
||
|
||
Your table is adopted essentially as written — the manuscript's claims-at-a-glance table now carries:
|
||
pre-merge disagreement predicts penalty (*empirical within the controlled grid*, with the boundary
|
||
clauses in the limits column); confidence weighting improves rank prediction (***not supported***);
|
||
functional beats all geometry (*not established — selected baselines only*); operator choice (*open*);
|
||
cliff, snowball, emergent DMIs (*hypotheses*, unchanged). We also added your snowball distinction
|
||
where the snowball is discussed: super-linear growth in incompatibility *count* does not by itself
|
||
entail a sharp *performance* cliff — that needs the count→effect-size→performance link, which the
|
||
analytic model supplies under its assumptions and any neural test must establish separately.
|
||
|
||
## 4. The two technical statements — corrected
|
||
|
||
**"Endpoints and chord are invariant."** Accepted; the ambiguity was ours. The SI proposition now
|
||
defines "chord" precisely as the α-linear interpolation **of the endpoint loss values** — the barrier
|
||
baseline, a function of endpoints only, which *is* invariant — and states explicitly that the
|
||
**weight-space interpolation path is generally not invariant** (that being precisely why alignment can
|
||
lower a barrier). It also now carries your second point: exact recovery of a permuted-and-rescaled
|
||
copy validates a special case and does not establish global optimality of the alignment for
|
||
independently trained networks — so the "removable" share is a lower bound and the "residual" an upper
|
||
bound, stated wherever the decomposition is used.
|
||
|
||
**"Frozen base pins the coordinate system, so failure is functional by construction."** Accepted, and
|
||
your replacement wording adopted verbatim across the module, configs, READMEs, and figure: *the shared
|
||
frozen base controls a major source of coordinate mismatch, allowing a cleaner test of
|
||
conflict-associated merging failure* — with the boundary stated (failures of delta-averaging can still
|
||
reflect nonlinear interaction, scaling, or capacity). "Conflict-associated," not "functional by
|
||
construction," throughout.
|
||
|
||
## 5. Chronology and reliability
|
||
|
||
**Chronology (now a section of the README).** Prospective: hypotheses, predictors, primary outcome and
|
||
falsifiers were in the config before the first grid ran. Adaptive: the compatible-overlap control was
|
||
added *after* geometry appeared to win, with its own pre-stated readings, run on the same seeds; no
|
||
existing rows were re-run or altered. Post hoc: the clustered-bootstrap/LOCO/multi-reference analyses
|
||
were added at your request after all data was collected. We agree this is transparent adaptive
|
||
experimentation, not wholly prospective confirmation, and it is labelled as such.
|
||
|
||
**Reliability.** The CI-width claim is withdrawn. The seed-level statement now reads: routing beat the
|
||
soup **in every seed** (3/3 paired, both metrics), directed selection beat the soup 3/3, and one seed
|
||
exhibited a catastrophic soup failure (0.071 overall, 0.000 worst-family) to which routing was immune
|
||
(0.262/0.225); seed-level sds (0.090 vs 0.023) are reported as an observation, with the explicit note
|
||
that three seeds do not support a variance estimate.
|
||
|
||
## 6. On your bottom line
|
||
|
||
We accept your formulation as the paper's claim for this line of work — it now closes the relevant
|
||
results section nearly verbatim: *we separated overlap, divergence, and conflict experimentally; in
|
||
this controlled setting, functional disagreement predicted merging damage when simple weight-distance
|
||
measures did not; the proposed epistasis refinement and the emergent-speciation mechanism remain
|
||
unconfirmed.* And we take the redirection about what would count next: not 0.5B → 7B alone, but
|
||
**generalisation to unfamiliar conflict structures** (a held-out conflict mechanism, and real rather
|
||
than constructed task pairs) and **a demonstrably better budget-matched merging decision**. Those two
|
||
now head the open-problems list, above the scale replication.
|
||
|
||
We would welcome a third pass if you have the appetite — particularly on whether the README's
|
||
conditional conclusion and the chronology section read at the right strength.
|