# The decisive experiment — predicting merge failure BEFORE merging The external review's bar (2026-08-11): population-genetic quantities must *predict*, not re-describe — forecast merge success **pre-merge** and beat existing predictors. Design: 39 parent pairs (0.5B LoRA children of one frozen base, 3 seeds) on **three axes decorrelated by construction** — `conflict` (contradictory conventions on shared ambiguous prompts, private budgets fixed), `compat` (the control: **same shared prompts, same convention** — task overlap *without* conflict; added after the first grid exposed a confound, see below), and `duration` (weight divergence with zero conflict, 1→12 epochs). Primary outcome (pre-registered): **merge penalty** = parent potential − merged achieved (the hybrid-load analogue). Figure: `llm_epistasis.png`. ### The league table (Spearman ρ vs merge penalty, full three-axis pool, n = 39) | pre-merge predictor | ρ | p | reading | |---|---|---|---| | raw functional disagreement (`dis_raw`) | **+0.460** | 0.003 | predicts | | operational epistasis (`epi_conf`, confidence-weighted) | **+0.446** | 0.004 | predicts | | gradient alignment at the base (cf. 2601.22285) | −0.347 | 0.03 | weakly informative | | LoRA-delta L2 distance (geometry) | +0.165 | 0.32 | uninformative | | LoRA-delta cosine (geometry) | +0.030 | 0.86 | uninformative | | cross-family accuracy (performance) | −0.005 | 0.98 | uninformative | **Headline: functional conflict, measured before merging, predicts merge failure; weight geometry does not.** The duration axis spans the same weight-divergence range as the conflict axis (L2 ≈ 2.4–4.0) at ~zero penalty, and the compat axis adds the same *data volumes* and overlap at ~zero penalty — so both geometric predictors collapse once overlap and volume are controlled. ### The control that did the work (`llm_epistasis_compat/`) In the first grid (conflict + duration only), `delta_cos` scored ρ = +0.60 — apparently the best predictor. That was an **artifact**: every shared-data pair in that pool was a conflicted pair, so geometry could win as a mere task-overlap/volume detector. The `compat` axis (overlap without conflict) exposes it: penalty ≈ 0.005 there, and the geometry correlations collapse (+0.60 → +0.03). The functional measures behave correctly on the control — parents trained on the same convention *agree* on the shared prompts (epi_conf: 0.46 conflict vs 0.23 compat, a 2× contrast; raw disagreement 0.72 vs 0.47, only 1.5× — the confidence weighting removes complementation noise from the *level*, giving the cleaner axis separation). ### Honest riders (pre-registered falsifier status) 1. The internal prediction that confidence weighting would beat raw disagreement **as a rank predictor is not confirmed**: `epi_conf` and `dis_raw` are statistically indistinguishable at n = 39 (the weighting does improve the conflict-vs-compat *contrast* in levels). The paper reports the functional-vs-geometric verdict, not a win for the refinement. 2. Correlations are moderate (|ρ| ≈ 0.45), bounded by 0.5B merge-outcome noise (soup merges carry large intrinsic seed variance — see `llm_moe_hard_seeds`); read as signs and ordering, not magnitudes. 7B replication is the natural firm-up. 3. Gradient alignment carries real signal (it differentiates conflicting conventions at the base) but less than the functional measures in this design. **Bottom line for the paper:** the framework's claim — *epistasis (functional conflict), not divergence, sets merge compatibility* — survives its designed falsification test at this tier: the operational conflict measures predict, the divergence measures do not, and the case was made honest by a control that first *broke our own experiment's* favourite-looking geometric predictor.