MachineSex/results/llm_moe_hard_seeds
Giorgio Gilestro a40ace1821 second review round: tempered claims, robust statistics, corrected technical statements
Analyses (figures/stats_llm_epistasis.py, committed + reproducible):
condition-clustered bootstrap CIs (functional measures exclude zero:
dis_raw [+0.04,+0.69], conf-weighted [+0.02,+0.68]; gradient alignment
[-0.59,-0.06]; geometry straddles zero), PAIRED predictor contrasts (not
individually significant — stated), leave-one-condition-out held-out
prediction (functional replicates, geometry ~0, performance baseline
unstable), three outcome references (ordering sensitive to reference —
reported, with the mechanism), between/within-axis decomposition
(within-conflict identification impossible by design; the compat axis
identifies), and seed-level paired reliability (routing/directed beat
soup 3/3 seeds incl. one catastrophic soup failure; CI-width fragility
claim withdrawn).

Renames and corrections: "decisive experiment" -> "controlled predictive
test"; "operational epistasis" -> "confidence-weighted functional
conflict (proposed proxy)"; "functional by construction" -> "controls a
major source of coordinate mismatch / conflict-associated" (module,
configs, READMEs, figures); SI proposition's "chord" defined precisely
(endpoint-loss interpolation, invariant) vs the path (not invariant) +
no-global-optimality caveat (removable = lower bound, residual = upper);
snowball count != performance cliff distinction added; claims table
gains four rows (grid finding / weighting NOT supported / functional-vs-
all-geometry not established / operator choice open); §1 ladder states
the prediction rung as a bounded small-model result.

paper/response-to-review-2.md: point-by-point, opening with the
bookkeeping correction (E13b/c were in the reviewed draft — revised
interpretation, not new results). READMEs rewritten around the four
analyses with the chronology (prospective/adaptive/post-hoc) disclosed.
151 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
2026-09-06 17:55:46 +01:00
..
manifest.json Phase 3: LLM-tier speciation + multi-seed firm-up of the recombination claims 2026-09-06 15:39:15 +01:00
README.md second review round: tempered claims, robust statistics, corrected technical statements 2026-09-06 17:55:46 +01:00
resolved_config.yaml Phase 3: LLM-tier speciation + multi-seed firm-up of the recombination claims 2026-09-06 15:39:15 +01:00

Multi-seed union-vs-fusion, hard benchmark (0.5B, 3 seeds)

Part of the multi-seed firm-up; full legend and per-seed table in results/llm_merge_seeds/README.md (panel B of its llm_seeds.png). Headline: routing beats fusion in every seed (3/3 paired, both metrics), and one seed exhibited a catastrophic soup failure (0.071 overall / 0.000 worst-family) that routing was immune to (0.262/0.225). With 3 seeds the variance contrast (sd 0.090 vs 0.023) is an observation, not an estimate.