MachineSex/results/llm_merge_seeds/README.md
Giorgio Gilestro 5a23ddaf2a Phase 3: LLM-tier speciation + multi-seed firm-up of the recombination claims
llm_speciation (new kind; src/llm/speciation.py): E13 in LLM weights.
LoRA children share the frozen base's coordinates, so merge failure is
functional by construction. CONFLICT (ambiguous sort prompts learned
under opposite conventions — the BDM structure): function-specific
hybrid breakdown — merged coherence 0.02-0.08 falls below BOTH parents
(~0.2) on the conflicted function; and in the de-confounded `add` design
(private budget fixed, conflict added on top; 3 seeds after a
single-seed pilot showed one anomalous point) the merge's private-family
accuracy shows NO trend with conflict — the damage is surgical, not
global. DURATION (over-trained disjoint specialists, 1->12 epochs): the
merge improves (0.84->0.94) and stays above the best parent — the MLP
"no emergent isolation" null generalises; relevant to the
expert-training-duration report (2607.11997), with the epistasis
prediction left to the decisive experiment.

Multi-seed firm-up (seeds threaded into specialist caches; `seeds:` list
support in the runner; fixed test sets): all three recombination claims
hold with CIs — merges beat every specialist (5 seeds, ties
0.647±0.027 > best spec 0.592±0.009; worst-family 0.28 vs <=0.16); union
0.274±0.026 > fusion 0.174±0.102 on hard (3 seeds); directed 0.221±0.026
> soup. NEW finding: fusion is seed-FRAGILE where headroom exists
(CI ±0.10) while routing/directed selection are stable (±0.026) — the
union/selection operators win on reliability, not just mean.

Figures (llm_speciation 3-panel; llm_seeds 3-panel with 95% CI), READMEs,
+1 convention test (150 green), make llm-speciation / llm-seeds targets.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
2026-09-06 15:39:15 +01:00

35 lines
2.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Multi-seed LLM recombination (0.5B) — the claims with error bars
PNAS work-order Phase 3: removes the "one seed" objection on the three LLM recombination claims.
Protocol: **test sets fixed** (seed 1000+i per family), **training seed varied** (specialists cache
per-seed as `spec_<family>[_hard]_s<seed>`), so across-seed variance is training variance only.
Figure: `llm_seeds.png` (this dir) aggregates all three experiments, 95% CI over seeds.
### (A) FisherMuller, easy benchmark, 5 seeds (`llm_merge_seeds`)
| model | overall | worst-family |
|---|---|---|
| merge_ties | **0.647 ± 0.027** | **0.282 ± 0.020** |
| merge_soup | 0.632 ± 0.042 | 0.278 ± 0.028 |
| best specialist (strings) | 0.592 ± 0.009 | 0.078 ± 0.011 |
| base | 0.277 | 0.150 |
Both merges beat every specialist overall (ties: non-overlapping CIs; soup: marginal at 0.5B, as in
the single-seed run — decisive at 7B) and the **worst-family signature is unambiguous**: merges ≈0.28
vs ≤0.16 for any parent — only recombined models are competent everywhere.
### (B) Union vs fusion, hard benchmark, 3 seeds (`llm_moe_hard_seeds/`)
Routing (union) 0.274 ± 0.026 overall / 0.238 ± 0.024 worst-family; fusion soup 0.174 ± 0.102 / 0.088
± 0.093; ties similar; best specialist 0.199 ± 0.026. Union beats fusion on both metrics — **and a new
finding: fusion is seed-FRAGILE on hard tasks (CI ±0.10) while routing is seed-stable (±0.026).**
Averaging's outcome depends on which specialist minima the seeds happened to find; selection-based
recombination is reliable. (Learned router still = oracle: lexically distinct families, known rider.)
### (C) Directed offspring selection, hard, 3 seeds (`llm_directed_hard_seeds/`)
directed_overall 0.221 ± 0.026 (> soup 0.174 ± 0.102 and > best specialist); directed_balanced
worst-family 0.158 ± 0.036 (> soup 0.088 ± 0.093). Directed selection both beats and **stabilises**
the a-priori soup; per-input routing (B) remains above any single global blend, as before.
**Read together:** all three recombination claims hold under seed replication, and the operator
ordering (route > directed-select > soup, on headroom tasks) is not only a mean effect but a
*variance* effect — the union/selection operators are the reliable ones. Base: Qwen2.5-0.5B-Instruct;
statistical (per-seed) reproducibility per blueprint §4.