MachineSex/results/llm_moe_hard_hpc/README.md
Giorgio Gilestro 84124de143 Manuscript revision and pending experiment work, snapshot before restructuring
Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 16:54:09 +01:00

5.4 KiB
Raw Permalink Blame History

llm_moe_hard_hpc — union vs fusion on HARD (unsaturated) tasks at 7B: the flip was a saturation artefact

Claim tested. llm_moe_hpc found fusion beating union at 7B (soup 0.87 > routing 0.84) and read it as "a capable base lets averaging compose." But the easy families were saturated (strings & arith at 1.00), so that flip could have been an artefact of no headroom rather than base capability. This run re-runs the union-vs-fusion contrast on the hard task variant (multi-step lists, Caesar ciphers / letter transforms, multi-step & larger arithmetic), where 7B is not saturated. One L40S (46 GB) GPU, Imperial CX3.

Setup. Base Qwen2.5-7B-Instruct, hard: true, fresh hard specialists (cached spec_*_hard), 200 test tasks/family, seed 1. Same five operators as llm_moe_hpc.

Results (accuracy — nothing saturated; arith ≈ 0.48, strings 0.67, lists 0.34)

operator lists strings arith overall worst-family router
best specialist (strings) 0.155 0.665 0.455 0.425 0.155
fuse: soup 0.390 0.300 0.485 0.392 0.300
fuse: ties 0.385 0.330 0.485 0.400 0.330
route: oracle 0.335 0.670 0.495 0.500 0.335 1.00
route: learned 0.335 0.670 0.495 0.500 0.335 1.00
max-merge 0.215 0.195 0.480 0.297 0.195

The finding: the 7B "fusion wins" flip was saturation, not capability

  • Union beats fusion again — decisively. Routing 0.500 > fusion 0.40 (soup 0.392 / ties 0.400), a 10-point margin, the same ordering as 0.5B. The easy-task flip (fusion > union at 7B) does not survive once the tasks are hard enough to leave headroom.
  • Fusion dilutes so badly it loses to the best single specialist. On hard tasks the strings skill (Caesar ciphers etc.) is fragile: the strings-specialist scores 0.665, but soup washes it out to 0.300 — so soup (0.392 overall) even trails the best single specialist (0.425). Routing keeps the specialist intact (strings 0.670) and wins. Dilution is severe exactly when the specialist's contribution is hard-won.
  • So "merge, don't average" is a HEADROOM law, not a base-size law. Union > fusion whenever there is room to lose to dilution — a weak base (0.5B) or hard tasks at a strong base (7B-hard). Fusion only wins in the degenerate corner where the tasks are so easy the strong base composes to the 1.00 ceiling (7B-easy). This corrects the llm_moe_hpc interpretation: base capability was a confound; the operative variable is task headroom.
  • Riders unchanged. Learned router still perfect (1.00, lexical families); max_merge still the weakest union (0.297, not input-adaptive).

Takeaway

On genuinely hard, unsaturated tasks, the union operator (routing) beats fusion at 7B by the same margin it does at 0.5B — E8's max > mean in real weights, robust across scale once you control for saturation. The llm_moe_hpc flip is re-read as a saturation artefact. Falsifier (not triggered): fusion matching/beating routing on hard tasks — instead fusion diluted below even the best specialist. Provenance in manifest.json (hard: true, L40S, torch 2.12.1 / transformers 5.13.0 / peft 0.19.1).

Seeds 13 (2026-09-11)

Seeds 23 were run on CX3 via hpc/llm_7b_seeds.pbs (seed 1 above was moved to s1/; the bundle layout is now s{seed}/). Fixed test sets, training seed varied. Per-seed values and mean ± 95% CI from figures/stats_llm_7b_seeds.py:

model       metric  n_seeds    s1    s2    s3  mean  ci95
best_specialist      overall        3 0.425 0.407 0.390 0.407 0.020
best_specialist worst_family        3 0.155 0.150 0.185 0.163 0.021
     merge_soup      overall        3 0.392 0.405 0.428 0.408 0.021
     merge_soup worst_family        3 0.300 0.345 0.340 0.328 0.028
     merge_ties      overall        3 0.400 0.427 0.438 0.422 0.022
     merge_ties worst_family        3 0.330 0.385 0.380 0.365 0.034
     moe_oracle      overall        3 0.500 0.498 0.510 0.503 0.007
     moe_oracle worst_family        3 0.335 0.360 0.455 0.383 0.072
    moe_learned      overall        3 0.500 0.498 0.510 0.503 0.007
    moe_learned worst_family        3 0.335 0.360 0.455 0.383 0.072
      max_merge      overall        3 0.297 0.342 0.403 0.347 0.061
      max_merge worst_family        3 0.195 0.175 0.240 0.203 0.038

                    contrast       metric  n_seeds     s1     s2    s3  mean  ci95 sign_agrees
     moe_oracle  merge_soup      overall        3  0.108  0.093 0.082 0.094 0.015         3/3
     moe_oracle  merge_soup worst_family        3  0.035  0.015 0.115 0.055 0.060         3/3
    moe_learned  merge_soup      overall        3  0.108  0.093 0.082 0.094 0.015         3/3
    moe_learned  merge_soup worst_family        3  0.035  0.015 0.115 0.055 0.060         3/3
merge_soup  best_specialist      overall        3 -0.033 -0.002 0.038 0.001 0.041         1/3
merge_soup  best_specialist worst_family        3  0.145  0.195 0.155 0.165 0.030         3/3

Reading: routing beats the weight-average in every seed (+0.094 ± 0.015 overall). The seed-1 observation that the soup fell below the best single specialist did not replicate (soup best specialist overall 0.033, 0.002, +0.038; mean +0.001): over three seeds the soup matches the best parent overall and beats it on worst-family (+0.165, 3/3). max_merge remains the weakest union (0.347).