MachineSex/results/llm_moe_hard_hpc/README.md
Giorgio Gilestro 84124de143 Manuscript revision and pending experiment work, snapshot before restructuring
Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 16:54:09 +01:00

76 lines
5.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# llm_moe_hard_hpc — union vs fusion on HARD (unsaturated) tasks at 7B: the flip was a saturation artefact
**Claim tested.** `llm_moe_hpc` found fusion beating union at 7B (soup 0.87 > routing 0.84) and read it
as "a capable base lets averaging compose." But the easy families were *saturated* (strings & arith at
1.00), so that flip could have been an artefact of no headroom rather than base capability. This run
re-runs the union-vs-fusion contrast on the **hard task variant** (multi-step lists, Caesar ciphers /
letter transforms, multi-step & larger arithmetic), where 7B is **not** saturated. One **L40S (46 GB)**
GPU, Imperial CX3.
**Setup.** Base **Qwen2.5-7B-Instruct**, `hard: true`, fresh hard specialists (cached `spec_*_hard`),
200 test tasks/family, seed 1. Same five operators as `llm_moe_hpc`.
### Results (accuracy — nothing saturated; arith ≈ 0.48, strings 0.67, lists 0.34)
| operator | lists | strings | arith | overall | worst-family | router |
|---|---|---|---|---|---|---|
| best specialist (strings) | 0.155 | 0.665 | 0.455 | 0.425 | 0.155 | — |
| fuse: soup | 0.390 | 0.300 | 0.485 | 0.392 | 0.300 | — |
| fuse: ties | 0.385 | 0.330 | 0.485 | 0.400 | 0.330 | — |
| **route: oracle** | 0.335 | **0.670** | 0.495 | **0.500** | 0.335 | 1.00 |
| **route: learned** | 0.335 | 0.670 | 0.495 | **0.500** | 0.335 | 1.00 |
| max-merge | 0.215 | 0.195 | 0.480 | 0.297 | 0.195 | — |
### The finding: the 7B "fusion wins" flip was saturation, not capability
- **Union beats fusion again — decisively.** Routing **0.500 > fusion 0.40 (soup 0.392 / ties 0.400)**,
a 10-point margin, the *same* ordering as 0.5B. The easy-task flip (fusion > union at 7B) does **not**
survive once the tasks are hard enough to leave headroom.
- **Fusion dilutes so badly it loses to the best single specialist.** On hard tasks the strings skill
(Caesar ciphers etc.) is fragile: the strings-specialist scores **0.665**, but soup washes it out to
**0.300** — so soup (0.392 overall) even trails the best *single* specialist (0.425). Routing keeps
the specialist intact (strings 0.670) and wins. Dilution is severe exactly when the specialist's
contribution is hard-won.
- **So "merge, don't average" is a HEADROOM law, not a base-size law.** Union > fusion whenever there
is room to lose to dilution — a weak base (0.5B) *or* hard tasks at a strong base (7B-hard). Fusion
only wins in the degenerate corner where the tasks are so easy the strong base composes to the 1.00
ceiling (7B-easy). This corrects the `llm_moe_hpc` interpretation: base capability was a confound;
the operative variable is task headroom.
- **Riders unchanged.** Learned router still perfect (1.00, lexical families); `max_merge` still the
weakest union (0.297, not input-adaptive).
### Takeaway
On genuinely hard, unsaturated tasks, the union operator (routing) beats fusion at 7B by the same
margin it does at 0.5B — E8's `max > mean` in real weights, robust across scale once you control for
saturation. The `llm_moe_hpc` flip is re-read as a saturation artefact. **Falsifier (not triggered):**
fusion matching/beating routing on hard tasks — instead fusion diluted below even the best specialist.
Provenance in `manifest.json` (`hard: true`, L40S, torch 2.12.1 / transformers 5.13.0 / peft 0.19.1).
## Seeds 13 (2026-09-11)
Seeds 23 were run on CX3 via `hpc/llm_7b_seeds.pbs` (seed 1 above was moved to `s1/`; the bundle
layout is now `s{seed}/`). Fixed test sets, training seed varied. Per-seed values and mean ± 95% CI
from `figures/stats_llm_7b_seeds.py`:
```
model metric n_seeds s1 s2 s3 mean ci95
best_specialist overall 3 0.425 0.407 0.390 0.407 0.020
best_specialist worst_family 3 0.155 0.150 0.185 0.163 0.021
merge_soup overall 3 0.392 0.405 0.428 0.408 0.021
merge_soup worst_family 3 0.300 0.345 0.340 0.328 0.028
merge_ties overall 3 0.400 0.427 0.438 0.422 0.022
merge_ties worst_family 3 0.330 0.385 0.380 0.365 0.034
moe_oracle overall 3 0.500 0.498 0.510 0.503 0.007
moe_oracle worst_family 3 0.335 0.360 0.455 0.383 0.072
moe_learned overall 3 0.500 0.498 0.510 0.503 0.007
moe_learned worst_family 3 0.335 0.360 0.455 0.383 0.072
max_merge overall 3 0.297 0.342 0.403 0.347 0.061
max_merge worst_family 3 0.195 0.175 0.240 0.203 0.038
contrast metric n_seeds s1 s2 s3 mean ci95 sign_agrees
moe_oracle merge_soup overall 3 0.108 0.093 0.082 0.094 0.015 3/3
moe_oracle merge_soup worst_family 3 0.035 0.015 0.115 0.055 0.060 3/3
moe_learned merge_soup overall 3 0.108 0.093 0.082 0.094 0.015 3/3
moe_learned merge_soup worst_family 3 0.035 0.015 0.115 0.055 0.060 3/3
merge_soup best_specialist overall 3 -0.033 -0.002 0.038 0.001 0.041 1/3
merge_soup best_specialist worst_family 3 0.145 0.195 0.155 0.165 0.030 3/3
```
Reading: routing beats the weight-average in every seed (+0.094 ± 0.015 overall). The seed-1 observation that the soup fell *below* the best single specialist did not replicate (soup best specialist overall 0.033, 0.002, +0.038; mean +0.001): over three seeds the soup matches the best parent overall and beats it on worst-family (+0.165, 3/3). `max_merge` remains the weakest union (0.347).