Clarity pass over the main text (36-item audit), Discussion rewrite and cut, acknowledgements, Souly et al. as ref 62, lettered SI panels, model section moved under Results; plus the untracked curriculum/society/compose/smol configs, runners, figures, stats and tests that the SI already cites. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
76 lines
5.4 KiB
Markdown
76 lines
5.4 KiB
Markdown
# llm_moe_hard_hpc — union vs fusion on HARD (unsaturated) tasks at 7B: the flip was a saturation artefact
|
||
|
||
**Claim tested.** `llm_moe_hpc` found fusion beating union at 7B (soup 0.87 > routing 0.84) and read it
|
||
as "a capable base lets averaging compose." But the easy families were *saturated* (strings & arith at
|
||
1.00), so that flip could have been an artefact of no headroom rather than base capability. This run
|
||
re-runs the union-vs-fusion contrast on the **hard task variant** (multi-step lists, Caesar ciphers /
|
||
letter transforms, multi-step & larger arithmetic), where 7B is **not** saturated. One **L40S (46 GB)**
|
||
GPU, Imperial CX3.
|
||
|
||
**Setup.** Base **Qwen2.5-7B-Instruct**, `hard: true`, fresh hard specialists (cached `spec_*_hard`),
|
||
200 test tasks/family, seed 1. Same five operators as `llm_moe_hpc`.
|
||
|
||
### Results (accuracy — nothing saturated; arith ≈ 0.48, strings 0.67, lists 0.34)
|
||
| operator | lists | strings | arith | overall | worst-family | router |
|
||
|---|---|---|---|---|---|---|
|
||
| best specialist (strings) | 0.155 | 0.665 | 0.455 | 0.425 | 0.155 | — |
|
||
| fuse: soup | 0.390 | 0.300 | 0.485 | 0.392 | 0.300 | — |
|
||
| fuse: ties | 0.385 | 0.330 | 0.485 | 0.400 | 0.330 | — |
|
||
| **route: oracle** | 0.335 | **0.670** | 0.495 | **0.500** | 0.335 | 1.00 |
|
||
| **route: learned** | 0.335 | 0.670 | 0.495 | **0.500** | 0.335 | 1.00 |
|
||
| max-merge | 0.215 | 0.195 | 0.480 | 0.297 | 0.195 | — |
|
||
|
||
### The finding: the 7B "fusion wins" flip was saturation, not capability
|
||
- **Union beats fusion again — decisively.** Routing **0.500 > fusion 0.40 (soup 0.392 / ties 0.400)**,
|
||
a 10-point margin, the *same* ordering as 0.5B. The easy-task flip (fusion > union at 7B) does **not**
|
||
survive once the tasks are hard enough to leave headroom.
|
||
- **Fusion dilutes so badly it loses to the best single specialist.** On hard tasks the strings skill
|
||
(Caesar ciphers etc.) is fragile: the strings-specialist scores **0.665**, but soup washes it out to
|
||
**0.300** — so soup (0.392 overall) even trails the best *single* specialist (0.425). Routing keeps
|
||
the specialist intact (strings 0.670) and wins. Dilution is severe exactly when the specialist's
|
||
contribution is hard-won.
|
||
- **So "merge, don't average" is a HEADROOM law, not a base-size law.** Union > fusion whenever there
|
||
is room to lose to dilution — a weak base (0.5B) *or* hard tasks at a strong base (7B-hard). Fusion
|
||
only wins in the degenerate corner where the tasks are so easy the strong base composes to the 1.00
|
||
ceiling (7B-easy). This corrects the `llm_moe_hpc` interpretation: base capability was a confound;
|
||
the operative variable is task headroom.
|
||
- **Riders unchanged.** Learned router still perfect (1.00, lexical families); `max_merge` still the
|
||
weakest union (0.297, not input-adaptive).
|
||
|
||
### Takeaway
|
||
On genuinely hard, unsaturated tasks, the union operator (routing) beats fusion at 7B by the same
|
||
margin it does at 0.5B — E8's `max > mean` in real weights, robust across scale once you control for
|
||
saturation. The `llm_moe_hpc` flip is re-read as a saturation artefact. **Falsifier (not triggered):**
|
||
fusion matching/beating routing on hard tasks — instead fusion diluted below even the best specialist.
|
||
Provenance in `manifest.json` (`hard: true`, L40S, torch 2.12.1 / transformers 5.13.0 / peft 0.19.1).
|
||
|
||
## Seeds 1–3 (2026-09-11)
|
||
|
||
Seeds 2–3 were run on CX3 via `hpc/llm_7b_seeds.pbs` (seed 1 above was moved to `s1/`; the bundle
|
||
layout is now `s{seed}/`). Fixed test sets, training seed varied. Per-seed values and mean ± 95% CI
|
||
from `figures/stats_llm_7b_seeds.py`:
|
||
```
|
||
model metric n_seeds s1 s2 s3 mean ci95
|
||
best_specialist overall 3 0.425 0.407 0.390 0.407 0.020
|
||
best_specialist worst_family 3 0.155 0.150 0.185 0.163 0.021
|
||
merge_soup overall 3 0.392 0.405 0.428 0.408 0.021
|
||
merge_soup worst_family 3 0.300 0.345 0.340 0.328 0.028
|
||
merge_ties overall 3 0.400 0.427 0.438 0.422 0.022
|
||
merge_ties worst_family 3 0.330 0.385 0.380 0.365 0.034
|
||
moe_oracle overall 3 0.500 0.498 0.510 0.503 0.007
|
||
moe_oracle worst_family 3 0.335 0.360 0.455 0.383 0.072
|
||
moe_learned overall 3 0.500 0.498 0.510 0.503 0.007
|
||
moe_learned worst_family 3 0.335 0.360 0.455 0.383 0.072
|
||
max_merge overall 3 0.297 0.342 0.403 0.347 0.061
|
||
max_merge worst_family 3 0.195 0.175 0.240 0.203 0.038
|
||
|
||
contrast metric n_seeds s1 s2 s3 mean ci95 sign_agrees
|
||
moe_oracle − merge_soup overall 3 0.108 0.093 0.082 0.094 0.015 3/3
|
||
moe_oracle − merge_soup worst_family 3 0.035 0.015 0.115 0.055 0.060 3/3
|
||
moe_learned − merge_soup overall 3 0.108 0.093 0.082 0.094 0.015 3/3
|
||
moe_learned − merge_soup worst_family 3 0.035 0.015 0.115 0.055 0.060 3/3
|
||
merge_soup − best_specialist overall 3 -0.033 -0.002 0.038 0.001 0.041 1/3
|
||
merge_soup − best_specialist worst_family 3 0.145 0.195 0.155 0.165 0.030 3/3
|
||
```
|
||
|
||
Reading: routing beats the weight-average in every seed (+0.094 ± 0.015 overall). The seed-1 observation that the soup fell *below* the best single specialist did not replicate (soup − best specialist overall −0.033, −0.002, +0.038; mean +0.001): over three seeds the soup matches the best parent overall and beats it on worst-family (+0.165, 3/3). `max_merge` remains the weakest union (0.347).
|