MachineSex/results/llm_moe_hard_hpc
Giorgio Gilestro 39f6c9f4df hard benchmark: the 7B "fusion wins / no headroom" results were saturation artefacts
The easy task families saturated 7B (strings & arith at 1.00), so the earlier
7B nulls — moe: fusion 0.87 > union 0.84; directed ~= soup — could not separate
"refinements don't help at scale" from "tasks too easy at 7B". Adds a hard task
variant (hard: true in tasks.py: multi-step lists, Caesar ciphers / letter
transforms, multi-step & larger arithmetic; same family labels and answer
formats, threaded through make_tasks/train_specialist/runners; hard specialists
cache separately as spec_*_hard) and re-runs both experiments at 7B on Imperial
CX3 (one L40S, 24 min, unsaturated: arith ~0.48, strings 0.67, lists 0.34).

Both nulls flip back to the 0.5B ordering:
- Union beats fusion again: routing 0.500 > fusion 0.40 (soup 0.392 / ties
  0.400), the same 10-pt margin as 0.5B. Fusion dilutes the fragile strings
  specialist so hard (0.665 -> soup 0.300) that soup even trails the best single
  specialist (0.425); routing keeps it intact (0.670).
- Directed selection beats soup again: 0.492 > 0.392 (+10 pts), recovering most
  of routing's benefit from one deployable merged model (lifts strings to 0.630).

Correction to the earlier interpretation: the llm_moe_hpc "regime flip" and the
llm_directed_hpc "no headroom" null were driven by TASK SATURATION, not base
capability. The operative variable is headroom — "merge, don't average" (union >
fusion) and "directed sex" (selection > single blend) hold whenever there is room
to lose to dilution: a weak base (0.5B) OR hard tasks at a strong base (7B-hard).
Fusion only wins in the degenerate corner where easy tasks let a strong base
compose to the 1.00 ceiling. Vindicates E8's max > mean in real 7B weights once
saturation is controlled.

Default (easy) task behaviour is unchanged (hard defaults False). +1 hard-task
test (131 green). Excludes the 0.5B smoke bundle (a pipeline gate, not a
deliverable). Results in results/llm_{moe,directed}_hard_hpc/ (parquet gitignored).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 19:13:26 +01:00
..
llm_moe.pdf hard benchmark: the 7B "fusion wins / no headroom" results were saturation artefacts 2026-07-05 19:13:26 +01:00
llm_moe.png hard benchmark: the 7B "fusion wins / no headroom" results were saturation artefacts 2026-07-05 19:13:26 +01:00
manifest.json hard benchmark: the 7B "fusion wins / no headroom" results were saturation artefacts 2026-07-05 19:13:26 +01:00
README.md hard benchmark: the 7B "fusion wins / no headroom" results were saturation artefacts 2026-07-05 19:13:26 +01:00
resolved_config.yaml hard benchmark: the 7B "fusion wins / no headroom" results were saturation artefacts 2026-07-05 19:13:26 +01:00

llm_moe_hard_hpc — union vs fusion on HARD (unsaturated) tasks at 7B: the flip was a saturation artefact

Claim tested. llm_moe_hpc found fusion beating union at 7B (soup 0.87 > routing 0.84) and read it as "a capable base lets averaging compose." But the easy families were saturated (strings & arith at 1.00), so that flip could have been an artefact of no headroom rather than base capability. This run re-runs the union-vs-fusion contrast on the hard task variant (multi-step lists, Caesar ciphers / letter transforms, multi-step & larger arithmetic), where 7B is not saturated. One L40S (46 GB) GPU, Imperial CX3.

Setup. Base Qwen2.5-7B-Instruct, hard: true, fresh hard specialists (cached spec_*_hard), 200 test tasks/family, seed 1. Same five operators as llm_moe_hpc.

Results (accuracy — nothing saturated; arith ≈ 0.48, strings 0.67, lists 0.34)

operator lists strings arith overall worst-family router
best specialist (strings) 0.155 0.665 0.455 0.425 0.155
fuse: soup 0.390 0.300 0.485 0.392 0.300
fuse: ties 0.385 0.330 0.485 0.400 0.330
route: oracle 0.335 0.670 0.495 0.500 0.335 1.00
route: learned 0.335 0.670 0.495 0.500 0.335 1.00
max-merge 0.215 0.195 0.480 0.297 0.195

The finding: the 7B "fusion wins" flip was saturation, not capability

  • Union beats fusion again — decisively. Routing 0.500 > fusion 0.40 (soup 0.392 / ties 0.400), a 10-point margin, the same ordering as 0.5B. The easy-task flip (fusion > union at 7B) does not survive once the tasks are hard enough to leave headroom.
  • Fusion dilutes so badly it loses to the best single specialist. On hard tasks the strings skill (Caesar ciphers etc.) is fragile: the strings-specialist scores 0.665, but soup washes it out to 0.300 — so soup (0.392 overall) even trails the best single specialist (0.425). Routing keeps the specialist intact (strings 0.670) and wins. Dilution is severe exactly when the specialist's contribution is hard-won.
  • So "merge, don't average" is a HEADROOM law, not a base-size law. Union > fusion whenever there is room to lose to dilution — a weak base (0.5B) or hard tasks at a strong base (7B-hard). Fusion only wins in the degenerate corner where the tasks are so easy the strong base composes to the 1.00 ceiling (7B-easy). This corrects the llm_moe_hpc interpretation: base capability was a confound; the operative variable is task headroom.
  • Riders unchanged. Learned router still perfect (1.00, lexical families); max_merge still the weakest union (0.297, not input-adaptive).

Takeaway

On genuinely hard, unsaturated tasks, the union operator (routing) beats fusion at 7B by the same margin it does at 0.5B — E8's max > mean in real weights, robust across scale once you control for saturation. The llm_moe_hpc flip is re-read as a saturation artefact. Falsifier (not triggered): fusion matching/beating routing on hard tasks — instead fusion diluted below even the best specialist. Provenance in manifest.json (hard: true, L40S, torch 2.12.1 / transformers 5.13.0 / peft 0.19.1).