MachineSex/results/llm_merge
Giorgio Gilestro 585264d0b4 llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive
Re-ran the specialist-merge experiment at a capable base (Qwen2.5-7B-Instruct,
200 tests/family) on one L40S GPU of Imperial's CX3 HPC (8 min walltime). The
two caveats the 0.5B prototype left marginal are now resolved:

- "exceeds every parent overall" is clean: both merges 0.87 vs best specialist
  0.77 (+10 pts), and above every specialist on every family.
- dilution vanishes: at 0.5B averaging diluted the lists-specialist
  (0.43->0.26); at 7B the merge beats it (0.62>0.57). Dilution was a
  small-model artefact -- a capable base composes rather than dilutes, which
  softens E4's "merge, don't average" once the parents are strong.

The figure title is now data-driven (reports ">" for 7B, "~" for 0.5B).
Adds the hpc/ smoke job script and the llm_merge walltime trim. Results synced
to results/llm_merge_hpc/ (parquet gitignored per the reproducibility contract).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 17:32:51 +01:00
..
llm_merge.pdf llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive 2026-07-05 17:32:51 +01:00
llm_merge.png llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive 2026-07-05 17:32:51 +01:00
manifest.json llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00
README.md llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00
resolved_config.yaml llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00

llm_merge — recombining specialist LLMs (the first real-LLM prototype; blueprint C2/C4)

Claim tested. The first step from toy models toward real language models: does the sexual- reproduction result — recombining decorrelated specialists yields a model that exceeds/retains what any single parent has (E8) — appear in real LoRA-adapted LLM weights? This is a prototype, run on a single 16 GB consumer GPU, not the full society.

Setup. Base model Qwen2.5-0.5B-Instruct (Apache-2.0). Three disjoint, procedurally-generated task families with an exact-match verifier (the "reality that says no"): lists (list ops), strings (string ops), arith (integer arithmetic), deliberately made hard so specialists decorrelate. One LoRA specialist is fine-tuned per family (~90 s for all three), then the base, each specialist, and two weight-space mergessoup (averaged LoRA deltas) and ties (sign-reconciled union) — are evaluated on a held-out mixed test set. Seed 1, 100 test tasks/family.

Results (accuracy)

model lists strings arith overall worst family
base 0.15 0.15 0.53 0.28 0.15
spec: lists 0.43 0.16 0.71 0.43 0.16
spec: strings 0.08 1.00 0.80 0.63 0.08
spec: arith 0.11 0.22 0.78 0.37 0.11
merge: soup 0.26 0.74 0.91 0.64 0.26
merge: ties 0.23 0.71 0.90 0.61 0.23

What holds, and what doesn't (honest)

  • Strong and robust — balance / "retains all specialties". The merges are the only models competent across all families: worst-family ≈ 0.25, versus < 0.16 for every single specialist (the best specialist, strings, is at 0.08 on its worst family). Each specialist spikes on its own family and is weak elsewhere; the merge is decent everywhere. This is the Fisher-Muller "a generalist assembled from specialists" signature, in real LLM weights.
  • Marginal / noisy — "exceeds any parent overall". On overall accuracy the merge only matches the best specialist (soup 0.64 vs strings-specialist 0.63; ties 0.61 is slightly below). At this scale (a 0.5 B model, 3 families, one seed) the strict "offspring exceed every parent" claim is not cleanly established.
  • The dilution caveat, visible in the flesh. On lists, the lists-specialist alone scores 0.43 but the merge only 0.230.26 — weight-averaging diluted that specialist's contribution. This is exactly Layer-1's "merge, don't average" concern (E4) appearing in real weights; the finer soup-vs-ties advantage is not resolved at K=3.

Takeaway

The pipeline runs end-to-end on real LLMs on a 16 GB GPU (specialise → verify → merge → evaluate), and the balance/retention half of the sexual-reproduction claim reproduces clearly. The stronger "exceeds every parent" claim is marginal at this toy scale and is the thing a larger run should firm up — more, cleaner-decorrelated families; a bigger base; multiple seeds; and a merge that resists dilution (e.g. per-task-family weighting, or the offspring-selection of "directed sex"). That scaling is the natural HPC step; this prototype de-risks the machinery and shows the first sign in real weights. Falsifier (partially triggered — reported honestly): a single specialist matches the merge on overall here; the merge's advantage is currently specific to cross-family balance.