# llm_merge — recombining specialist LLMs (the first real-LLM prototype; blueprint C2/C4) **Claim tested.** The first step from toy models toward real language models: does the sexual- reproduction result — recombining decorrelated specialists yields a model that exceeds/retains what any single parent has (E8) — appear in real LoRA-adapted LLM weights? This is a **prototype**, run on a single 16 GB consumer GPU, not the full society. **Setup.** Base model **Qwen2.5-0.5B-Instruct** (Apache-2.0). Three *disjoint*, procedurally-generated task families with an **exact-match verifier** (the "reality that says no"): `lists` (list ops), `strings` (string ops), `arith` (integer arithmetic), deliberately made hard so specialists decorrelate. One **LoRA specialist** is fine-tuned per family (~90 s for all three), then the base, each specialist, and two weight-space **merges** — `soup` (averaged LoRA deltas) and `ties` (sign-reconciled union) — are evaluated on a held-out mixed test set. Seed 1, 100 test tasks/family. ### Results (accuracy) | model | lists | strings | arith | overall | **worst family** | |---|---|---|---|---|---| | base | 0.15 | 0.15 | 0.53 | 0.28 | 0.15 | | spec: lists | 0.43 | 0.16 | 0.71 | 0.43 | 0.16 | | spec: strings | 0.08 | **1.00** | 0.80 | 0.63 | 0.08 | | spec: arith | 0.11 | 0.22 | 0.78 | 0.37 | 0.11 | | **merge: soup** | 0.26 | 0.74 | 0.91 | 0.64 | **0.26** | | **merge: ties** | 0.23 | 0.71 | 0.90 | 0.61 | **0.23** | ### What holds, and what doesn't (honest) - **Strong and robust — balance / "retains all specialties".** The merges are the *only* models competent across **all** families: worst-family ≈ **0.25**, versus **< 0.16** for every single specialist (the best specialist, strings, is at 0.08 on its worst family). Each specialist spikes on its own family and is weak elsewhere; the merge is decent everywhere. This is the Fisher-Muller "a generalist assembled from specialists" signature, in real LLM weights. - **Marginal / noisy — "exceeds any parent overall".** On *overall* accuracy the merge only *matches* the best specialist (soup 0.64 vs strings-specialist 0.63; ties 0.61 is slightly below). At this scale (a 0.5 B model, 3 families, one seed) the strict "offspring exceed every parent" claim is not cleanly established. - **The dilution caveat, visible in the flesh.** On `lists`, the lists-specialist alone scores 0.43 but the merge only 0.23–0.26 — weight-averaging *diluted* that specialist's contribution. This is exactly Layer-1's "merge, don't average" concern (E4) appearing in real weights; the finer soup-vs-ties advantage is not resolved at K=3. ### Takeaway The pipeline runs end-to-end on real LLMs on a 16 GB GPU (specialise → verify → merge → evaluate), and the **balance/retention** half of the sexual-reproduction claim reproduces clearly. The stronger "exceeds every parent" claim is marginal at this toy scale and is the thing a larger run should firm up — more, cleaner-decorrelated families; a bigger base; multiple seeds; and a merge that resists dilution (e.g. per-task-family weighting, or the offspring-selection of "directed sex"). That scaling is the natural HPC step; this prototype de-risks the machinery and shows the first sign in real weights. **Falsifier (partially triggered — reported honestly):** a single specialist matches the merge on *overall* here; the merge's advantage is currently specific to cross-family *balance*.