llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive
Re-ran the specialist-merge experiment at a capable base (Qwen2.5-7B-Instruct, 200 tests/family) on one L40S GPU of Imperial's CX3 HPC (8 min walltime). The two caveats the 0.5B prototype left marginal are now resolved: - "exceeds every parent overall" is clean: both merges 0.87 vs best specialist 0.77 (+10 pts), and above every specialist on every family. - dilution vanishes: at 0.5B averaging diluted the lists-specialist (0.43->0.26); at 7B the merge beats it (0.62>0.57). Dilution was a small-model artefact -- a capable base composes rather than dilutes, which softens E4's "merge, don't average" once the parents are strong. The figure title is now data-driven (reports ">" for 7B, "~" for 0.5B). Adds the hpc/ smoke job script and the llm_merge walltime trim. Results synced to results/llm_merge_hpc/ (parquet gitignored per the reproducibility contract). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
38bb252c20
commit
585264d0b4
12 changed files with 140 additions and 22 deletions
|
|
@ -344,3 +344,11 @@ C3 vertical claim deferred.*
|
|||
overall-exceeds needs scale (bigger base/more families/seeds/dilution-resistant merge) = HPC step.
|
||||
- Python 3.14 + transformers 5.13 OK; note transformers-5.x apply_chat_template returns a dict.
|
||||
`make env-llm`/`make llm`; `figures/plot_llm_merge.py`, README, `tests/test_llm.py` (+3, 125 green).
|
||||
|
||||
**2026-07-05 — LLM merge 7B firm-up on Imperial CX3 (`llm_merge_hpc`): marginal sign → decisive.** ✅
|
||||
- Ran on one L40S (46 GB) via `/imperial-hpc` runbook; 8 min walltime; Qwen2.5-7B-Instruct, 200 tests/family.
|
||||
- **Both merges 0.87 overall > best specialist 0.77** (decisive +10 pts) and beat every specialist on
|
||||
every family; worst-family 0.62 vs ≤0.57. Both 0.5B caveats resolved: overall-exceeds is now clean,
|
||||
and dilution VANISHES (merge 0.62 > lists-spec 0.57 on lists) — dilution was a small-model artefact.
|
||||
- `results/llm_merge_hpc/` (README legend, data-driven figure title). Next refinement: module-level
|
||||
union-preserving recombination (MoE-expert/adapter-union = real-weight E8 max-merge), not delta-avg.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue