Re-ran the specialist-merge experiment at a capable base (Qwen2.5-7B-Instruct, 200 tests/family) on one L40S GPU of Imperial's CX3 HPC (8 min walltime). The two caveats the 0.5B prototype left marginal are now resolved: - "exceeds every parent overall" is clean: both merges 0.87 vs best specialist 0.77 (+10 pts), and above every specialist on every family. - dilution vanishes: at 0.5B averaging diluted the lists-specialist (0.43->0.26); at 7B the merge beats it (0.62>0.57). Dilution was a small-model artefact -- a capable base composes rather than dilutes, which softens E4's "merge, don't average" once the parents are strong. The figure title is now data-driven (reports ">" for 7B, "~" for 0.5B). Adds the hpc/ smoke job script and the llm_merge walltime trim. Results synced to results/llm_merge_hpc/ (parquet gitignored per the reproducibility contract). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
20 lines
No EOL
468 B
JSON
20 lines
No EOL
468 B
JSON
{
|
|
"experiment": "llm_merge_hpc",
|
|
"master_seed": 1,
|
|
"git_commit": null,
|
|
"python": "3.11.13",
|
|
"libraries": {
|
|
"numpy": "2.4.6",
|
|
"scipy": "1.17.1",
|
|
"pandas": "3.0.3",
|
|
"pyarrow": "24.0.0",
|
|
"torch": "2.12.1",
|
|
"transformers": "5.13.0",
|
|
"peft": "0.19.1"
|
|
},
|
|
"rows": 30,
|
|
"results_sha256": "6cc0a07c66ba92a379d895d6d6707591aced48f06eee895bb4f6c15d12e6e588",
|
|
"layer": "2",
|
|
"tier": "llm",
|
|
"base_model": "Qwen/Qwen2.5-7B-Instruct"
|
|
} |