Re-ran the specialist-merge experiment at a capable base (Qwen2.5-7B-Instruct, 200 tests/family) on one L40S GPU of Imperial's CX3 HPC (8 min walltime). The two caveats the 0.5B prototype left marginal are now resolved: - "exceeds every parent overall" is clean: both merges 0.87 vs best specialist 0.77 (+10 pts), and above every specialist on every family. - dilution vanishes: at 0.5B averaging diluted the lists-specialist (0.43->0.26); at 7B the merge beats it (0.62>0.57). Dilution was a small-model artefact -- a capable base composes rather than dilutes, which softens E4's "merge, don't average" once the parents are strong. The figure title is now data-driven (reports ">" for 7B, "~" for 0.5B). Adds the hpc/ smoke job script and the llm_merge walltime trim. Results synced to results/llm_merge_hpc/ (parquet gitignored per the reproducibility contract). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
24 lines
364 B
YAML
24 lines
364 B
YAML
experiment: llm_merge_hpc
|
|
seed: 1
|
|
n_replicates: 1
|
|
source_config:
|
|
experiment: llm_merge_hpc
|
|
kind: llm_merge
|
|
seed: 1
|
|
n_replicates: 1
|
|
base_model: Qwen/Qwen2.5-7B-Instruct
|
|
families:
|
|
- lists
|
|
- strings
|
|
- arith
|
|
n_train: 800
|
|
n_test: 200
|
|
epochs: 3
|
|
lora:
|
|
r: 16
|
|
alpha: 32
|
|
merges:
|
|
- soup
|
|
- ties
|
|
output:
|
|
dir: results/llm_merge_hpc
|