Manuscript revision and pending experiment work, snapshot before restructuring

Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
This commit is contained in:
Giorgio Gilestro 2026-09-13 16:54:09 +01:00
parent e4804adabc
commit 84124de143
450 changed files with 52813 additions and 1202 deletions

View file

@ -40,3 +40,26 @@ Together with `llm_moe_hard_hpc` this completes the correction — **both** "mer
the tasks are hard enough to leave room, not weak-base-only effects. **Falsifier (not triggered):**
directed offspring ≤ uniform soup — instead they beat it by 10 points. Provenance in `manifest.json`
(`hard: true`, L40S, torch 2.12.1 / transformers 5.13.0 / peft 0.19.1).
## Seeds 13 (2026-09-11)
Seeds 23 were run on CX3 via `hpc/llm_7b_seeds.pbs` (seed 1 above was moved to `s1/`; the bundle
layout is now `s{seed}/`). Fixed test sets, training seed varied. Per-seed values and mean ± 95% CI
from `figures/stats_llm_7b_seeds.py`:
```
model metric n_seeds s1 s2 s3 mean ci95
merge_soup overall 3 0.392 0.405 0.428 0.408 0.021
merge_soup worst_family 3 0.300 0.345 0.340 0.328 0.028
directed_overall overall 3 0.492 0.480 0.473 0.482 0.011
directed_overall worst_family 3 0.380 0.405 0.430 0.405 0.028
directed_balanced overall 3 0.492 0.480 0.473 0.482 0.011
directed_balanced worst_family 3 0.380 0.405 0.430 0.405 0.028
contrast metric n_seeds s1 s2 s3 mean ci95 sign_agrees
directed_overall merge_soup overall 3 0.10 0.075 0.045 0.073 0.031 3/3
directed_overall merge_soup worst_family 3 0.08 0.060 0.090 0.077 0.017 3/3
directed_balanced merge_soup overall 3 0.10 0.075 0.045 0.073 0.031 3/3
directed_balanced merge_soup worst_family 3 0.08 0.060 0.090 0.077 0.017 3/3
```
Reading: directed selection beats the a-priori soup in every seed (+0.073 ± 0.031 overall, +0.077 ± 0.017 worst-family). The two objectives selected the same offspring in all three seeds.

View file

@ -0,0 +1,26 @@
{
"experiment": "llm_directed_hard_hpc",
"master_seed": 1,
"git_commit": null,
"python": "3.11.13",
"libraries": {
"numpy": "2.4.6",
"scipy": "1.17.1",
"pandas": "3.0.3",
"pyarrow": "24.0.0",
"torch": "2.12.1",
"transformers": "5.13.0",
"peft": "0.19.1"
},
"rows": 35,
"results_sha256": "a5e45785a8701c5e1cdcfc64bc1482dffe4e68f8c6978e9d07dc8f04d40fbab6",
"layer": "2",
"tier": "llm",
"base_model": "Qwen/Qwen2.5-7B-Instruct",
"hard": true,
"directed": {
"n_candidates": 24,
"concentration": 0.5,
"n_val": 100
}
}

View file

@ -0,0 +1,25 @@
experiment: llm_directed_hard_hpc
seed: 1
n_replicates: 1
source_config:
experiment: llm_directed_hard_hpc
kind: llm_directed
seed: 1
n_replicates: 1
base_model: Qwen/Qwen2.5-7B-Instruct
hard: true
families:
- lists
- strings
- arith
n_train: 800
n_val: 100
n_test: 200
n_candidates: 24
concentration: 0.5
epochs: 3
lora:
r: 16
alpha: 32
output:
dir: results/llm_directed_hard_hpc

View file

@ -0,0 +1,26 @@
{
"experiment": "llm_directed_hard_hpc",
"master_seed": 2,
"git_commit": null,
"python": "3.11.7",
"libraries": {
"numpy": "2.4.6",
"scipy": "1.17.1",
"pandas": "3.0.3",
"pyarrow": "24.0.0",
"torch": "2.12.1",
"transformers": "5.16.1",
"peft": "0.20.0"
},
"rows": 35,
"results_sha256": "27ced93ac5ec7b6af3a5034d97d45326bde182b9894c5abdaf5f867405d978b3",
"layer": "2",
"tier": "llm",
"base_model": "Qwen/Qwen2.5-7B-Instruct",
"hard": true,
"directed": {
"n_candidates": 24,
"concentration": 0.5,
"n_val": 100
}
}

View file

@ -0,0 +1,25 @@
experiment: llm_directed_hard_hpc
seed: 2
n_replicates: 1
source_config:
experiment: llm_directed_hard_hpc
kind: llm_directed
seed: 2
n_replicates: 1
base_model: Qwen/Qwen2.5-7B-Instruct
hard: true
families:
- lists
- strings
- arith
n_train: 800
n_val: 100
n_test: 200
n_candidates: 24
concentration: 0.5
epochs: 3
lora:
r: 16
alpha: 32
output:
dir: results/llm_directed_hard_hpc/s2

View file

@ -0,0 +1,26 @@
{
"experiment": "llm_directed_hard_hpc",
"master_seed": 3,
"git_commit": null,
"python": "3.11.7",
"libraries": {
"numpy": "2.4.6",
"scipy": "1.17.1",
"pandas": "3.0.3",
"pyarrow": "24.0.0",
"torch": "2.12.1",
"transformers": "5.16.1",
"peft": "0.20.0"
},
"rows": 35,
"results_sha256": "ed2b40be721b464690960744b0a72ef2f39fc6d7623cc32f6128741e8d241699",
"layer": "2",
"tier": "llm",
"base_model": "Qwen/Qwen2.5-7B-Instruct",
"hard": true,
"directed": {
"n_candidates": 24,
"concentration": 0.5,
"n_val": 100
}
}

View file

@ -0,0 +1,25 @@
experiment: llm_directed_hard_hpc
seed: 3
n_replicates: 1
source_config:
experiment: llm_directed_hard_hpc
kind: llm_directed
seed: 3
n_replicates: 1
base_model: Qwen/Qwen2.5-7B-Instruct
hard: true
families:
- lists
- strings
- arith
n_train: 800
n_val: 100
n_test: 200
n_candidates: 24
concentration: 0.5
epochs: 3
lora:
r: 16
alpha: 32
output:
dir: results/llm_directed_hard_hpc/s3