Manuscript revision and pending experiment work, snapshot before restructuring

Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
This commit is contained in:
Giorgio Gilestro 2026-09-13 16:54:09 +01:00
parent e4804adabc
commit 84124de143
450 changed files with 52813 additions and 1202 deletions

View file

@ -0,0 +1,34 @@
# Generation-0 gate, second configuration (prereg v3 §4a): the composed target is MATH-500
# (competition maths, level >= 3, numeric answers), and merge weights are *selected* on a disjoint
# validation split rather than fixed at 0.5/0.5.
#
# Why: on GSM8k-Hard the code parent alone reaches 0.427, because once code removes the arithmetic
# burden the base's own reasoning suffices — so maths is not scarce and E8's premise fails. MetaMathQA
# is built from GSM8K *and* MATH, so MATH-500 tests reasoning the specialist has and the base lacks.
# Founders are shared with the first gate (same experiment name), so this costs evaluation only.
experiment: llm_compose_gate
kind: llm_compose
base_model: Qwen/Qwen2.5-1.5B
seed: 1
generations: 0
arms: [dry]
target: math500
n_hard: 120 # test split (level>=3 pool is 271; 70% test / 30% val, disjoint)
n_hard_val: 50 # val split, screens the merge weights only
merge_weights: [[0.5, 0.5], [0.3, 0.7], [0.2, 0.8]]
n_gsm8k: 100
n_mbpp: 80
n_probe: 40
k_inherit: 300
epochs: 3
conf_gate: 0.85
g: 0.10
spec_train: 1200
spec_epochs: 3
max_new_tokens: 320
batch_size: 16
score_batch_size: 8
train_batch_size: 2
train_max_len: 448
lora: {r: 16, alpha: 32}
output: {dir: results/llm_compose_gate_math500}