Clarity pass over the main text (36-item audit), Discussion rewrite and cut, acknowledgements, Souly et al. as ref 62, lettered SI panels, model section moved under Results; plus the untracked curriculum/society/compose/smol configs, runners, figures, stats and tests that the SI already cites. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
34 lines
1.2 KiB
YAML
34 lines
1.2 KiB
YAML
# Generation-0 gate, second configuration (prereg v3 §4a): the composed target is MATH-500
|
|
# (competition maths, level >= 3, numeric answers), and merge weights are *selected* on a disjoint
|
|
# validation split rather than fixed at 0.5/0.5.
|
|
#
|
|
# Why: on GSM8k-Hard the code parent alone reaches 0.427, because once code removes the arithmetic
|
|
# burden the base's own reasoning suffices — so maths is not scarce and E8's premise fails. MetaMathQA
|
|
# is built from GSM8K *and* MATH, so MATH-500 tests reasoning the specialist has and the base lacks.
|
|
# Founders are shared with the first gate (same experiment name), so this costs evaluation only.
|
|
experiment: llm_compose_gate
|
|
kind: llm_compose
|
|
base_model: Qwen/Qwen2.5-1.5B
|
|
seed: 1
|
|
generations: 0
|
|
arms: [dry]
|
|
target: math500
|
|
n_hard: 120 # test split (level>=3 pool is 271; 70% test / 30% val, disjoint)
|
|
n_hard_val: 50 # val split, screens the merge weights only
|
|
merge_weights: [[0.5, 0.5], [0.3, 0.7], [0.2, 0.8]]
|
|
n_gsm8k: 100
|
|
n_mbpp: 80
|
|
n_probe: 40
|
|
k_inherit: 300
|
|
epochs: 3
|
|
conf_gate: 0.85
|
|
g: 0.10
|
|
spec_train: 1200
|
|
spec_epochs: 3
|
|
max_new_tokens: 320
|
|
batch_size: 16
|
|
score_batch_size: 8
|
|
train_batch_size: 2
|
|
train_max_len: 448
|
|
lora: {r: 16, alpha: 32}
|
|
output: {dir: results/llm_compose_gate_math500}
|