# Generation-0 gate, second configuration (prereg v3 §4a): the composed target is MATH-500 # (competition maths, level >= 3, numeric answers), and merge weights are *selected* on a disjoint # validation split rather than fixed at 0.5/0.5. # # Why: on GSM8k-Hard the code parent alone reaches 0.427, because once code removes the arithmetic # burden the base's own reasoning suffices — so maths is not scarce and E8's premise fails. MetaMathQA # is built from GSM8K *and* MATH, so MATH-500 tests reasoning the specialist has and the base lacks. # Founders are shared with the first gate (same experiment name), so this costs evaluation only. experiment: llm_compose_gate kind: llm_compose base_model: Qwen/Qwen2.5-1.5B seed: 1 generations: 0 arms: [dry] target: math500 n_hard: 120 # test split (level>=3 pool is 271; 70% test / 30% val, disjoint) n_hard_val: 50 # val split, screens the merge weights only merge_weights: [[0.5, 0.5], [0.3, 0.7], [0.2, 0.8]] n_gsm8k: 100 n_mbpp: 80 n_probe: 40 k_inherit: 300 epochs: 3 conf_gate: 0.85 g: 0.10 spec_train: 1200 spec_epochs: 3 max_new_tokens: 320 batch_size: 16 score_batch_size: 8 train_batch_size: 2 train_max_len: 448 lora: {r: 16, alpha: 32} output: {dir: results/llm_compose_gate_math500}