# The composition campaign, seed 1 (prereg v3, amended after the generation-0 sweep of 2026-09-07). # # Arms. The gen-0 sweep found that the merge *weighting* dominates the operator: the a-priori 0.5/0.5 # blend fails under both operators (surplus -0.020 cat, -0.093 linear) while a selected weight passes # (+0.080 linear at 0.2/0.8, +0.027 cat at 0.3/0.7). Weights are therefore chosen each generation on a # disjoint validation split (E10, directed recombination) in every arm, and the operator is an # explicit per-arm setting: # dry — linear operator, no grounding [H2, H3, H5: does composition survive drift?] # grounded — linear operator, g = 0.10 [H4: does immigration arrest it?] # dry_cat — concatenation operator, no grounding [H6, revised: does the operator ordering hold # across generations, or only at gen 0?] experiment: llm_compose kind: llm_compose base_model: Qwen/Qwen2.5-1.5B seed: 1 generations: 6 arms: - dry - grounded - dry_cat g: 0.1 n_hard: 150 n_gsm8k: 150 n_mbpp: 100 n_probe: 60 k_inherit: 300 epochs: 3 conf_gate: 0.85 spec_train: 1200 spec_epochs: 3 max_new_tokens: 320 batch_size: 16 score_batch_size: 4 train_batch_size: 2 train_max_len: 448 resume: true lora: r: 16 alpha: 32 output: dir: results/llm_compose/s1 arm_ops: dry: linear grounded: linear dry_cat: cat n_hard_val: 60 merge_weights: - - 0.5 - 0.5 - - 0.3 - 0.7 - - 0.2 - 0.8 - - 0.1 - 0.9