src/neural/recombine.py mirrors Layer-1 run_coverage but trains K_T specialist RNNs on assignments from the exact shared-switch retention construction (K_T/rho/q clean; union matches the closed form), then recombines the measured teacher distributions two ways: mean (naive pooling) vs oracle-guided max-merge (per-mode strongest teacher, M2N2-style), each followed by size-n resampling. Result (8 reps): at rho=0, union rises 0.49->0.96 (supply matches closed form); analytic surviving_max rises 0.043->0.087 while surviving_mean stays flat ~0.045 — the conservation law (averaging cancels the union gain, max-merge realises it). At rho=1 (identical teachers) union and max are flat. The lesson holds in the neural setting; trained-weight columns show the same signs but noisier (smoothing inflates baseline; deep tail barely clears n=200 resampling). torch-gated test added. 93 tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
44 lines
685 B
YAML
44 lines
685 B
YAML
experiment: recombination
|
|
seed: 20260704
|
|
n_replicates: 8
|
|
source_config:
|
|
experiment: recombination
|
|
kind: recombination
|
|
seed: 20260704
|
|
n_replicates: 8
|
|
synthetic:
|
|
K: 256
|
|
R: 1
|
|
tail: zipf
|
|
zipf_s: 1.3
|
|
tail_frac: 0.5
|
|
tail_threshold: 0.001
|
|
style_len: 3
|
|
style_vocab: 5
|
|
id_base: 2
|
|
model:
|
|
kind: rnn
|
|
hidden: 128
|
|
embed: 24
|
|
epochs: 22
|
|
lr: 0.002
|
|
batch_size: 256
|
|
n_eval: 12000
|
|
coverage:
|
|
n: 200
|
|
q: 0.5
|
|
retain_thresh: 0.001
|
|
region_specialisation: false
|
|
sweep:
|
|
- param: K_T
|
|
values:
|
|
- 1
|
|
- 2
|
|
- 3
|
|
- 5
|
|
- param: rho
|
|
values:
|
|
- 0.0
|
|
- 1.0
|
|
output:
|
|
dir: results/recombination
|