MachineSex/tasks/prereg-llm-society-v4.md
Giorgio Gilestro 84124de143 Manuscript revision and pending experiment work, snapshot before restructuring
Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 16:54:09 +01:00

45 KiB
Raw Blame History

Pre-registration — llm_curriculum v4: does a society accumulate more than its members?

Status: draft for GG review, 2026-09-08. Supersedes prereg-llm-compose-v3.md (run; H1 passed, H2H5 null and uninterpretable). Nothing runs until §5's gates pass and GG signs off §8.


0. Why v4: v3 measured the wrong thing

v3 had a fixed skill set. Two founders were trained once and every later generation was a lossy copy, so the experiment's ceiling was its own generation 0 and no outcome could have shown capability climbing. It answered "do ancestral skills degrade?", a retention question. The paper's claim (C3) is that capability climbs — each specialty re-earned and exceeded. GG, 2026-09-08: "are models learning NEW skills at EACH generation? or are we just seeing if the ancestral skills degrade? because that was not the problem being addressed… we need to ground this into continual learning."

The specific technical fault: v3 trained a fresh LoRA from the base each generation, so knowledge survived only through the data channel. That is Weismannian — nothing acquired is inherited as structure. v4's children start from their parent's adapter, which is the actual Lamarckian channel and the precondition for accumulation.

Pattern across v2 → v3 → v4, recorded so it stops recurring: each design was checked against the mechanism (drift, immigration, recombination) and never against the claim. §5's gate G0 exists solely to check the claim is reachable before any compute is spent.


1. Design

Curriculum. Nine task families from v2's calibrated set (llm_society_v2_calib_a2: base ≤ 0.40, specialist ≥ 0.60, pairwise conflict < 0.41 — strings, setops, numtheory, mixedtoken, digits, alphabet, prime, wordlen, roman). Three lineages, nine generations. Each lineage sees all nine families in a different order (a cyclic Latin square), so at generation t every lineage has met t families but different ones. Complementarity is maximal early and decays to zero by generation 9 — a shape the analysis can test, not just a condition it assumes.

One generation, per lineage:

  1. Acquire — the environment presents the next family; train on n_new verified real examples.
  2. Inherit — training starts from the parent's adapter, not the base (Lamarckian transmission).
  3. Maintain — old families are kept alive by n_replay real examples (grounding = immigration), or by self-generated answers (dry), or not at all, depending on arm.
  4. Recombine — merge with the arm's partner (contemporary, ancestor, or nobody), weights chosen on a held-out validation split (directed recombination, carried over from v3).

Two external-information channels, deliberately separated — v3 conflated them. Acquisition is a capability the population never had (novel allele; moves the frontier). Replay is re-supply of a capability already present (immigration proper; fights loss, never advances). They fight different diseases and must be separate factors.

2. Arms

arm recombines with old skills maintained by isolates
isolated nobody real replay asexual continual learning — the drift/forgetting baseline
society a decorrelated contemporary real replay the treatment
society_dry contemporary self-generated only replay's contribution (E2 immigration)
seed_bank its own ancestor at t3 real replay see below — this is not a throwaway control

The seed-bank arm is a substantive comparison, not a null. My first reading was that an ancestor is your own lineage and therefore highly correlated (ρ→1), so E8 predicts it buys nothing. Working through the curriculum shows that is wrong: your t3 ancestor knows exactly the families you learned three generations ago and have since been forgetting. It carries temporal complementarity where a contemporary carries spatial complementarity. So the arms pose a real question with predictions pulling opposite ways:

  • E8 (decorrelation is the fuel): the contemporary is more decorrelated → larger union → society wins.
  • E12 (merge compatibility): the ancestor is same-lineage, so no BatesonDobzhanskyMuller incompatibilities have had time to accumulate → it merges more safelyseed_bank wins.

Which dominates is not obvious from the framework, and the answer is directly translational: when a model forgets, is it better recovered from a peer who knows something else, or from your own earlier checkpoint? Nobody has posed that as a population-genetic question.

Dropped (GG, 2026-09-08): a no_acquisition arm — v3 already is that experiment.

3. The mono-generational baselines, and why they are the falsifier

The difficulty GG named is pairing a multigenerational design against single-shot SoTA. Rather than work around it, it becomes the spine: does structure across time beat the same compute spent all at once? Three references, all at matched total training examples:

baseline what it is role
sequential one model, all nine families in sequence the standard continual-learning baseline
single_shot_merge three specialists trained from base in parallel (three families each), merged once at the end actual SoTA — LoRA Soups / TIES / model soup
joint one model trained on all nine jointly the conventional ceiling

Budget accounting (must be equal, and is checked in the artifact). A lineage trains 9 generations × (n_new + n_replay) examples; three lineages give 27 × (n_new + n_replay). Each single-shot specialist gets 9 × (n_new + n_replay) so three of them match exactly; sequential and joint receive the same total. Compute per arm is recorded in the manifest.

If society does not beat single_shot_merge at matched budget, iterating buys nothing and the multigenerational framing is decoration. That is the claim worth staking, and the merging literature has never tested it because every paper in it merges once.

4. Hypotheses

Primary outcome: cumulative capability — accuracy of the population's best model on all nine families, at each generation. Secondary: per-family forgetting curves, union-exceedance at each merge, ρ between partners, and the complementarity decay predicted by the Latin-square design.

prediction (source) threshold falsified if
H1 Capability climbs: society at generation 9 exceeds its own generation 1 by ≥ 0.15 ≥ 0.15 in ≥ 2 of 3 seeds flat or declining → the design still cannot show accumulation (this is G0 restated as a result)
H2 FisherMuller (E7): society > isolated on cumulative capability at generation 9 ≥ +0.08, 3/3 seeds positive recombination adds nothing over isolated continual learning
H3 The multigenerational claim: society > single_shot_merge at matched budget ≥ +0.05 iterating buys nothing; the framing is decorative and the paper should say so
H4 Immigration (E2): society > society_dry ≥ +0.08 self-generated replay suffices; grounding is not load-bearing here
H5 Spatial vs temporal complementarity: societyseed_bank, direction not pre-committed (§2 gives arguments both ways) report with CI either way — (this is a measurement, not a gated prediction)
H6 Complementarity decays by construction, so the society isolated gap is largest at intermediate generations and shrinks by generation 9 peak gap at 3 ≤ t ≤ 6 a monotone gap → the advantage is not coming from complementarity, and the mechanism story is wrong

H6 is the design's internal check: the Latin square makes complementarity a known function of generation, so the framework predicts the shape of the advantage, not just its sign.

5. Gates — G0 is the one that would have caught v2 and v3

gate what pass criterion
G0 — can capability climb at all? One lineage, 3 generations, 3 families, isolated settings. Measure cumulative accuracy over families seen. Generation 3 exceeds generation 1 by ≥ 0.10. If capability cannot accumulate in the simplest arm, no outcome of the full design is interpretable — stop.
G1 — does inheritance transmit? Adapter-continued training on family 2 starting from the family-1 adapter family-2 accuracy ≥ 0.6 × a from-scratch specialist's
G2 — does forgetting occur? Same, measuring family-1 accuracy after learning family 2 without replay family-1 accuracy drops ≥ 0.15 (if nothing is forgotten, replay and recombination have nothing to fix)
G3 — does recombination combine? Merge two lineages holding disjoint families at generation 3 merged model ≥ 0.8 × each parent's accuracy on that parent's own families
G4 — budget parity Recorded example counts across all arms and baselines equal to within 2%

G1 and G2 must both pass: transmission without forgetting means nothing decays; forgetting without transmission means nothing accumulates. The experiment needs the tension.

6. Cost

Base: Qwen2.5-1.5B (base, not Instruct — v3's C1 measured that instruction-tuned checkpoints already hold these skills). Per generation-lineage: train (n_new + n_replay ≈ 400 examples × 3 epochs) plus evaluation on nine families × 60 items. ≈ 6 min. Society arms: 3 lineages × 9 generations × 4 arms ≈ 108 generation-lineages ≈ 11 GPU-h; baselines ≈ 2 GPU-h; three seeds ≈ 40 GPU-h total. Seed 1 local overnight, seeds 23 as a CX3 array — the same split that worked last night.

7. Engineering

Reused unchanged from v3: execution/verifier infrastructure, LoRA training, merge operators, directed weight selection on a disjoint validation split, checkpoint/resume, the figure and stats scaffolding. Reused from v2: the nine calibrated families and their difficulty/conflict measurements.

New: adapter-continued training (inherit the parent's weights rather than a fresh LoRA); the Latin-square curriculum scheduler; replay buffers per lineage; the ancestor registry for seed_bank; budget accounting in the manifest; the three mono-generational baselines.

8a. Gate record (2026-09-08) and the v5 curriculum

G0 passes. Three families, three lineages: all_families (mean over every family in the curriculum — the accumulation metric) climbs 0.739 → 0.950 (isolated) and 0.611 → 0.939 (society). The design can show accumulation, which v3 structurally could not. A labelling hazard caught on the first line: retention_seen (mean over families taught) starts near 1 and can only fall — watching it would have recreated v3's error. Both are recorded; the primary is all_families.

G2 fails on the v2 families, twice. Three-family gate with replay: nothing forgotten (every family only rises). Nine-family single-lineage probe with replay off (results/llm_curriculum_g2): mean drop across families learned before the last is only +0.074, and it is carried by one family — mixedtoken 0.80 → 0.15 (+0.65), oscillating violently throughout (0.18, 0.20, 0.27, 0.80, 0.60, 0.37, 0.42, 0.20, 0.15) — while two families improve through positive transfer (numtheory 0.12, alphabet 0.13) and the rest are within ±0.08. Forgetting here is a single pairwise-interference event between confusable counting families (the cluster v2's conflict measure flagged at 0.44), not a general pressure a population could smooth.

G3 is negative on the v2 families. society isolated = 0.128, 0.022, 0.011 across the three gate generations. Merging costs and cannot pay, because a partner can only contribute what the recipient lacks and nothing was lacking.

The base reference quantifies the format confound. Qwen2.5-1.5B base on all nine v2 families: 0.094 (strings 0.02, setops 0.02, alphabet 0.00, wordlen 0.00, prime 0.02, digits 0.07, roman 0.10, mixedtoken 0.12, numtheory 0.52). Training on one family lifts the nine-family mean to 0.417. Most of the apparent accumulation is a one-off format acquisition shared by all nine families.

One root, three faults. The v2 families were built for a specialisation experiment and calibrated for low mutual conflict so merging would be safe. That makes them (i) format-homogeneous — one family teaches the convention for all; (ii) too compatible — no interference, hence no forgetting; (iii) too easy — the base is unformatted, not incapable. They are not a curriculum, and no arrangement of arms fixes that. The same lesson as v3's task pairing: the paradigmatic continual-learning benchmarks use naturally heterogeneous tasks because those interfere, differ in format, and exceed a small base.

v5 curriculum (src/llm/curriculum_data.py, 2026-09-08). Eleven candidates from public datasets, each with a disjoint train/test split and its own verifier — gsm8k (number), mbpp (code, executed), boolq (yes/no), mnli (3-way label), sst2 (sentiment word), csqa (AE), arc (AD), winogrande (1/2), squad (extractive span, normalised EM over aliases), nq_open (short text, aliases), hellaswag (AD). All eleven self-verify 40/40 and reject garbage 0/40. Answer shapes span five forms (code, number, word, letter, phrase) against v2's one. Selection rule, fixed before running: the C1 band (base ≤ 0.40, specialist ≥ 0.60) from curriculum_v5_calib (one specialist per candidate, evaluated on every family — the full transfer matrix), then a single-lineage zero-replay probe over the chosen set with mean forgetting ≥ 0.15 and not carried by a single family (max single-family share of the total drop ≤ 50%). Nine survivors form the curriculum; if fewer than nine pass, L and G shrink to match and the Latin square is recomputed.

v5 stage A (results/llm_curriculum_v5_calib, founders at 300 × 3, Qwen2.5-1.5B base). Base on all eleven: 0.011 (mbpp 0.08, gsm8k 0.02, the rest 0.00) — a non-instruct base on real tasks, as expected; the C1 base floor is moot. Specialist × family matrix (own-family on the diagonal):

specialist own mean off-family character
mnli 0.82 0.46 permissive — lifts most others
arc 0.77 0.21
hellaswag 0.72 0.27
squad 0.68 0.34 permissive
boolq 0.65 0.04 destructive — zeroes others
csqa 0.55 0.36 permissive
sst2 0.38 0.23
winogrande 0.38 0.00 destructive — 0.00 on all ten others
mbpp 0.18 0.08 destructive
nq_open 0.18 0.36
gsm8k 0.10 0.42 permissive (0.92 on sst2, 0.83 on arc; 0.10 on gsm8k itself)

Five pass C1 at this budget. Mean off-family transfer 0.25 against own-family 0.49 — half, where v2's one family lifted all nine to near-own level. Two kinds of specialist, which is what a curriculum needs: format-permissive ones (mnli, csqa, squad, gsm8k) teach general instruction-following and lift other families — the residual format-transfer effect, now bounded and measurable; format-destructive ones (winogrande, boolq, mbpp) learn one narrow output form and erase the rest. The destructive group is the forgetting mechanism made visible: a lineage that meets winogrande loses what it held, and replay or a partner who did not just learn winogrande is what can restore it — E8 with something to act on.

The six failures were under-trained, not unlearnable: every specialist that has worked in this project (v2, v3) had 1200 examples; these had 300 (gsm8k reached 0.540.62 at 1200 in v3). Pass 2 (curriculum_v5_calib_b, the six at 1200 × 3) decides six families or nine. Budget consequence, pre-noted: if 1200 is what a family needs, n_new in the campaign rises accordingly and §6's cost scales by ~4× on the training term.

v5 stage A pass 2 (results/llm_curriculum_v5_calib_b, the six failures at 1200 × 3) — the "under-trained" hypothesis is refuted. winogrande 0.38 → 0.57, sst2 0.38 → 0.52, mbpp 0.18 → 0.22, gsm8k 0.10 → 0.07, nq_open 0.18 → 0.07, csqa 0.55 → 0.18 (chance on 5-way; off-family transfer collapsed to 0.00). Three got worse with four times the data — a training instability of the fresh-adapter learning rate on these tasks, not a data shortage. gsm8k is a data-source issue: v3's 0.540.62 came from MetaMathQA's augmented chain-of-thought, not raw GSM8K. None of the six clears 0.60.

Selection (2026-09-08): six families. mnli 0.82, arc 0.77, hellaswag 0.72, squad 0.68, boolq 0.65 pass C1; winogrande 0.57 is the sixth. Amendment: the C1 specialist floor is relaxed 0.60 → 0.55 for one family so that F is divisible by L = 3 (the pre-registered shrink rule needs F ∈ {3, 6, 9}). winogrande is also the right sixth on the merits: it is the most format-destructive specialist in the matrix (0.00 on every other family), i.e. the strongest forgetting pressure available — the mechanism the design exists to test. Curriculum: L = 3, F = 6, G = 6; complementarity 1.0 at t = 2, 0 at t = 6. Founders at 300 examples (the budget that passed; 1200 destabilised). Campaign n_new = 300.

Stage B running: curriculum_v5_g2 — one lineage, zero replay, the six in sequence. Pass criterion unchanged: mean drop ≥ 0.15 across families learned before the last, no single family > 50% of the total.

v5 stage B — G2 on the six (results/llm_curriculum_v5_g2, one lineage, zero replay). Order mnli → arc → hellaswag → squad → boolq → winogrande. Drops (learned → final): mnli +0.27, arc +0.03, hellaswag 0.05, squad +0.57, boolq +0.05. Mean +0.173 — magnitude gate (≥ 0.15) PASSES. All-families 0.575 → 0.725 (gen 3) → 0.625 (gen 5): acquisition then loss as the destructive families arrive — the tension the arms need. Concentration criterion (≤ 50% in one family) MISSES at 62% (squad). Recorded as a marginal miss, with the reasons it does not reproduce the v4 failure the criterion was written against: two families forgotten (not one pair), the destruction lands exactly where the transfer matrix predicted (winogrande's option format erases span and 3-way label; letter formats survive), and the probe tested one order where the campaign's Latin square gives each lineage a different one — so different families are forgotten in different lineages, which is the complementarity recombination acts on. Recommendation: go, pending GG.

8b. The v5 campaign result, and the two-kinds-of-variation measurement (2026-09-08)

Campaign (3 seeds, 6 families, 3 lineages, 6 generations; results/llm_curriculum_v5/). Best model per arm at the final generation, mean over seeds: sequential (one model, no population) 0.802 · isolated (population, never merges) 0.796 · joint (multi-task ceiling) 0.748 · seed_bank (merges with its own ancestor at t3) 0.663 · single_shot_merge 0.549 (0.125 / 0.758 / 0.764 — the huge variance is which families landed in which allopatric split) · society_dry 0.307 · society (merges with a contemporary) 0.269. Budget parity within 6%.

Every pre-registered hypothesis fails, consistently across all three seeds: society isolated 0.527 ± 0.092 (3/3 negative); society single-shot 0.280 ± 0.363 (unresolved, huge variance); society society_dry 0.038 ± 0.073 (replay policy irrelevant once the merge channel dominates). The one uncommitted contrast resolves decisively: seed_bank society = +0.394 ± 0.089, 3/3 positive — merging with your own past beats merging with a peer, so the compatibility argument beats the decorrelation argument.

Mechanism, identified and isolated. Two of the six families (boolq, winogrande) are answer-format destroyers — the calibration matrix measured them at 0.000.04 mean off-family. A lineage that learns one propagates it through the merge into partners that never trained on it; because merged offspring continue the lineage, the damage compounds (society lineage 0, gen 2→3: five families fall together while boolq alone rises). An ancestor cannot transmit a family the lineage never met, which is exactly why the seed-bank arm holds.

Scope limits, stated plainly. (i) Merging was obligate — no veto, no option to keep the parent unchanged. (ii) There is no selection between lineages: all three persist regardless of fitness, so the design has transmission, acquisition, gene flow and immigration but no differential reproduction. It is a gene-flow experiment, not a natural-selection one, and is therefore not a test of the composed-society claim. (iii) The operator was linear averaging, which the merging literature ranks below concatenation — but concatenation doubles adapter rank per merge, so iterated merging faces a capacity constraint single-shot merging never meets (16 → 1024 over six generations). That constraint is itself a finding about iteration.

Two kinds of variation are opposite in sign (/tmp/paralleldiv.py, 2026-09-08). Three adapters on the same family differing only in seed and data draw: accuracies 0.762 / 0.800 / 0.312 (one run simply failed — training instability); pairwise output disagreement 0.237 between the two good ones; weight cosine +0.006 (near-orthogonal); either-right 0.887 vs both-right 0.675. Merging the two good ones gives 0.887 — +0.087 over the better parent, landing exactly on the either-right ceiling. Same base, same linear operator, same scale, same evaluation as the collapsing arms.

So the framework's single decorrelation parameter conflates two quantities that behave oppositely:

  • decorrelation in what parents know → risk (0.53 measured);
  • decorrelation in how parents encode the same knowledge → benefit (+0.087, at the ceiling).

Equal weights are best for same-skill merging (0.887 at 0.5/0.5 vs 0.863 at 0.3/0.7), the reverse of skill composition where asymmetric weights won decisively (v3: 0.507 at 0.2/0.8 vs 0.333 at 0.5/0.5). The optimal merge weight is a signal of which regime the merge is in.

Design consequence for v6 (GG, 2026-09-08 — selection is needed to generalise the hypothesis). Three arms — never merge · complementary partners (different skills) · parallel partners (same skills, different seed/data draw) — with selection added in two places: a veto ("keep the parent unchanged" is always a candidate offspring) and population selection (score parents and offspring together, keep the best; the failed 0.312 run is exactly what it should discard). Roughly doubles evaluation cost per generation: ≈ 10 GPU-h per seed, ≈ 30 for three.

8c. Mechanism probes, 2026-09-08 — two of my explanations retracted

Four cheap probes run after the v5 campaign, chasing why cross-lineage merging collapsed. They retract two explanations I had given and leave a third standing.

Probe 1 — same-skill variation (/tmp/paralleldiv.py). Three adapters, one family (arc), same data distribution, differing only in seed and draw. Accuracies 0.762 / 0.800 / 0.312 (one run simply failed). Pairwise output disagreement 0.237; weight cosine +0.006 (near-orthogonal); either-right 0.887 vs both-right 0.675. Merging the two good ones: 0.887 — +0.087 over the better parent, exactly at the either-right ceiling. Equal weights beat 0.3/0.7 (0.887 vs 0.863), the reverse of skill composition, where asymmetric weights won.

Probe 2 — signal/noise and the inbred-lines cross (/tmp/inbred.py). Across-seed decomposition: signal power 17.8 vs noise 34.9 for arc, 19.0 vs 37.1 for boolq — and correcting for the K=3 sample mean's own noise puts the true signal near 6.2, i.e. ~85% of a LoRA's weight change is run-specific and arbitrary. That is why raw weight distance measures mostly noise and predicted nothing in Fig. 3C-D. Variance-weighted overlap surfaces ~4x more of the real overlap (0.0015 -> 0.0066) but different skills stay near-orthogonal even in signal directions. Denoising before crossing helps modestly and on both skills at once: raw mix 0.825 -> denoised mix 0.850 (+0.025), while the denoised singles are no better alone (arc 0.825 vs 0.838). The inbred-line signature: averaging within a line does not improve the line, it makes it cleaner to cross.

RETRACTION 1 — "a destructive skill propagates through the merge and destroys partners" is wrong. A single merge of clean single-skill adapters is excellent: arc alone 0.838 (boolq 0.000), boolq alone 0.800 (arc 0.300), merged 50/50 = arc 0.863 / boolq 0.787, mean 0.825 vs 0.550 for the best parent. Merging is protective — it stops either adapter dominating the output format. Consistent with LoRA Soups rather than contradicting it.

Probe 3 — iterated merging (/tmp/decay.py), five chained merges, three weight schemes. arc retention: convex [0.5,0.5] 1.02 · selfish [0.8,0.4] 1.02 · additive [1.0,1.0] 0.52 (arc 0.867 -> 0.450, incoming skills 0.033). RETRACTION 2 — geometric signal dilution is not the mechanism. Convex merging loses nothing over five rounds even though the first adapter's coefficient falls to 1/32. The scheme that preserves signal coefficients is the only one that collapses, because the accumulated change grows without bound and leaves the region where the base still functions. The operative constraint on iterated merging is bounding total drift from the base, not preserving signal.

Probe 4 — the scaling dose-response (/tmp/scale.py), prompted by GG asking the obvious control: does dividing the change vector by 30 retain the skill? Base (no adapter) 0.000; scale 1 0.838; 1/2 0.875; 1/4 0.850; 1/8 0.863; 1/16 0.450; 1/32 0.000; 1/30 0.000.

This forces a reinterpretation of Probe 3. At the coefficient arc actually held after five merges (1/32) the adapter alone delivers nothing. So the 0.883 measured in the chain was never arc's residual. What propagates through a merge is the answer format, supplied by whichever partner has enough weight to carry it: arc needs "a single letter", and it collapsed at round 4 (partner squad, free-text spans, arc 0.567) and recovered at round 5 (partner hellaswag, single letter A-D, arc 0.883). The chain measured format compatibility with the dominant partner, not skill retention.

What survives, and what it implies. (i) A sharp effectiveness threshold at ~1/8: an adapter works at full strength down to an eighth and collapses below it, so useful merge depth is ~3 rounds at convex weights, not 5. (ii) The additive collapse stands (it did not depend on the misreading). (iii) The curriculum result becomes coherent for the first time: the two families that destroyed everything, boolq and winogrande, are the ones with the most idiosyncratic output formats. If format is what propagates, a partner carrying a dominant format overwrites the ability to answer anything else. That is a claim about output conventions, not weight geometry — and it is consistent with Fig. 3C-D, where functional conflict predicted merge damage (rho 0.45) and weight geometry did not (0.03).

Still untested, and now the leading candidate for the curriculum collapse: continued training on top of merged weights. Probe 3 chained merges without ever training between them and lost nothing; the curriculum merges then trains, every generation. Test: repeat the chain with a fine-tune on the next family after each merge, and see whether that alone reproduces the collapse.

8d. Per-skill scaling thresholds, and the bespoke-weights negative (2026-09-08)

Dose-response per skill (/tmp/thresholds.py), accuracy vs adapter scale, base = 0.000 on all six:

skill 1 1/2 1/4 1/8 1/16 1/32 own optimum
arc 0.87 0.88 0.88 0.92 0.42 0.00 1/8
squad 0.72 0.73 0.78 0.65 0.10 0.00 1/4
hellaswag 0.75 0.82 0.80 0.63 0.07 0.00 1/2
boolq 0.78 0.78 0.78 0.37 0.00 0.00 11/4 (flat)
winogrande 0.55 0.55 0.55 0.37 0.30 0.00 11/4 (flat)
mnli 0.40 0.40 0.68 0.48 0.48 0.18 1/4

Three facts. (i) Thresholds are skill-specific (GG predicted this): boolq dies at 1/8 where arc, squad and hellaswag are still at full strength, so merge depth in a population is set by the weakest skill. (ii) Cliffs are sharp — full effectiveness right up to the edge, then near-total loss in one halving; there is no graceful degradation to trade against. (iii) Four of six skills are BETTER scaled down — mnli 0.40 -> 0.68 at 1/4, hellaswag 0.75 -> 0.82 at 1/2, squad 0.72 -> 0.78 at 1/4, arc 0.87 -> 0.92 at 1/8. These adapters are over-trained at full strength; attenuation recovers accuracy with no retraining.

Denoising does NOT move the threshold. arc raw 0.87/0.88/0.88/0.92/0.42/0.00 vs denoised 0.87/0.83/0.88/0.87/0.35/0.00; boolq raw 0.78/0.78/0.78/0.37/0.00 vs denoised 0.78/0.78/0.77/0.38/0.03. Identical cliffs. So the limit is signal magnitude, not signal-to-noise: averaging leaves signal at full strength and only removes noise, and therefore cannot buy merge depth. Denoising remains worth doing for cross-skill merge quality (+0.025, §8c) and for knowing what a line contains — not for depth.

Bespoke per-skill merge weights — tested and NEGATIVE (/tmp/bespoke.py). All six skills merged into one model:

scheme sum mean
six separate adapters, full strength 0.678
six separate, each at its own optimum 0.755
uniform convex (1/6) 1.00 0.708
uniform 0.25 1.50 0.686
bespoke: cliff (lowest viable per skill) 1.12 0.689
bespoke: optimum (best-accuracy per skill) 1.62 0.686
additive (1.0 each) 6.00 0.156

All three non-uniform schemes cluster at 0.6860.689, below plain equal weighting. Why the inference failed: solo dose-response curves do not transfer to the multi-way case. A skill's effective strength in a merge is set by its coefficient relative to the other five — six output formats compete for one model — so raising one skill's absolute weight starves the others. Clearest in squad: 0.65 at uniform 1/6, but 0.470.55 whenever given a larger absolute weight alongside others. The curves are sound; the inference from them to merge weights was not.

The two results worth keeping. (a) One merged model beats six separate specialists on their own tasks — 0.708 vs 0.678 — with mnli the clearest case (0.40 alone, 0.67 merged, because merging dilutes it to near its optimum and undoes the over-training). The merge is doing compression plus incidental regularisation, which is a more honest description than "combining capabilities". (b) The largest free win needs no merging at all: attenuating each specialist to its own optimum takes the separate-models baseline from 0.678 to 0.755, the best number in the table — one scalar per adapter, no retraining. Merging then costs ~5 points and saves five models: a real engineering trade, honestly stated.

8e. The last candidate eliminated — merge-then-train is the best procedure tested (2026-09-08)

Test (/tmp/trainmerge.py). Two chains, identical partners and order. Control: merge only. Test: merge, then continue-training on the partner's family (300) plus replay across everything seen (150 split) — i.e. the v5 sequence. The control reproduced the earlier chain exactly at all five rounds, so the comparison is clean.

round partner merge-only arc merge+train arc merge-only partner merge+train partner
1 boolq 0.883 0.933 0.733 0.850
2 winogrande 0.900 0.933 0.550 0.717
3 mnli 0.850 0.917 0.467 0.867
4 squad 0.567 0.883 0.733 0.767
5 hellaswag 0.883 0.833 0.800 0.850

Training on merged weights is not the mechanism — it is a substantial improvement. Better on the tracked skill in four of five rounds, better on the incoming skill in all five, and it absorbs the round-4 format shock that dropped the control to 0.567. Merge-then-train holds the old skill near its ceiling and acquires the new one far better than merging alone (mnli 0.867 vs 0.467).

All three proposed mechanisms for the v5 collapse are now refuted, each by direct test: a destructive skill propagating through merges (§8c — a single merge is protective); geometric signal dilution (§8c — five convex merges lose nothing); training on merged weights (here — it helps). The one structural difference left is that these chains merge clean single-skill adapters, whereas v5 merged accumulating lineages carrying up to six skills each in rank 16 — i.e. capacity. Recorded as unexplained rather than attributed: three mechanisms have been proposed and refuted, and a fourth guess would not have earned its place.

Consequence for the plan (§8b/§8c). The veto arm was scheduled to make the v5 negative interpretable. With three mechanisms refuted and merge-then-train shown sound in isolation, the v5 result is not reportable whatever a veto arm shows — recommendation is to stop, not to run it.

8f. The veto arm (GG overruled my recommendation to skip it — correctly) 2026-09-08

Change: identical to v5's society arm except that "keep the parent unchanged" is scored on the same validation split as the merge candidates, and wins if no weighting beats it. One bit of selection. Everything else — operator, weight grid, curriculum, replay, seeds — unchanged.

Seed 1 result: a single veto converts total collapse into a healthy trajectory.

generation 0 1 2 3 4 5
veto (declinable) 0.719 0.756 0.744 0.781 0.781 0.783
society (obligate) 0.689 0.719 0.700 0.597 0.439 0.211
isolated (never merges) 0.625 0.700 0.728 0.753 0.789 0.814

The veto decisions are structured, and this is the substantive finding. Merges declined, per generation (of 3): 1, 1, 1, 3, 3, 3 — 67% overall, and from generation 3 onward every lineage declines every merge, unanimously. Median validation gain when accepted: +0.042. That timing tracks complementarity, which is 1.00 through generation 1, 0.80 at 2, then 0.67 and falling: the population discovers on its own that recombination has stopped paying and stops — H6's predicted shape of the advantage, reached from the opposite direction.

GG's caveat, and it is the right reading (2026-09-08): once merges are always declined the arm is literally the isolated arm, so the comparison at the end is between "merged early, then stopped" and "never merged". Early merging gives a large lead (+0.094 at generation 0) but isolated overtakes at generation 4 and finishes higher (0.814 vs 0.783) — the early merges leave a residual cost that never-merging avoids. So the claim is not "the veto fixes recombination". It is: recombination pays only while partners differ, a population can detect when that stops, and even then it ends slightly behind never having merged.

Consequence for reportability. This makes the v5 negative interpretable and no longer vulnerable to "you forced merging, of course it broke": obligate recombination collapses (0.211), one bit of selection rescues it (0.783), and never recombining is still marginally best (0.814). Risk, remedy, and honest limit — the language-model rung beside Fig. 5A.

My error, recorded: I recommended skipping this experiment on the grounds that v5 was unreportable whatever it showed. That judged the experiment by whether it would rescue a conclusion I had already written off, rather than by what it would measure. The per-generation veto rate is information neither other arm could produce, and it is the most interesting thing in the arm.

Replication: seeds 2-3 submitted to CX3 as array 4007703 (hpc/llm_veto.pbs), pairing against the existing v5 isolated/society/seed_bank runs for those seeds.

Control worth considering if the seeds hold: a forced stop at generation 3, to separate "the veto's timing is smart" from "any early merging then stopping does this". The veto's stopping point coincides with complementarity falling below 0.8, which is principled rather than arbitrary — but that is an observation, not a test.

8g. The vocabulary substrate: contamination screen and the prior-art problem (2026-09-09)

GG's proposal: replace task families with content — teach 100 words of a language the model does not speak, one word-set per modifier, so thousands of words yield hundreds of modifiers and the generation count rises tenfold. Retention is trivially measurable ("what does X mean in English?").

Contamination screen — the first measurement was invalid. Generating an answer and string-matching it gave Italian 0.092, French 0.050, Basque 0.050, Welsh 0.017, Zulu 0.008, pseudo-words 0.000. French tying Basque is impossible if the quantity measured were knowledge, so the probe was measuring whether a base model obeys "answer with one English word" — the same instruction-following floor that gives 0.011 on the task families. Re-run as an 8-way forced choice over candidate translations scored by likelihood (domain-conditional PMI, chance 0.125), which needs no instruction-following:

language generated forced choice verdict
French 0.050 0.950 fully known
Italian 0.092 0.908 fully known
Welsh 0.017 0.508 half known
Basque 0.050 0.483 half known
Zulu 0.008 0.142 at chance — genuinely unknown
pseudo-words 0.000 0.158 floor (cycling 20 nouns inflates this slightly)

Only Zulu is clean among natural languages; pseudo-words are clean by construction and unlimited in supply. Probe: /tmp/contam2.py.

Prior art makes the experiment-as-framed a reproduction. WikiBigEdit (arXiv:2503.05683) runs 506K factual QA pairs across 8 sequential timesteps. Locate-then-edit methods (ROME, MEMIT) collapse within the first few hundred updates, but their LoRA + merging baseline — a fresh adapter per timestep, interpolated into the accumulated adapter at weight 0.25 — is stable across the whole benchmark and beats every dedicated editing method past ~100K updates. That is our isolated-plus- attenuated-merge arm, run three orders of magnitude further, and it does not collapse. Separately, arXiv:2506.14126 finds that over-training experts harms merging via late-stage memorisation, which is the published version of our §8d observation that attenuating four of six adapters was free gain — cite it, do not claim it.

Consequence for the diagnosis. If a single lineage accumulates 500K disjoint facts by fresh-adapter-plus-interpolation without collapsing, capacity is not what stopped v5 at generation 3-4 with six families. The difference between the two settings is that WikiBigEdit's content is homogeneous QA in one output format, whereas v5's families conflict at the output (label vs span vs number vs code). The remaining candidate is interference between competing output formats.

In the population-genetic frame the two are distinct: a new word-set is a new locus, and adding loci is cheap; two families demanding different output formats for the same input shape are competing alleles at one locus, and that is what collapses. A pure vocabulary curriculum is all loci and no allelic competition, so it would run to a hundred generations and confirm only that capacity is ample — removing precisely the variable that produced the phenomenon.

The collision sweep I proposed here is also occupied — do not run it either. In Praise of Stubbornness (arXiv:2502.04390) sweeps exactly this: non-contradictory updates integrate safely, while contradictory ones destroy up to 80% of unrelated knowledge with as few as 10-100 facts, consistently across model scales, and the authors conclude explicitly that the cause is conflict rather than capacity. Interference and Retention in Continual Learning (arXiv:2607.09202) supplies the theory: disjoint task supports make forgetting structurally eliminable, conflicting overlap imposes an unavoidable distortion floor. Both the measurement and its formalisation exist.

What this buys us anyway: the v5 collapse now has a cause. Three of our own explanations were retracted (§8c, §8d) and capacity was the standing suspect. Between 2502.04390 and 2607.09202 the mechanism is settled and it is allelic conflict, not locus exhaustion — the six families conflict at the output, and contradictory updates corrupt disproportionately and non-locally. WikiBigEdit is the mirror control: homogeneous single-format content accumulates to 506K facts without collapsing. The collapse we could not explain is a known, characterised, independently replicated phenomenon.

The gap that survives. Every merging paper in the landscape still merges once — GENOME (the ACL 2026 population-evolution paper) evolves a population toward a single target task, with no collapse, forgetting or grounding analysis. The iterated reproduction loop is still unoccupied, we have run it, and v5's negative answer (obligate recombination collapses by generation 3-4; veto-gated recombination merely matches isolation) is now interpretable rather than mysterious. No further LLM compute is required to state it.

8. Decisions for GG

  1. Three lineages × nine generations × nine families (complementarity maximal at t=3, zero at t=9), or fewer families and more generations per family?
  2. Replay budget — fixed n_replay split across all families seen so far (so per-family replay thins as the curriculum grows, which is realistic and makes forgetting a live pressure), or fixed per-family (constant protection, more compute)?
  3. Whether single_shot_merge gets the directed weight selection the society arms use, or plain uniform soup as published. I would give it the same selection, so the comparison isolates iteration rather than handing the society a free operator advantage.

8h. Two controls for the declinable merge, pre-registered before running (2026-09-11)

Both were identified in the 2026-09-11 manuscript review as the weakest hedges in the six-generation population section. Code: merge_until and orders config keys in src/llm/curriculum.py; configs curriculum_v5_stop3.yaml, curriculum_v5_decor.yaml; stats figures/stats_llm_curriculum.py.

Control 1 — forced stop at generation 3 (llm_curriculum_v5_stop3). The v5 society arm with recombination switched off from generation 3 (merge_until: 3, allow_veto: false). Rationale: in the seed-1 veto run lineages declined 1/3 of merges at generations 02 and 3/3 at 35, so this is the matched fixed schedule. Readout: best-lineage all-family accuracy at generation 5, paired per seed (3 seeds) against veto, isolated and society.

  • stop3 ≈ veto (within ±0.03 in every seed): the veto's outcome is explained by when it stopped; the paper keeps "the population found the schedule by itself" and drops any claim that per-decision evaluation adds value beyond timing.
  • stop3 < veto in every seed: the early declines avoided specific harmful merges; the modifier reading strengthens.
  • stop3 ≈ society (collapsed): three obligate merges already carry the format destroyers; stopping is not enough, screening is required.

Control 2 — decorrelated curriculum (llm_curriculum_v5_decor). Same six families, G = 6, but every lineage starts with mnli, then diverges maximally, then converges, so partner complementarity by generation is 0.00, 0.67, 0.70, 0.58, 0.33, 0.00 (Latin square: 1.00, 1.00, 0.80, 0.67, 0.33, 0.00). Arms: isolated and society with allow_veto: true. Primary readout, pooled over both curricula (2 × 6 generations × 3 seeds = 36 points of mean veto_used): partial Spearman correlation of the fraction declined with complementarity, controlling for generation (rank-regress both on generation, correlate residuals), seed-clustered bootstrap CI.

  • Modifier hypothesis: partial ρ(declined, complementarity | generation) < 0 with CI excluding 0.
  • Adapter-age hypothesis: that partial ρ ≈ 0 while partial ρ(declined, generation | complementarity) > 0.
  • Known ambiguity, stated in advance: at generation 0 of the new curriculum all lineages hold the same family from different training draws, and §8b measured that such same-skill merges gain +0.087 (encoding decorrelation). A low decline rate at generation 0 therefore does not test complementarity; the primary test is the pooled partial correlation, not that point.
  • Secondary, descriptive: whether the new curriculum's veto arm finishes level with its own isolated arm, as in the Latin square (0.792 vs 0.796).

Seed 1 of each control runs locally (batch 24 / train batch 2, as the v5 seed-1 runs); seeds 23 on CX3 (hpc/llm_curriculum_controls.pbs, batch 48 / train batch 4, as the v5 seeds 23 runs).

§8h outcome (2026-09-11, 3 seeds each; figures/stats_llm_curriculum.py)

  • Control 1, forced stop: stop3 0.793 vs veto 0.792 vs isolated 0.796 (per-seed veto stop3: 0.008, 0.006, +0.011). First branch: the veto's outcome is explained by when it stopped.
  • Control 2, decorrelated curriculum: partial ρ(declined, complementarity | generation) = 0.067, CI (0.211, +0.088); partial ρ(declined, generation | complementarity) = +0.31. Adapter-age branch: declines track generation, not complementarity. The Latin-square ρ = 0.57 was carried by generation. Decor veto 0.790 = decor isolated 0.790.
  • Manuscript consequence: the recombination-modifier / reduction-principle reading of Fig. 4B is withdrawn; the declinable merge remains the mechanism that avoided the obligate-merge collapse at no cost against never merging, and the forced-stop control shows a fixed schedule does the same.