MachineSex/tasks/prereg-llm-society-v2.md
Giorgio Gilestro 84124de143 Manuscript revision and pending experiment work, snapshot before restructuring
Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 16:54:09 +01:00

42 KiB
Raw Permalink Blame History

Pre-registration — llm_society v2: the composed society at LLM scale

Status (2026-09-07, 21:30): calibrated; NO-GO at 0.5B (GG, §12a); 7B plan in §13 awaiting scope. Supersedes the design in workorder-llm-society.md (v1). Calibration record and every amendment are in §4a; the campaign was not submitted. Read §1 (what v1 got wrong), §4a (what calibration found), §12a (the decision), §13 (what next).

Why a v2. The v1 campaign (3 CX3 seeds landed 2026-09-07, results/llm_society_campaign/; seed 1 still running locally) did not reproduce E11: no_grounding degraded (0.575 → 0.436, worst arm in every seed), but full also declined (→ 0.506) and no_sex was flattest (0.558). The conformitytruth gap did not separate the arms. Read against the framework, v1 had three structural faults that the theory would have predicted, plus one power fault. All four are diagnosed in §1 and designed out in §3. The point of this document is to make the remaining predictions explicit before spending the compute, so the campaign can fail informatively.


0. The question and the claims it tests

Does a finite population of LLM agents under the four composed operators — grounded evaluation, directed recombination, diversity-preserving selection, lossy inheritance — climb to capability that no founder had and hold it, while each ablation fails in its own way? This is E11 at the language-model tier: the paper's "open — the stated gap" cell (Fig. 1A).

Claims exercised, and the analytic experiment each rests on:

Claim Analytic source LLM prediction (§5)
Recombination assembles a genotype no parent had (vertical) E8 (FisherMuller, unbounded parents) H1
Blending conserves the single-parent level; only union realises the gain E4 (conservation law), llm_moe H2
Ungrounded selection → self-consumption → confident, unfit consensus E11 (conformity mechanism) H3
Without recombination, capability is capped at the best founder E8 control (ρ=1) + no mutation operator here H4
Greedy selection collapses diversity faster; QD holds it E5, E11 H5
A capability survives inheritance only if observed often enough E2 (per-item floor 1e^{mp}) H6

Explicit non-goal. Grounding here is E11's selection-channel grounding (g·fitness + (1g)·conformity), not E2's immigration into the training mix. No verified answer ever enters any child's training data, in any arm. Fig. 1A already states that LLM-scale training-mix grounding is established in prior work and not re-run; this campaign does not change that. A negative here is evidence against the selection mechanism only.


1. What v1 got wrong, read through the framework

# Fault What the theory says Evidence in v1 Fix (§3)
F1 Near-clone founders. 8 agents over 3 families → 3 lists-, 3 strings-, 2 arith-specialists differing only by task draws. E8 control: recombining ρ=1 parents buys nothing (flat at 6 for any K). Pigeonhole: 4 parents from 3 families always contains a same-family pair. no_sexfull: merging near-clones is pure perturbation cost. L = 12 disjoint families, one founder per family, ρ = 0 by construction; verified at gen 0.
F2 Combinatorial space too small. 3 skills → 2³ = 8 competence states; founders occupy 3 of them. E8/E11 use L = 12 (4096 genotypes). The vertical claim needs room for a child to be new. Best possible gain over a founder was tiny. L = 12 → the best founder holds 1/12 of the space.
F3 Blending operator in the dilution regime. Sex = 2-parent linear LoRA merge at 0.5B. E4 conservation law; llm_moe 0.5B: soup dilutes lists 0.43 → 0.26. Headroom law: dilution wherever there is room to lose. full declined while no_sex held. Reproduction by union-preserving recombination (confidence-routed union of parents' answers → distil). Linear merge kept as an explicit control arm (sex_linear) — H2.
F4 Truncation before breeding. Top-4 of 8 selected as parents; children bred only from them. E11 selects on survival over the pooled parents + offspring, never on breeding eligibility. Truncating parents discards half the alleles at gen 1 with no mutation operator to restore them (E6: loss is permanent). Half the families were unreachable after gen 1. Survival selection over the pool (§3.4). Elitism becomes emergent, as in E11.
F5 Underpowered evaluation. n_test = 40 → SE 0.079 per measurement. Every contrast except vs no_grounding sat inside one SE. n_test = 240 (20/family) → SE 0.032 overall.
F6 Weak grounding contrast. g = 0.5 vs E11's 0.85; G = 10 vs 80. The conformity gap in E11 needs the population to converge; g = 0.5 leaves conformity with half the vote even in full. Gap flat in all arms. g = 0.85; G = 12 (§7 explains why 12 suffices here).
F7 Transmission floor never measured. n_inherit = 600 over 3 families chosen by feel. E2: an item survives only if it is observed enough in the inheritance sample — the per-item floor. Pilot v1 measured a ~25%/gen "distillation tax" and fixed it by doubling data, without asking where the floor was. Calibration C2 measures the retention curve and sets n_inherit from it.

2. Theoretical predictions → design constraints

Each constraint below is derived, not chosen.

2.1 Decorrelation (E8, E4). Union coverage of K parents is U = ρq + (1ρ)(1(1q)^K); the gain over a single parent is proportional to (1ρ). Founders must therefore be as decorrelated as the task space allows: one family each, no shared training items, and the gen-0 behavioural-distance matrix must show no pair below 0.5 disagreement (gate C1c).

2.2 Combinatorial headroom (E8). With one family per founder, q = 1/L. The doubling bound for 2-parent recombination gives ≥ ⌈log₂ L⌉ = 4 generations to reach full coverage under lossless inheritance; with per-generation retention r per family the plateau is set by r, not L. So L = 12 gives headroom; G must exceed 4 by enough to see the plateau: G = 12.

2.3 Operator (E4, llm_moe, llm_directed). At 0.5B on unsaturated families the framework predicts linear blending dilutes and union preserves. The society's reproduction operator must be union-preserving or the experiment re-measures a known result. The union is implemented in the inheritance data, not in weight space: for each inheritance prompt the child learns the answer of whichever parent is more confident (mean token log-probability of its own answer). This is E4's max operator applied per item, it is verifier-free (legal in the no_grounding arm), and it is directed sex in E10's sense — mate choice by complementarity plus per-item selection. Gate C4 checks that confidence tracks competence (the routing precondition); gate C3 checks the union child beats the linear child on a single 2-founder cross before any campaign money is spent.

2.4 Selection acts on survival (E11). E11 pools N parents with n_off offspring and keeps the top N by score + novelty·λ. Reproducing that exactly gives: elitism for free (a strong parent survives by out-scoring its children), no gen-1 truncation, and a directly comparable selection intensity (keep 12 of 24 = top ½; E11 keeps 60 of 180 = top ⅓ — pre-noted as a difference).

2.5 Conformity must be decoupled from truth for H3 to be testable. E11 initialises random genotypes, so its consensus is uninformative at gen 0. In the LLM, consensus is the modal answer over agents. With one expert per family, 11 of 12 agents answer any given family's prompt at roughly base level, so the modal answer ≈ the base model's answer, and conformity rewards being base-like. Prediction: consensus accuracy at gen 0 ≈ base overall (gate C5 measures it; it must be < 0.35, i.e. well below the best founder's own-family accuracy, otherwise conformity is a truth proxy and the no_grounding arm cannot fail by the predicted mechanism — see §6 F-alt).

2.6 The inheritance floor (E2). A family survives distillation only if the child sees enough correct examples of it. Retention as a function of examples-per-family, r(k), is measurable (C2). Set k* = min k with r ≥ 0.85, n_inherit = L·k*. Consequence for a diluted skill at accuracy a: effective correct examples = a·k*, so skills are lost at the distillation step, not the merge step (H6, a mechanism prediction that distinguishes this account from "merging destroyed it").

2.7 Headroom, not size (the llm_moe_hard lesson). Families must be unsaturated for the specialist (≤ 0.9) and non-trivial for the base (≥ 0.05, ≤ 0.4) at 0.5B. C1 rejects families outside that band. The three existing easy families straddle it (base lists 0.15 / strings 0.15 / arith 0.53; specialists 0.43 / 1.00 / 0.91) — strings saturates, arith's base is high. Both are candidates for replacement, decided by C1 not by preference.

2.8 No cross-family conflict (E12, llm_epistasis, llm_speciation). Recombination helps on additive landscapes and hurts under functional conflict. Families must not share a prompt shape with different answer conventions (e.g. two families that both say "sort" meaning different orders). Gate C1b: pairwise confidence-weighted functional conflict between founders (the llm_epistasis measure) below the grid's no-conflict axis level.


3. Design

3.1 Base and agents. Qwen2.5-0.5B-Instruct (frozen). Agent = LoRA r=16 α=32 on attention + MLP projections. N = 12.

3.2 Task space. L = 12 disjoint procedurally generated families, exact-match verifier, chosen from the candidate list in §4 C1 by the calibration band. Fixed pools per family: test 20 (reporting only), val 10 (grounded selection signal), conformity 10 fresh per generation (prompts only). Inheritance pool: k* per family, fresh per generation.

3.3 Founders. Agent i is a specialist on family i, trained on 600 × 3 epochs of that family only, seed master·100 + i. Gate C1c: gen-0 pairwise behavioural distance ≥ 0.5 for every pair.

3.4 One generation.

  1. Produce & score. Every agent in the current population answers test, val, and conformity pools. Consensus = modal normalised answer per conformity prompt. fitness = val accuracy; conformity = agreement rate with consensus; novelty = mean behavioural distance to the rest. score = g·fitness + (1g)·conformity, g = 0.85 (0 in no_grounding).
  2. Mate. Form n_off = N pairs by complementarity (greedy maximum behavioural distance, each agent used at most twice), restricted to the current population.
  3. Recombine (sex ON). For each pair, each parent answers the inheritance pool once (with mean token log-prob). Union child data: per prompt, the answer of the more confident parent. In sex_linear: 6 Dirichlet-weighted linear LoRA merges screened on the arm's own signal (val fitness, or conformity in ungrounded arms), winner answers the pool. Sex OFF: the child's data is a single parent's answers (parent chosen by score-proportional sampling).
  4. Inherit. Each child is a fresh LoRA trained from the base on its data, epochs from C2.
  5. Survive. Pool = N parents + N children. Keep the top N by score + λ·novelty (λ = 0.3; λ = 0 in no_diversity). Scores for children computed after training, on the same pools.

3.5 Arms (5).

arm g recombination survival
full 0.85 union (confidence-routed) QD (λ=0.3)
no_grounding 0 union; routing unchanged (verifier-free) QD on conformity
no_sex 0.85 none (single-parent redistillation) QD
no_diversity 0.85 union greedy (λ=0)
sex_linear 0.85 2-parent linear merge, 6 screened candidates QD

sex_linear is the H2 control and is the first arm to drop if budget forces it (§9).

3.6 Generations and seeds. G = 12. Seeds 14 (fixed before running; not re-drawn). Reporting uses the verifier in every arm; the verifier enters the loop only where g > 0.


4. Calibration phase (local 16 GB GPU, ≈ 34 h, before any submission)

Each gate has a pre-set pass criterion. Fail → stop, fix, re-run the gate. No campaign until all pass.

Gate What Pass criterion Cost
C1a family band Base and specialist (600×3) accuracy on each of ~15 candidate families (100 test items each) Keep families with base ∈ [0.05, 0.40] and specialist ∈ [0.60, 0.90]; need ≥ 12 ~60 min
C1b no conflict Pairwise confidence-weighted functional conflict between the 12 founders (from llm_epistasis) Every pair below the compat axis level of the epistasis grid ~15 min
C1c decorrelation Gen-0 behavioural-distance matrix on 120 mixed prompts min pairwise disagreement ≥ 0.5 (with C1b)
C2 transmission floor Distil a child from a founder's own answers with k ∈ {25, 50, 100, 200} examples of its family (rest of the pool mixed), 2 and 3 epochs; measure retained own-family accuracy ÷ founder accuracy Choose k* = min k with retention ≥ 0.85 at the chosen epochs; if no k ≤ 200 passes, the design is infeasible at 0.5B — stop ~45 min
C3 operator One cross (two founders): union-distil child vs best-of-6 linear-merge-distil child; both families' accuracy Union child ≥ 0.85 × each parent on that parent's family; union ≥ linear on the minimum of the two. If linear ≥ union, H2 is already falsified — record it and reconsider the operator before the campaign ~20 min
C4 routing precondition For each founder: mean log-prob on own-family answers vs off-family; AUC AUC ≥ 0.7 for ≥ 10 of 12 founders (with C2)
C5 consensus anchoring Consensus accuracy over the 12 founders at gen 0, 120 prompts < 0.35 (conformity is not a truth proxy) ~5 min

C2 also fixes the cost model (§9) — n_inherit = 12·k*.


5. Pre-registered hypotheses, thresholds, falsifiers

Primary outcome metric: best-agent overall test accuracy at generation G (deployed capability, elite included), reported with the best newborn (child trained that generation) alongside, so a "climb" carried by a surviving founder is visible as such. Secondary: per-family accuracy of the best agent (the competence genotype), behavioural diversity, consensus accuracy, conformitytruth gap. Reference level B₀ = best founder overall at gen 0 (≈ (0.7 + 11·base)/12 ≈ 0.24 if base ≈ 0.2; measured, not assumed).

Prediction (from) Quantitative threshold Falsified if
H1 vertical climb E8: union recombination of decorrelated one-family founders assembles multi-family agents; sigmoidal, most of the climb in gens 15, plateau set by r full best-agent at G ≥ B₀ + 0.20 and best newborn at G ≥ B₀ + 0.15; best agent competent (≥ 0.6) on ≥ 6 of 12 families; in ≥ 3 of 4 seeds best-agent gain < 0.10 in ≥ 2 seeds
H2 operator E4 conservation law in the dilution regime full sex_linear ≥ 0.10 at G (paired, per seed); sex_linear best agent competent on ≤ 3 families sex_linearfull in ≥ 2 seeds. Pre-stated regime caveat: this ordering is predicted to invert at 7B on easy tasks (headroom law); a 7B follow-up would test that, not this.
H3 self-consumption E11 + §2.5: conformity rewards base-likeness; ungrounded selection regresses the population to base and homogenises it no_grounding best agent at G ≤ B₀ + 0.05; conformitytruth gap (no_grounding full) ≥ 0.30 at G; consensus accuracy in no_grounding non-increasing no_groundingfull 0.05 on best agent, or gap difference < 0.10
H4 sex necessity E8 ρ=1 control + no skill-acquisition operator without recombination no_sex best agent at G ≤ B₀ + 0.05 in every seed (a ceiling, stronger than E11's ~1-point effect) no_sex gains ≥ 0.10 over B₀ in any seed → an unmodelled acquisition route exists (base competence amplified by self-distillation); report it
H5 diversity E5/E11: greedy converges earliest; QD holds AUC of behavioural diversity full > no_diversity in ≥ 3 seeds; no_diversity diversity < 0.1 by gen ≤ 6. Low power on best fitness pre-declared (E11: 0.78 vs 0.74) no ordering in diversity AUC
H6 where skills die E2 floor: loss occurs at distillation when correct examples/family < k*·a For families lost between t and t+1 in full, the source (union answer set) accuracy on that family at t is ≥ 0.6 in ≤ 20% of cases — i.e. skills that were competently supplied are retained; skills die because they arrived diluted ≥ 40% of lost families were supplied at ≥ 0.6 → the distillation channel itself is lossy beyond the calibrated floor; revisit C2

Analysis is per-seed paired contrasts (4 seeds), reported as mean ± 95% CI and sign count. No metric introduced after unblinding is called a result. All rows above are also plotted whether or not they pass.


6. Anticipated failure modes and how each is handled

  • F-alt (conformity anchored to truth). If C5 shows consensus accuracy ≥ 0.35, the no_grounding arm may fail by drift rather than by confident-wrong consensus (v1 workorder falsifier 3). Then H3's gap threshold is not testable; run anyway, report the observed signature, and say so.
  • Terminal degeneration. A source that emits < 8 usable answers: copy the parent unchanged (existing sentinel). Count and report occurrences per arm.
  • Family extinction is permanent (E6). No mutation operator reintroduces a lost family. Pre-state: the number of families alive in the population is itself a reported curve; full is predicted to hold ≥ 10 of 12 to G, no_diversity fewer.
  • Router failure (C4 fails). Fall back to self-consistency routing (two samples, prefer the parent whose answers agree); re-run C4. If still failing, the union operator has no verifier-free implementation at this scale — record and consider 7B.
  • Elite lock-in. A founder that survives to G on score alone makes "best agent" flat. Best newborn is co-primary for exactly this reason.
  • Screening noise in sex_linear. 6 candidates on 120 val items (SE 0.046) — adequate for choosing among merges that differ by ≥ 0.1, which is the dilution scale.
  • Queue / wall-time loss. Per-generation checkpoint (rows flushed to parquet; adapters on disk) and --resume; each PBS array element = one (seed, arm), ≤ 6 h.
  • Environment drift. Never uv sync on a machine with a running job (tasks/lessons.md).

7. Power

SE of an overall accuracy at 240 items, p ≈ 0.5: 0.032. Per-family at 20 items: 0.11 (per-family readouts are descriptive only). Predicted effects: H1 ≥ 0.20, H2 ≥ 0.10, H3 ≥ 0.20 on best agent and ≥ 0.30 on the gap, H4 a ceiling — all ≥ 3 SE. H5 on best fitness is predicted small and is not powered; its diversity readout is. Four seeds give a sign test at p = 1/16 one-sided for a 4/4 outcome; the primary analysis is the paired mean and CI, the sign count is descriptive.

Why G = 12 suffices: the climb needs ⌈log₂ 12⌉ = 4 doublings; the E11 no_grounding crash occurred by generation 6 of 80; diversity collapse in E11's ablations by 1015. Twelve generations covers every predicted transition with margin to see the plateau.


8. Analysis and figure (fixed now)

Figure, E11 layout plus one panel: (A) best-agent and best-newborn overall accuracy per arm over generations, B₀ dashed; (B) behavioural diversity; (C) conformitytruth gap; (D) competence heat-map — families × generations for the full best agent, with sex_linear beside it. Mean ± 95% CI over seeds. Script figures/plot_llm_society.py, reading only the committed bundles.

Statistics: per-seed paired contrasts at G for H1H4; AUC contrast for H5; the supplied-vs-retained tabulation for H6. figures/stats_llm_society.py.


9. Compute and schedule

Measured anchor: v1 seed (N=8, G=10, 4 arms, n_inherit=600, 3 epochs) ≈ 3.3 L40S-hours.

Per generation-arm at v2 defaults (N=12, 24-agent pool eval on 480 prompts; 12 parents answering the inheritance pool once each with log-probs; 12 children trained at 12·k* × epochs):

k* from C2 training inheritance answers eval per gen-arm per (seed, arm), G=12 campaign, 4 seeds × 5 arms
50 ~6 min ~3 min ~4 min ~13 min ~2.6 h ~52 L40S-h
100 ~12 min ~5 min ~4 min ~21 min ~4.2 h ~84 L40S-h

Submitted as a 20-element PBS array (seed × arm), 6 h wall-time each, checkpointed. Dropping sex_linear saves 20%. The local GPU runs one seed's full + no_grounding in parallel as a hedge.

7B. ≈ 4× per operation → ~200340 L40S-h for the full grid: not a first shot. Pre-registered role for 7B: a 12 seed confirmation of full vs no_grounding and of the H2 inversion, run only if 0.5B passes H1 and H3. On the easy families 7B is predicted to be in the composition regime, so sex_linear should catch up with full there — a positive prediction of the headroom law, not a replication.

Schedule. Calibration day 1 (local). GG gate on C-results. Engineering (§10) days 12. Campaign submission day 23; wall-clock ≈ 1 day if the queue cooperates. Analysis + figure day 4.


10. Engineering checklist (before submission)

  • tasks.py: ~15 candidate families with verifier formats (int / int-list / lowercase word / uppercase word); FAMILIES becomes config-driven.
  • society.py: survival-over-pool selection replacing parent truncation; complementarity pairing over the whole population; confidence-routed union inheritance (needs generate() to return mean token log-prob); sex_linear arm; per-generation parquet flush + --resume; log the source per-family accuracy before distillation (for H6) and families-alive per generation.
  • Tests for the pure pieces (union routing, pooled survival, pairing constraint) — extend the 155.
  • configs/llm/society_v2_calib_*.yaml, society_v2_s{1..4}.yaml; hpc/llm_society_v2.pbs array.
  • figures/plot_llm_society.py, figures/stats_llm_society.py written before unblinding, against the smoke bundle.
  • Smoke: N=4, L=4, G=2, all 5 arms, exit 0, figure renders.

11. Outcome → manuscript

Outcome What changes in the paper
H1 ∧ H3 pass (H2, H4, H5 whatever they are) Fig. 1A cell "open — the stated gap" → filled; new figure (§8) enters as the LLM tier of the society; Table S2 row; the Discussion's prescriptive claim gains its LLM instantiation.
H1 fails, H2 passes The vertical claim does not transfer at 0.5B but the operator law does: report as a bounded negative in SI with the H6 diagnosis; Fig. 1A cell becomes "tested at 0.5B: operator law holds, climb does not"; 7B confirmation becomes the open item.
H3 fails (no self-consumption signature) with C5 passed The selection-channel grounding mechanism does not transfer; state it, keep the E11 result as biological-model-only; nothing prescriptive at LLM scale.
C-gates fail No campaign. The gate result itself goes in the SI as the reason the tier was not run.

12. Decisions (GG, 2026-09-07)

  1. Scale: 0.5B full grid; 7B only as the gated confirmation of §9. Decided.
  2. sex_linear arm: dropped from the first campaign — founders are cached and shared, so it can be appended later at ~20% of the grid cost, and the operator law is already established by llm_moe (Fig. 3B). The code path stays (arm_settings("sex_linear")); H2 is therefore deferred, not tested, in this campaign. Four arms × four seeds = 16 array elements.
  3. Family candidates: no vetoes; calibration C1 decides membership. Seventeen candidates are implemented in src/llm/families.py (the three originals + fourteen new). Word-order reversal and run-length encoding were dropped at implementation because the verifier cannot score multi-word or alphanumeric answers; sortletters, caesar and charfreq use random pseudo-words so a 600-item training set cannot cover the test space.
  4. Go/no-go gate after calibration — GG reviews the C-table before anything is submitted. Standing.

4a. Calibration record and amendments

Stage A, pass 1 (2026-09-07, results/llm_society_v2_calib_a, founders 600×3). Only 6 of 17 families in band: setops, numtheory, mixedtoken, vectors, digits, alphabet. Out of band: lists and binary under-trained (specialist 0.44); strings 0.91, roman 0.97, prime 1.00 above the specialist ceiling; arith base 0.53 (the base already knows it — violates the sole-expert premise); liststats and charfreq weak (0.38; charfreq also fails the routing AUC, 0.44); sortletters 0.20, caesar 0.01, progression 0.10 unlearnable at this budget. C1b max pairwise conflict 0.348 (just under the gate), every top pair involving a failed specialist answering confidently wrong; C1c min distance 0.65. Smoke passed on all four arms (figure + stats script exercised).

Amendments before pass 2 — recorded here because they change what §4 promised:

  1. Specialist upper bound 0.90 → 1.00. The bound encoded the llm_moe_hard headroom lesson, which concerns fusion composing to a ceiling so that soup matches routing. The society uses union inheritance, and the quantities under test are transmission and assembly, for which a founder at 0.97 is not a problem. The base bound (≤ 0.40) is kept strict: it protects the sole-expert premise and C5's consensus decoupling. Consequence: strings, roman, prime become eligible; arith stays out unless the count forces it (then flagged).
  2. Founder budget 600×3 → 1200×3, uniform, to recover the under-trained lists and binary.
  3. Prompt spaces enlarged for roman (1999), binary (1511), prime (≤ 400) so a 600-item training set cannot cover the test space (pass 1: 300 / 200 / 210 unique of 600).
  4. Three candidates addedwordlen, lettercount, sumeven (counting and filtered sums; verifier-safe; large prompt spaces). Three dropped without retraining: sortletters, caesar, progression (specialist ≤ 0.20).

None of these touches a hypothesis, a threshold, or the campaign design; they change which families are eligible. Pass 2 is configs/llm/society_v2_calib_a2.yaml.

Stage A, pass 2 (results/llm_society_v2_calib_a2, founders 1200×3). Nine in band: strings 1.00, setops 0.85, numtheory 0.83, mixedtoken 0.74, digits 0.97, alphabet 0.97, prime 0.96, wordlen 0.92, lettercount 0.77 (specialist accuracies; all bases 0.080.30). Roman misses only the base floor (base 0.04, specialist 0.99). The borderline families did not converge with more training — vectors 0.61 → 0.53, lists 0.44 → 0.50, binary 0.44 → 0.22, charfreq 0.38 → 0.55, liststats 0.38 → 0.47 — so they are noise-level at 0.5B, not under-trained; sumeven is unlearnable (0.06). C1c min distance 0.68 (pass). C4: 9 of 10 candidates have confidence AUC ≥ 0.7 (mixedtoken 0.62).

C1b fails among the ten. Max pairwise conflict 0.439 (wordlen × lettercount); five pairs ≥ 0.35, all among the counting families (wordlen, lettercount, digits, strings, mixedtoken): two specialists that both confidently answer "How many …?" with different small integers is precisely E12's conflicting-convention hazard, and under confidence-routed union the wrong one can win the prompt. The largest subset with every pair < 0.35 has seven members. Dropping lettercount (the hub, 7 conflicts) leaves nine with max 0.396.

Re-deriving the gate from the grid, rather than from a midpoint. 0.35 was chosen as midway between the grid's no-conflict axis (0.200.26) and its conflict axis (0.460.52). The grid's own outcome data say where the break is: P(merge penalty > 0.02) is 0.29 for epi_conf < 0.35, 0.22 for [0.35, 0.41), and 0.77 for ≥ 0.41 (n = 17 / 9 / 13). Below 0.41 the measure does not predict a penalty; above it, it does. A gate at 0.41 is therefore the data-derived boundary, and the nine-family set passes it (max 0.396).

Options put to GG (2026-09-07 evening):

  1. L = 9, gate 0.41 — strings, setops, numtheory, mixedtoken, digits, alphabet, prime, wordlen, roman. Two amendments: gate 0.35 → 0.41 (grid-derived, above), base floor waived for roman (the floor screened for unlearnable tasks; roman's specialist at 0.99 settles that). H1's family threshold scales to ≥ 5 of 9. Recommended.
  2. L = 7, gate 0.35 as written — alphabet, mixedtoken, numtheory, prime, roman, setops, strings. Weaker combinatorics (128 states, ~2.8 doublings) but no gate amendment.
  3. Prompt tags per family to suppress off-family confidence, then re-calibrate everything (~1.5 h). Removes the conflict by construction; also makes routing trivially lexical (the llm_moe rider).

GG decision (2026-09-07, 20:30): option 1. L = 9: strings, setops, numtheory, mixedtoken, digits, alphabet, prime, wordlen, roman. C1b gate amended 0.35 → 0.41 (grid-derived); base floor waived for roman. Consequential edits to §3/§5/§7: N = 9 agents; n_test 27/family (243 overall, SE 0.032); n_val 13/family; n_conf 13/family; H1's family threshold ≥ 5 of 9; B₀ ≈ (spec + 8·base)/9. Stage B launched on the nine (society_v2_calib_b.yaml; probe families setops / alphabet / digits; cross setops × alphabet).

Stage B, C2 transmission (results/llm_society_v2_calib_b) — FAILS as pre-registered. Retention of a founder's own family in a child distilled from the founder's own answers, k examples per family × 9 families, 2 or 3 epochs:

k setops alphabet digits
25 0.310.38 0.300.34 0.590.69
50 0.670.73 0.500.64 ~0.71
100 0.620.67 0.570.67
150 0.610.73 0.730.81

No k ≤ 150 reaches 0.85; the curve is flattening. This is not the E2 observation floor — the source supplied the family at 0.771.00 accuracy, so the items were observed. It is interference: a one-family founder's inheritance data is one competent family and eight families of confident garbage, and a fresh LoRA fits all nine. That is the mechanism behind pilot v1's "distillation tax", now measured at 2040% per generation. Under the pre-registered rule the ungated design is infeasible at 0.5B.

Proposed amendment (measured before adoption, C2b — society_v2_calib_c2b.yaml): confidence-gated inheritance. The child learns only the prompts its source is confident on (exp mean token log-prob ≥ τ). Verifier-free; identical in every arm; makes the child agnostic rather than wrong off-expertise — E8's founder model. C2b measures retention by τ, the Youden τ* separating own- from off-family confidence (calibration uses family labels; the campaign uses the fixed τ), and the off-family harm of ungated inheritance. Adoption requires GG's sign-off because it changes §3.4 step 4. Implemented as conf_gate in society_v2.py (default None = ungated).

C2b (results/llm_society_v2_calib_b_transmission_conf) — the gate passes at τ = 0.5, and the mechanism is two-part. Source answers a 300-per-family pool (2700 prompts); child keeps prompts with source confidence ≥ τ; 3 epochs. Retention (child own-family ÷ founder):

ungated τ = 0.5 τ = 0.7 τ = 0.85
setops 0.69 0.93 0.93 0.94
alphabet 0.86 0.85 0.88 0.82
digits 0.89 0.83 0.76
mean 0.81 0.87 0.86

Two drivers, not one. (i) Pool size: at k = 300 alphabet and digits already retain ≥ 0.86 ungated (they were 0.730.81 at k = 150). (ii) Confidence gating rescues the family pool size does not (setops 0.69 → 0.93) and is neutral-to-slightly-negative where the specialist is confident everywhere (digits: off-family confidence median 0.67 vs 0.45 for setops — the gate discards data without separating). Youden τ* ≈ 0.950.97 in all three (own-family confidence median 1.00), so τ = 0.5 is a mild gate keeping 5075% of the pool. Off-family harm: not confirmed. Ungated children score at base off-family (0.190.27 vs base 0.200.22), not below it; gated children slightly above. Interference costs own-family retention; it does not push off-expertise competence below the prior.

Adopted for the campaign (pending GG go/no-go on the full table): k_inherit = 300, conf_gate = 0.5, epochs = 3. Mean retention 0.87 ≥ 0.85 meets the pre-registered C2 criterion under the amended channel. Cost consequence (§9): per-generation inheritance answers 2700 prompts per parent and children train on ~13002000 kept examples → ≈ 30 min per generation-arm, ≈ 6 h per (seed, arm) at G = 12, ≈ 96 L40S-h for 16 elements; PBS walltime raised to 8 h.

Stage B, C3 cross setops × alphabet (results/llm_society_v2_calib_b_cross) — operator half passes, retention half fails for C2's reason. Parents 0.84 / 0.96. Union child 0.57 / 0.62 — holds both families, routed 53% of prompts to the alphabet parent. Linear-blend child (best of 6 screened) 0.66 / 0.30 — keeps one family and loses the other: E4's dilution, in the operator the campaign dropped. Union ≥ linear on the minimum (0.57 vs 0.30) ✓. Union ≥ 0.85 × parent ✗ (0.68× / 0.65×) — the same transmission tax as C2. The cross is re-run under the confidence gate if C2b passes (conf_gate is now a cross-stage option; both children gated by their own source's confidence).

C3 re-run under the gate (results/llm_society_v2_calib_b_cross_gated, k = 300, τ = 0.5) — fails for a NEW reason. Union child setops 0.46 / alphabet 0.73 (0.55× / 0.76× of parents) — both held, both diluted. Linear child 0.80 / 0.22 — one skill at 0.95×, the other lost. Read together with C2b (a one-skill child retains 0.870.93 under the same gate and budget): a single skill transmits; a two-skill child dilutes each skill by 2545% even from union-preserved data. This is E4's conservation law relocated from the sample budget to the learning budget — a fixed adapter (r = 16) and fixed epochs split across skills. Under it H1 (≥ 5 families at ≥ 0.6 in one agent) is predicted to fail by construction, whatever the operators do. F8 in the fault ledger.

C3b (before deciding anything): does capacity or budget lift the two-skill child? Three variants of the gated cross — 6 epochs at r = 16; r = 64 (α = 128) at 3 epochs; both. Pass criterion as C3: union child ≥ 0.85 × parent on each family. If one passes, the campaign adopts that child budget (cost re-estimated). If none passes, the vertical claim cannot be tested at 0.5B with self-distilled inheritance, and the honest options are a reduced campaign (H3H5 only, which do not need multi-skill children) or 7B.

C3b results (results/llm_society_v2_calib_c3b_*; union child accuracy and ×parent):

child budget setops alphabet verdict
r16, 3 ep, τ 0.5 (C3 gated) 0.46 (0.55×) 0.73 (0.76×) fail
r16, 6 ep, τ 0.5 0.73 (0.87×) 0.79 (0.82×) at the gate within noise (SE ≈ 0.06 on the ratio)
r64, 3 ep, τ 0.5 0.72 (0.90×) 0.58 (0.67×) fail; r64 founders weaker (0.80/0.87) and less confident (35% routed to alphabet)
r64, 6 ep, τ 0.5 0.35 (0.44×) 0.50 (0.57×) fail — overfits

Budget, not capacity, is the lever; rank stays 16. The kept count explains the residual: at τ = 0.5 the union child kept 1941 of 2700 prompts, of which only ~600 are its two competent families — the mild gate passes the max of two parents' confidences, so ~70% of the child's data is confident garbage. C2b's own table had τ = 0.85 as the best mean retention (0.88) at a third of the data. C3c (last calibration run): the cross at τ = 0.85, 3 and 6 epochs.

Prediction update carried into the go/no-go, whatever C3c says. A two-skill child retains ≈ 0.85× per skill at best; H1 as written (best agent ≥ B₀ + 0.20, ≥ 5 of 9 families at ≥ 0.6) needs five or six skills co-resident at ≈ 0.6 in one r = 16 adapter, which the calibration does not support at 0.5B. The realistic bar the calibration does support — H1: children holding 23 families beat every founder on overall accuracy (≥ B₀ + 0.05) and the best agent climbs monotonically for ≥ 3 generations — is recorded now, before the campaign, as the primary vertical readout, with H1 kept as the stretch criterion. H3, H4, H5 do not need multi-skill agents and are unchanged.

C3c results (results/llm_society_v2_calib_c3c_*, τ = 0.85): 3 epochs → union child 0.72 / 0.79 (0.86× / 0.82×), kept 1041 of 2700; 6 epochs → 0.69 / 0.75. The tight gate reproduces the 6-epoch mild-gate retention at half the training, and more budget beyond that buys nothing: a two-skill child plateaus at ≈ 0.85× / 0.8× of its parents. Final child budget: τ = 0.85, 3 epochs, k = 300, r = 16. Cost re-estimate (§9): inheritance answers (9 parents × 2700 prompts with log-probs) now dominate at ≈ 11 min per generation-arm; training ≈ 9 min; evaluation ≈ 3 min → ≈ 25 min per generation-arm, ≈ 5 h per (seed, arm), ≈ 80 L40S-h for 16 elements. Walltime 8 h.

Stage B, C5 consensus (results/llm_society_v2_calib_b_consensus) — passes. Consensus accuracy over the nine founders at gen 0 = 0.31 (< 0.35); corr(conformity, own accuracy) = +0.29; min pairwise distance 0.63. Conformity is not a truth proxy here, so H3 is testable by the predicted mechanism.

12a. Go/no-go (GG, 2026-09-07, 21:30): NO-GO at 0.5B; plan 7B.

Grounds: the calibration passed C1, C4, C5, and — under the amended inheritance channel — C2, but C3's retention half exposed a ceiling no budget moves: a two-skill child holds each skill at ≈ 0.85× / 0.8× of its parents, and the vertical claim needs five or six skills co-resident in one r = 16 adapter. At 0.5B the society experiment could test H3H5 but not the claim the paper's stated gap is about. Nothing is submitted. What today produced is a measured transmission ceiling for self-distilled LoRA inheritance at 0.5B — three mechanistically distinct limits (near-clone founders; interference from confident off-expertise answers, 2040%/generation, removable by a confidence gate; the multi-skill learning-budget plateau) — and it goes in the SI as the reason the 0.5B tier was not run (§11, row 4).

13. The 7B plan (for GG review; nothing runs without a go)

What changes at 7B, and why it is not a re-run. The premise "one founder is the sole expert on its family" requires a base that cannot do the family. Qwen2.5-7B-Instruct already scores 0.99 on easy arith, 0.69 on easy strings, and will be high on roman / binary / setops / digits; the base bound (≤ 0.40) will exclude most of the current nine. The hard variants exist for three families only (7B base: lists 0.34, strings 0.67, arith 0.48). Phase 0 is therefore task design: ≥ 9 disjoint families with 7B base ≤ 0.4 and specialist ≥ 0.8 — multi-step, cipher, and compositional variants of the current generators — plus the C1 band and conflict gate re-run at 7B. This is a day of work before any GPU time, and it cannot be done on the 16 GB local card (7B training needs the L40S), so every calibration step goes through the CX3 queue (183 queued at last check).

Cost anchors (L40S, from llm_merge_hpc / llm_moe_hpc / llm_hard): 7B generation ≈ 10 prompts/s (0.5B ≈ 40); 7B LoRA SFT ≈ 15 example-passes/s (0.5B ≈ 60). Per generation-arm at N = 9, k = 300, τ = 0.85, 3 epochs: inheritance answers 9 × 2700 / 10 ≈ 40 min; evaluation 18 × 373 / 10 ≈ 11 min; training 9 × (≈ 1000 × 3 / 15) ≈ 31 min → ≈ 80 min per generation-arm.

scope elements G per element total L40S-h walltime
Phase 1 — calibration A + B at 7B 2 jobs ~2 h each ~4 4 h
Headline: full vs no_grounding, 3 seeds 6 10 ~13 h ~80 2 × 8 h with resume, or one 16 h
H3 + H4: 3 arms, 3 seeds 9 10 ~13 h ~120 as above
Full grid: 4 arms, 4 seeds 16 12 ~16 h ~260 3 × 8 h with resume, or one 24 h

Checkpoint/resume already makes multi-requeue elements safe. A vLLM generation path would cut the dominant 40-minute term by 510× but adds a dependency and a second code path; noted, not proposed.

Gates carried over unchanged: C1 band (base ≤ 0.40 strict, specialist ≥ 0.60; upper bound 1.0), C1b conflict < 0.41, C1c distance ≥ 0.5, C2/C2b retention ≥ 0.85 (gated channel), C3 union ≥ 0.85× per family on a two-founder cross, C4 AUC ≥ 0.7 for ≥ 8 of 9, C5 consensus < 0.35. The 7B-specific prediction that decides whether to proceed past Phase 1: with the larger adapter margin at 7B, the two-skill cross should clear 0.85× on both families at 3 epochs. If it does not, the multi-skill plateau is not a 0.5B artefact and the vertical claim should be pursued with a different inheritance channel (e.g. inheriting weights, not answers — which is what llm_merge_hpc already showed composes at 7B) rather than with more scale.

Hypotheses: H1 restored as written (≥ 5 of 9 families at ≥ 0.6, ≥ B₀ + 0.20) — that is the point of going to 7B; H1 kept as the fallback readout; H3H6 unchanged; H2 deferred.

Decisions for GG before Phase 0 starts: (i) scope row from the table; (ii) whether Phase 0 task design is worth the day, given the alternative in the prediction paragraph above; (iii) whether the SI text for the 0.5B ceiling (§11 row 4) is drafted now or after 7B.

14. Build log

  • 2026-09-07 — families.py (17 candidates, all self-verifying and deterministic), society_ops.py (pooled survival, capped complementary mating, confidence-routed union, score-proportional single parent), society_v2.py (kind: llm_society_v2; per-generation checkpoint + resume; founder lock for concurrent arm-jobs; H6 source diagnostics; families-alive), calibrate.py (kind: llm_society_calib, stages families / transmission / cross / consensus), configs (society_v2_calib_a/b, society_v2_smoke), hpc/llm_society_v2.pbs (16-element seed × arm array), figures/plot_llm_society.py (the §8 layout, written before unblinding), 9 new pure tests (164 green). v1 code path untouched and still green. Smoke → calibration A launched locally.