MachineSex/tasks/prereg-llm-society-v2.md
Giorgio Gilestro 84124de143 Manuscript revision and pending experiment work, snapshot before restructuring
Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 16:54:09 +01:00

563 lines
42 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Pre-registration — `llm_society` v2: the composed society at LLM scale
**Status (2026-09-07, 21:30): calibrated; NO-GO at 0.5B (GG, §12a); 7B plan in §13 awaiting scope.**
Supersedes the design in `workorder-llm-society.md` (v1). Calibration record and every amendment are
in §4a; the campaign was not submitted. Read §1 (what v1 got wrong), §4a (what calibration found),
§12a (the decision), §13 (what next).
**Why a v2.** The v1 campaign (3 CX3 seeds landed 2026-09-07, `results/llm_society_campaign/`;
seed 1 still running locally) did not reproduce E11: `no_grounding` degraded (0.575 → 0.436, worst
arm in every seed), but `full` also declined (→ 0.506) and `no_sex` was flattest (0.558). The
conformitytruth gap did not separate the arms. Read against the framework, v1 had three
*structural* faults that the theory would have predicted, plus one power fault. All four are
diagnosed in §1 and designed out in §3. The point of this document is to make the remaining
predictions explicit *before* spending the compute, so the campaign can fail informatively.
---
## 0. The question and the claims it tests
Does a finite population of LLM agents under the four composed operators — grounded evaluation,
directed recombination, diversity-preserving selection, lossy inheritance — climb to capability that
no founder had and hold it, while each ablation fails in its own way? This is E11 at the language-model
tier: the paper's "open — the stated gap" cell (Fig. 1A).
Claims exercised, and the analytic experiment each rests on:
| Claim | Analytic source | LLM prediction (§5) |
|---|---|---|
| Recombination assembles a genotype no parent had (vertical) | E8 (FisherMuller, unbounded parents) | H1 |
| Blending conserves the single-parent level; only union realises the gain | E4 (conservation law), `llm_moe` | H2 |
| Ungrounded selection → self-consumption → confident, unfit consensus | E11 (conformity mechanism) | H3 |
| Without recombination, capability is capped at the best founder | E8 control (ρ=1) + no mutation operator here | H4 |
| Greedy selection collapses diversity faster; QD holds it | E5, E11 | H5 |
| A capability survives inheritance only if observed often enough | E2 (per-item floor 1e^{mp}) | H6 |
**Explicit non-goal.** Grounding here is E11's *selection-channel* grounding
(`g·fitness + (1g)·conformity`), not E2's *immigration into the training mix*. No verified answer
ever enters any child's training data, in any arm. Fig. 1A already states that LLM-scale training-mix
grounding is established in prior work and not re-run; this campaign does not change that. A
negative here is evidence against the *selection* mechanism only.
---
## 1. What v1 got wrong, read through the framework
| # | Fault | What the theory says | Evidence in v1 | Fix (§3) |
|---|---|---|---|---|
| F1 | **Near-clone founders.** 8 agents over 3 families → 3 lists-, 3 strings-, 2 arith-specialists differing only by task draws. | E8 control: recombining ρ=1 parents buys **nothing** (flat at 6 for any K). Pigeonhole: 4 parents from 3 families always contains a same-family pair. | `no_sex``full`: merging near-clones is pure perturbation cost. | L = 12 disjoint families, **one founder per family**, ρ = 0 by construction; verified at gen 0. |
| F2 | **Combinatorial space too small.** 3 skills → 2³ = 8 competence states; founders occupy 3 of them. | E8/E11 use L = 12 (4096 genotypes). The vertical claim needs room for a child to be *new*. | Best possible gain over a founder was tiny. | L = 12 → the best founder holds 1/12 of the space. |
| F3 | **Blending operator in the dilution regime.** Sex = 2-parent *linear* LoRA merge at 0.5B. | E4 conservation law; `llm_moe` 0.5B: soup dilutes lists 0.43 → 0.26. Headroom law: dilution wherever there is room to lose. | `full` declined while `no_sex` held. | Reproduction by **union-preserving recombination** (confidence-routed union of parents' answers → distil). Linear merge kept as an explicit control arm (`sex_linear`) — H2. |
| F4 | **Truncation before breeding.** Top-4 of 8 selected as parents; children bred only from them. | E11 selects on *survival over the pooled parents + offspring*, never on breeding eligibility. Truncating parents discards half the alleles at gen 1 with no mutation operator to restore them (E6: loss is permanent). | Half the families were unreachable after gen 1. | Survival selection over the pool (§3.4). Elitism becomes emergent, as in E11. |
| F5 | **Underpowered evaluation.** `n_test` = 40 → SE 0.079 per measurement. | — | Every contrast except vs `no_grounding` sat inside one SE. | `n_test` = 240 (20/family) → SE 0.032 overall. |
| F6 | **Weak grounding contrast.** g = 0.5 vs E11's 0.85; G = 10 vs 80. | The conformity gap in E11 needs the population to converge; g = 0.5 leaves conformity with half the vote even in `full`. | Gap flat in all arms. | g = 0.85; G = 12 (§7 explains why 12 suffices here). |
| F7 | **Transmission floor never measured.** `n_inherit` = 600 over 3 families chosen by feel. | E2: an item survives only if it is *observed* enough in the inheritance sample — the per-item floor. Pilot v1 measured a ~25%/gen "distillation tax" and fixed it by doubling data, without asking where the floor was. | — | Calibration C2 measures the retention curve and sets `n_inherit` from it. |
---
## 2. Theoretical predictions → design constraints
Each constraint below is derived, not chosen.
**2.1 Decorrelation (E8, E4).** Union coverage of K parents is U = ρq + (1ρ)(1(1q)^K); the
gain over a single parent is proportional to (1ρ). Founders must therefore be as decorrelated as the
task space allows: one family each, no shared training items, and the gen-0 behavioural-distance
matrix must show no pair below 0.5 disagreement (gate C1c).
**2.2 Combinatorial headroom (E8).** With one family per founder, q = 1/L. The doubling bound for
2-parent recombination gives ≥ ⌈log₂ L⌉ = 4 generations to *reach* full coverage under lossless
inheritance; with per-generation retention r per family the plateau is set by r, not L. So L = 12
gives headroom; G must exceed 4 by enough to see the plateau: G = 12.
**2.3 Operator (E4, `llm_moe`, `llm_directed`).** At 0.5B on unsaturated families the framework
predicts linear blending dilutes and union preserves. The society's reproduction operator must be
union-preserving or the experiment re-measures a known result. The union is implemented in the
*inheritance data*, not in weight space: for each inheritance prompt the child learns the answer of
whichever parent is more confident (mean token log-probability of its own answer). This is E4's
`max` operator applied per item, it is verifier-free (legal in the `no_grounding` arm), and it is
directed sex in E10's sense — mate choice by complementarity plus per-item selection. Gate C4 checks
that confidence tracks competence (the routing precondition); gate C3 checks the union child beats the
linear child on a single 2-founder cross before any campaign money is spent.
**2.4 Selection acts on survival (E11).** E11 pools N parents with n_off offspring and keeps the top
N by `score + novelty·λ`. Reproducing that exactly gives: elitism for free (a strong parent survives
by out-scoring its children), no gen-1 truncation, and a directly comparable selection intensity
(keep 12 of 24 = top ½; E11 keeps 60 of 180 = top ⅓ — pre-noted as a difference).
**2.5 Conformity must be decoupled from truth for H3 to be testable.** E11 initialises random
genotypes, so its consensus is uninformative at gen 0. In the LLM, consensus is the modal answer
over agents. With one expert per family, 11 of 12 agents answer any given family's prompt at roughly
base level, so the modal answer ≈ the base model's answer, and conformity rewards *being base-like*.
Prediction: consensus accuracy at gen 0 ≈ base overall (gate C5 measures it; it must be < 0.35, i.e.
well below the best founder's own-family accuracy, otherwise conformity is a truth proxy and the
`no_grounding` arm cannot fail by the predicted mechanism see §6 F-alt).
**2.6 The inheritance floor (E2).** A family survives distillation only if the child sees enough
*correct* examples of it. Retention as a function of examples-per-family, r(k), is measurable
(C2). Set k* = min k with r 0.85, `n_inherit` = L·k*. Consequence for a *diluted* skill at accuracy
a: effective correct examples = a·k*, so skills are lost at the distillation step, not the merge step
(H6, a mechanism prediction that distinguishes this account from "merging destroyed it").
**2.7 Headroom, not size (the `llm_moe_hard` lesson).** Families must be unsaturated for the
specialist (≤ 0.9) and non-trivial for the base (≥ 0.05, 0.4) at 0.5B. C1 rejects families outside
that band. The three existing easy families straddle it (base lists 0.15 / strings 0.15 / arith 0.53;
specialists 0.43 / 1.00 / 0.91) strings saturates, arith's base is high. Both are candidates for
replacement, decided by C1 not by preference.
**2.8 No cross-family conflict (E12, `llm_epistasis`, `llm_speciation`).** Recombination helps on
additive landscapes and hurts under functional conflict. Families must not share a prompt shape with
different answer conventions (e.g. two families that both say "sort" meaning different orders). Gate
C1b: pairwise confidence-weighted functional conflict between founders (the `llm_epistasis` measure)
below the grid's no-conflict axis level.
---
## 3. Design
**3.1 Base and agents.** Qwen2.5-0.5B-Instruct (frozen). Agent = LoRA r=16 α=32 on attention + MLP
projections. N = 12.
**3.2 Task space.** L = 12 disjoint procedurally generated families, exact-match verifier, chosen from
the candidate list in §4 C1 by the calibration band. Fixed pools per family: test 20 (reporting only),
val 10 (grounded selection signal), conformity 10 fresh per generation (prompts only). Inheritance pool:
k* per family, fresh per generation.
**3.3 Founders.** Agent i is a specialist on family i, trained on 600 × 3 epochs of that family only,
seed `master·100 + i`. Gate C1c: gen-0 pairwise behavioural distance 0.5 for every pair.
**3.4 One generation.**
1. *Produce & score.* Every agent in the current population answers test, val, and conformity pools.
Consensus = modal normalised answer per conformity prompt. `fitness` = val accuracy; `conformity`
= agreement rate with consensus; `novelty` = mean behavioural distance to the rest.
`score = g·fitness + (1g)·conformity`, g = 0.85 (0 in `no_grounding`).
2. *Mate.* Form n_off = N pairs by complementarity (greedy maximum behavioural distance, each agent
used at most twice), restricted to the current population.
3. *Recombine (sex ON).* For each pair, each parent answers the inheritance pool once (with mean
token log-prob). Union child data: per prompt, the answer of the more confident parent. In
`sex_linear`: 6 Dirichlet-weighted linear LoRA merges screened on the arm's own signal (val
fitness, or conformity in ungrounded arms), winner answers the pool. Sex OFF: the child's data is a
single parent's answers (parent chosen by score-proportional sampling).
4. *Inherit.* Each child is a fresh LoRA trained from the base on its data, `epochs` from C2.
5. *Survive.* Pool = N parents + N children. Keep the top N by `score + λ·novelty` (λ = 0.3; λ = 0 in
`no_diversity`). Scores for children computed after training, on the same pools.
**3.5 Arms (5).**
| arm | g | recombination | survival |
|---|---|---|---|
| `full` | 0.85 | union (confidence-routed) | QD (λ=0.3) |
| `no_grounding` | 0 | union; routing unchanged (verifier-free) | QD on conformity |
| `no_sex` | 0.85 | none (single-parent redistillation) | QD |
| `no_diversity` | 0.85 | union | greedy (λ=0) |
| `sex_linear` | 0.85 | 2-parent linear merge, 6 screened candidates | QD |
`sex_linear` is the H2 control and is the first arm to drop if budget forces it 9).
**3.6 Generations and seeds.** G = 12. Seeds 14 (fixed before running; not re-drawn). Reporting
uses the verifier in every arm; the verifier enters the *loop* only where g > 0.
---
## 4. Calibration phase (local 16 GB GPU, ≈ 34 h, before any submission)
Each gate has a pre-set pass criterion. Fail → stop, fix, re-run the gate. No campaign until all pass.
| Gate | What | Pass criterion | Cost |
|---|---|---|---|
| **C1a** family band | Base and specialist (600×3) accuracy on each of ~15 candidate families (100 test items each) | Keep families with base ∈ [0.05, 0.40] and specialist ∈ [0.60, 0.90]; need ≥ 12 | ~60 min |
| **C1b** no conflict | Pairwise confidence-weighted functional conflict between the 12 founders (from `llm_epistasis`) | Every pair below the `compat` axis level of the epistasis grid | ~15 min |
| **C1c** decorrelation | Gen-0 behavioural-distance matrix on 120 mixed prompts | min pairwise disagreement ≥ 0.5 | (with C1b) |
| **C2** transmission floor | Distil a child from a founder's *own* answers with k ∈ {25, 50, 100, 200} examples of its family (rest of the pool mixed), 2 and 3 epochs; measure retained own-family accuracy ÷ founder accuracy | Choose k* = min k with retention ≥ 0.85 at the chosen epochs; if no k ≤ 200 passes, the design is infeasible at 0.5B — stop | ~45 min |
| **C3** operator | One cross (two founders): union-distil child vs best-of-6 linear-merge-distil child; both families' accuracy | Union child ≥ 0.85 × each parent on that parent's family; union ≥ linear on the *minimum* of the two. If linear ≥ union, H2 is already falsified — record it and reconsider the operator before the campaign | ~20 min |
| **C4** routing precondition | For each founder: mean log-prob on own-family answers vs off-family; AUC | AUC ≥ 0.7 for ≥ 10 of 12 founders | (with C2) |
| **C5** consensus anchoring | Consensus accuracy over the 12 founders at gen 0, 120 prompts | < 0.35 (conformity is not a truth proxy) | ~5 min |
C2 also fixes the cost model 9) `n_inherit` = 12·k*.
---
## 5. Pre-registered hypotheses, thresholds, falsifiers
Primary outcome metric: **best-agent overall test accuracy** at generation G (deployed capability,
elite included), reported with the **best newborn** (child trained that generation) alongside, so a
"climb" carried by a surviving founder is visible as such. Secondary: per-family accuracy of the best
agent (the competence genotype), behavioural diversity, consensus accuracy, conformitytruth gap.
Reference level **B₀** = best founder overall at gen 0 (≈ (0.7 + 11·base)/12 0.24 if base 0.2;
measured, not assumed).
| | Prediction (from) | Quantitative threshold | Falsified if |
|---|---|---|---|
| **H1** vertical climb | E8: union recombination of decorrelated one-family founders assembles multi-family agents; sigmoidal, most of the climb in gens 15, plateau set by r | `full` best-agent at G B + 0.20 and best newborn at G B + 0.15; best agent competent (≥ 0.6) on 6 of 12 families; in 3 of 4 seeds | best-agent gain < 0.10 in 2 seeds |
| **H2** operator | E4 conservation law in the dilution regime | `full` `sex_linear` 0.10 at G (paired, per seed); `sex_linear` best agent competent on 3 families | `sex_linear` `full` in 2 seeds. *Pre-stated regime caveat:* this ordering is predicted to **invert** at 7B on easy tasks (headroom law); a 7B follow-up would test that, not this. |
| **H3** self-consumption | E11 + §2.5: conformity rewards base-likeness; ungrounded selection regresses the population to base and homogenises it | `no_grounding` best agent at G B + 0.05; conformitytruth gap (`no_grounding` `full`) 0.30 at G; consensus accuracy in `no_grounding` non-increasing | `no_grounding` `full` 0.05 on best agent, **or** gap difference < 0.10 |
| **H4** sex necessity | E8 ρ=1 control + no skill-acquisition operator without recombination | `no_sex` best agent at G B + 0.05 in every seed (a *ceiling*, stronger than E11's ~1-point effect) | `no_sex` gains 0.10 over B in any seed an unmodelled acquisition route exists (base competence amplified by self-distillation); report it |
| **H5** diversity | E5/E11: greedy converges earliest; QD holds | AUC of behavioural diversity `full` > `no_diversity` in ≥ 3 seeds; `no_diversity` diversity < 0.1 by gen 6. **Low power on best fitness pre-declared** (E11: 0.78 vs 0.74) | no ordering in diversity AUC |
| **H6** where skills die | E2 floor: loss occurs at distillation when correct examples/family < ka | For families lost between t and t+1 in `full`, the *source* (union answer set) accuracy on that family at t is 0.6 in 20% of cases i.e. skills that were competently supplied are retained; skills die because they arrived diluted | 40% of lost families were supplied at 0.6 the distillation channel itself is lossy beyond the calibrated floor; revisit C2 |
Analysis is per-seed paired contrasts (4 seeds), reported as mean ± 95% CI and sign count. No metric
introduced after unblinding is called a result. All rows above are also plotted whether or not they
pass.
---
## 6. Anticipated failure modes and how each is handled
- **F-alt (conformity anchored to truth).** If C5 shows consensus accuracy 0.35, the `no_grounding`
arm may fail by drift rather than by confident-wrong consensus (v1 workorder falsifier 3). Then H3's
gap threshold is not testable; run anyway, report the observed signature, and say so.
- **Terminal degeneration.** A source that emits < 8 usable answers: copy the parent unchanged (existing
sentinel). Count and report occurrences per arm.
- **Family extinction is permanent (E6).** No mutation operator reintroduces a lost family. Pre-state:
the number of families alive in the population is itself a reported curve; `full` is predicted to
hold 10 of 12 to G, `no_diversity` fewer.
- **Router failure (C4 fails).** Fall back to self-consistency routing (two samples, prefer the parent
whose answers agree); re-run C4. If still failing, the union operator has no verifier-free
implementation at this scale record and consider 7B.
- **Elite lock-in.** A founder that survives to G on score alone makes "best agent" flat. Best newborn
is co-primary for exactly this reason.
- **Screening noise in `sex_linear`.** 6 candidates on 120 val items (SE 0.046) adequate for choosing
among merges that differ by 0.1, which is the dilution scale.
- **Queue / wall-time loss.** Per-generation checkpoint (rows flushed to parquet; adapters on disk)
and `--resume`; each PBS array element = one (seed, arm), 6 h.
- **Environment drift.** Never `uv sync` on a machine with a running job (`tasks/lessons.md`).
---
## 7. Power
SE of an overall accuracy at 240 items, p 0.5: 0.032. Per-family at 20 items: 0.11 (per-family
readouts are descriptive only). Predicted effects: H1 0.20, H2 0.10, H3 0.20 on best agent and
0.30 on the gap, H4 a ceiling all 3 SE. H5 on best fitness is predicted small and is not powered;
its diversity readout is. Four seeds give a sign test at p = 1/16 one-sided for a 4/4 outcome; the
primary analysis is the paired mean and CI, the sign count is descriptive.
Why G = 12 suffices: the climb needs log 12 = 4 doublings; the E11 `no_grounding` crash occurred
by generation 6 of 80; diversity collapse in E11's ablations by 1015. Twelve generations covers
every predicted transition with margin to see the plateau.
---
## 8. Analysis and figure (fixed now)
Figure, E11 layout plus one panel: **(A)** best-agent and best-newborn overall accuracy per arm over
generations, B dashed; **(B)** behavioural diversity; **(C)** conformitytruth gap; **(D)** competence
heat-map families × generations for the `full` best agent, with `sex_linear` beside it. Mean ± 95% CI
over seeds. Script `figures/plot_llm_society.py`, reading only the committed bundles.
Statistics: per-seed paired contrasts at G for H1H4; AUC contrast for H5; the supplied-vs-retained
tabulation for H6. `figures/stats_llm_society.py`.
---
## 9. Compute and schedule
Measured anchor: v1 seed (N=8, G=10, 4 arms, n_inherit=600, 3 epochs) 3.3 L40S-hours.
Per generation-arm at v2 defaults (N=12, 24-agent pool eval on 480 prompts; 12 parents answering the
inheritance pool once each with log-probs; 12 children trained at 12·k* × epochs):
| k* from C2 | training | inheritance answers | eval | per gen-arm | per (seed, arm), G=12 | campaign, 4 seeds × 5 arms |
|---|---|---|---|---|---|---|
| 50 | ~6 min | ~3 min | ~4 min | ~13 min | ~2.6 h | **~52 L40S-h** |
| 100 | ~12 min | ~5 min | ~4 min | ~21 min | ~4.2 h | **~84 L40S-h** |
Submitted as a 20-element PBS array (seed × arm), 6 h wall-time each, checkpointed. Dropping
`sex_linear` saves 20%. The local GPU runs one seed's `full` + `no_grounding` in parallel as a hedge.
**7B.** 4× per operation ~200340 L40S-h for the full grid: not a first shot. Pre-registered role
for 7B: a **12 seed confirmation of `full` vs `no_grounding` and of the H2 inversion**, run only if
0.5B passes H1 and H3. On the *easy* families 7B is predicted to be in the composition regime, so
`sex_linear` should catch up with `full` there a positive prediction of the headroom law, not a
replication.
**Schedule.** Calibration day 1 (local). GG gate on C-results. Engineering 10) days 12. Campaign
submission day 23; wall-clock 1 day if the queue cooperates. Analysis + figure day 4.
---
## 10. Engineering checklist (before submission)
- [ ] `tasks.py`: ~15 candidate families with verifier formats (int / int-list / lowercase word /
uppercase word); `FAMILIES` becomes config-driven.
- [ ] `society.py`: survival-over-pool selection replacing parent truncation; complementarity pairing
over the whole population; confidence-routed union inheritance (needs `generate()` to return
mean token log-prob); `sex_linear` arm; per-generation parquet flush + `--resume`; log the
*source* per-family accuracy before distillation (for H6) and families-alive per generation.
- [ ] Tests for the pure pieces (union routing, pooled survival, pairing constraint) extend the 155.
- [ ] `configs/llm/society_v2_calib_*.yaml`, `society_v2_s{1..4}.yaml`; `hpc/llm_society_v2.pbs` array.
- [ ] `figures/plot_llm_society.py`, `figures/stats_llm_society.py` written **before** unblinding,
against the smoke bundle.
- [ ] Smoke: N=4, L=4, G=2, all 5 arms, exit 0, figure renders.
---
## 11. Outcome → manuscript
| Outcome | What changes in the paper |
|---|---|
| H1 H3 pass (H2, H4, H5 whatever they are) | Fig. 1A cell "open the stated gap" filled; new figure 8) enters as the LLM tier of the society; Table S2 row; the Discussion's prescriptive claim gains its LLM instantiation. |
| H1 fails, H2 passes | The vertical claim does not transfer at 0.5B but the operator law does: report as a bounded negative in SI with the H6 diagnosis; Fig. 1A cell becomes "tested at 0.5B: operator law holds, climb does not"; 7B confirmation becomes the open item. |
| H3 fails (no self-consumption signature) with C5 passed | The selection-channel grounding mechanism does not transfer; state it, keep the E11 result as biological-model-only; nothing prescriptive at LLM scale. |
| C-gates fail | No campaign. The gate result itself goes in the SI as the reason the tier was not run. |
---
## 12. Decisions (GG, 2026-09-07)
1. **Scale:** 0.5B full grid; 7B only as the gated confirmation of §9. *Decided.*
2. **`sex_linear` arm:** **dropped from the first campaign** founders are cached and shared, so it
can be appended later at ~20% of the grid cost, and the operator law is already established by
`llm_moe` (Fig. 3B). The code path stays (`arm_settings("sex_linear")`); H2 is therefore
*deferred*, not tested, in this campaign. Four arms × four seeds = 16 array elements.
3. **Family candidates:** no vetoes; calibration C1 decides membership. Seventeen candidates are
implemented in `src/llm/families.py` (the three originals + fourteen new). Word-order reversal and
run-length encoding were dropped at implementation because the verifier cannot score multi-word or
alphanumeric answers; `sortletters`, `caesar` and `charfreq` use random pseudo-words so a 600-item
training set cannot cover the test space.
4. **Go/no-go gate after calibration** GG reviews the C-table before anything is submitted.
*Standing.*
## 4a. Calibration record and amendments
**Stage A, pass 1 (2026-09-07, `results/llm_society_v2_calib_a`, founders 600×3).** Only **6 of 17**
families in band: setops, numtheory, mixedtoken, vectors, digits, alphabet. Out of band: lists and
binary under-trained (specialist 0.44); strings 0.91, roman 0.97, prime 1.00 above the specialist
ceiling; arith base 0.53 (the base already knows it violates the sole-expert premise); liststats and
charfreq weak (0.38; charfreq also fails the routing AUC, 0.44); sortletters 0.20, caesar 0.01,
progression 0.10 unlearnable at this budget. C1b max pairwise conflict 0.348 (just under the gate),
every top pair involving a *failed* specialist answering confidently wrong; C1c min distance 0.65.
Smoke passed on all four arms (figure + stats script exercised).
**Amendments before pass 2 — recorded here because they change what §4 promised:**
1. **Specialist upper bound 0.90 → 1.00.** The bound encoded the `llm_moe_hard` headroom lesson,
which concerns *fusion composing to a ceiling* so that soup matches routing. The society uses
union inheritance, and the quantities under test are transmission and assembly, for which a
founder at 0.97 is not a problem. The **base** bound (≤ 0.40) is kept strict: it protects the
sole-expert premise and C5's consensus decoupling. Consequence: strings, roman, prime become
eligible; arith stays out unless the count forces it (then flagged).
2. **Founder budget 600×3 → 1200×3, uniform**, to recover the under-trained lists and binary.
3. **Prompt spaces enlarged** for roman (1999), binary (1511), prime (≤ 400) so a 600-item training
set cannot cover the test space (pass 1: 300 / 200 / 210 unique of 600).
4. **Three candidates added** `wordlen`, `lettercount`, `sumeven` (counting and filtered sums;
verifier-safe; large prompt spaces). Three dropped without retraining: sortletters, caesar,
progression (specialist 0.20).
None of these touches a hypothesis, a threshold, or the campaign design; they change which families
are *eligible*. Pass 2 is `configs/llm/society_v2_calib_a2.yaml`.
**Stage A, pass 2 (`results/llm_society_v2_calib_a2`, founders 1200×3).** Nine in band: strings 1.00,
setops 0.85, numtheory 0.83, mixedtoken 0.74, digits 0.97, alphabet 0.97, prime 0.96, wordlen 0.92,
lettercount 0.77 (specialist accuracies; all bases 0.080.30). Roman misses only the base *floor*
(base 0.04, specialist 0.99). The borderline families did **not** converge with more training
vectors 0.61 0.53, lists 0.44 0.50, binary 0.44 0.22, charfreq 0.38 0.55, liststats
0.38 0.47 so they are noise-level at 0.5B, not under-trained; sumeven is unlearnable (0.06).
C1c min distance 0.68 (pass). C4: 9 of 10 candidates have confidence AUC 0.7 (mixedtoken 0.62).
**C1b fails among the ten.** Max pairwise conflict 0.439 (wordlen × lettercount); five pairs 0.35,
all among the *counting* families (wordlen, lettercount, digits, strings, mixedtoken): two specialists
that both confidently answer "How many …?" with different small integers is precisely E12's
conflicting-convention hazard, and under confidence-routed union the wrong one can win the prompt.
The largest subset with every pair < 0.35 has **seven** members. Dropping lettercount (the hub, 7
conflicts) leaves nine with max 0.396.
**Re-deriving the gate from the grid, rather than from a midpoint.** 0.35 was chosen as midway between
the grid's no-conflict axis (0.200.26) and its conflict axis (0.460.52). The grid's own outcome data
say where the break is: P(merge penalty > 0.02) is **0.29 for epi_conf < 0.35, 0.22 for
[0.35, 0.41), and 0.77 for 0.41** (n = 17 / 9 / 13). Below 0.41 the measure does not predict a
penalty; above it, it does. A gate at **0.41** is therefore the data-derived boundary, and the
nine-family set passes it (max 0.396).
**Options put to GG (2026-09-07 evening):**
1. *L = 9, gate 0.41* strings, setops, numtheory, mixedtoken, digits, alphabet, prime, wordlen,
roman. Two amendments: gate 0.35 0.41 (grid-derived, above), base floor waived for roman (the
floor screened for unlearnable tasks; roman's specialist at 0.99 settles that). H1's family
threshold scales to 5 of 9. *Recommended.*
2. *L = 7, gate 0.35 as written* alphabet, mixedtoken, numtheory, prime, roman, setops, strings.
Weaker combinatorics (128 states, ~2.8 doublings) but no gate amendment.
3. *Prompt tags per family* to suppress off-family confidence, then re-calibrate everything (~1.5 h).
Removes the conflict by construction; also makes routing trivially lexical (the `llm_moe` rider).
**GG decision (2026-09-07, 20:30): option 1.** L = 9: strings, setops, numtheory, mixedtoken, digits,
alphabet, prime, wordlen, roman. C1b gate amended 0.35 0.41 (grid-derived); base floor waived for
roman. Consequential edits to §357: N = 9 agents; `n_test` 27/family (243 overall, SE 0.032);
`n_val` 13/family; `n_conf` 13/family; H1's family threshold 5 of 9; B (spec + 8·base)/9.
Stage B launched on the nine (`society_v2_calib_b.yaml`; probe families setops / alphabet / digits;
cross setops × alphabet).
**Stage B, C2 transmission (`results/llm_society_v2_calib_b`) — FAILS as pre-registered.** Retention of
a founder's own family in a child distilled from the founder's own answers, k examples per family × 9
families, 2 or 3 epochs:
| k | setops | alphabet | digits |
|---|---|---|---|
| 25 | 0.310.38 | 0.300.34 | 0.590.69 |
| 50 | 0.670.73 | 0.500.64 | ~0.71 |
| 100 | 0.620.67 | 0.570.67 | |
| 150 | 0.610.73 | 0.730.81 | |
No k 150 reaches 0.85; the curve is flattening. This is **not** the E2 observation floor the
source *supplied* the family at 0.771.00 accuracy, so the items were observed. It is interference: a
one-family founder's inheritance data is one competent family and eight families of confident
garbage, and a fresh LoRA fits all nine. That is the mechanism behind pilot v1's "distillation tax",
now measured at 2040% per generation. Under the pre-registered rule the ungated design is infeasible
at 0.5B.
**Proposed amendment (measured before adoption, C2b `society_v2_calib_c2b.yaml`):
confidence-gated inheritance.** The child learns only the prompts its source is confident on
(exp mean token log-prob τ). Verifier-free; identical in every arm; makes the child *agnostic*
rather than *wrong* off-expertise E8's founder model. C2b measures retention by τ, the Youden τ*
separating own- from off-family confidence (calibration uses family labels; the campaign uses the
fixed τ), and the off-family harm of ungated inheritance. Adoption requires GG's sign-off because it
changes §3.4 step 4. Implemented as `conf_gate` in `society_v2.py` (default None = ungated).
**C2b (`results/llm_society_v2_calib_b_transmission_conf`) the gate passes at τ = 0.5, and the
mechanism is two-part.** Source answers a 300-per-family pool (2700 prompts); child keeps prompts with
source confidence τ; 3 epochs. Retention (child own-family ÷ founder):
| | ungated | τ = 0.5 | τ = 0.7 | τ = 0.85 |
|---|---|---|---|---|
| setops | 0.69 | **0.93** | 0.93 | 0.94 |
| alphabet | 0.86 | 0.85 | 0.88 | 0.82 |
| digits | 0.89 | 0.83 | 0.76 | |
| mean | 0.81 | **0.87** | 0.86 | |
Two drivers, not one. (i) **Pool size**: at k = 300 alphabet and digits already retain 0.86
ungated (they were 0.730.81 at k = 150). (ii) **Confidence gating** rescues the family pool size does
not (setops 0.69 0.93) and is neutral-to-slightly-negative where the specialist is confident
everywhere (digits: off-family confidence median 0.67 vs 0.45 for setops the gate discards data
without separating). Youden τ* 0.950.97 in all three (own-family confidence median 1.00), so
τ = 0.5 is a *mild* gate keeping 5075% of the pool. **Off-family harm: not confirmed.** Ungated children
score at base off-family (0.190.27 vs base 0.200.22), not below it; gated children slightly above.
Interference costs own-family retention; it does not push off-expertise competence below the prior.
**Adopted for the campaign (pending GG go/no-go on the full table): `k_inherit` = 300,
`conf_gate` = 0.5, `epochs` = 3.** Mean retention 0.87 0.85 meets the pre-registered C2 criterion
under the amended channel. Cost consequence 9): per-generation inheritance answers 2700 prompts per
parent and children train on ~13002000 kept examples 30 min per generation-arm, 6 h per
(seed, arm) at G = 12, ** 96 L40S-h** for 16 elements; PBS walltime raised to 8 h.
**Stage B, C3 cross setops × alphabet (`results/llm_society_v2_calib_b_cross`) operator half
passes, retention half fails for C2's reason.** Parents 0.84 / 0.96. Union child 0.57 / 0.62 holds
*both* families, routed 53% of prompts to the alphabet parent. Linear-blend child (best of 6 screened)
0.66 / 0.30 keeps one family and loses the other: E4's dilution, in the operator the campaign
dropped. Union linear on the minimum (0.57 vs 0.30) ✓. Union 0.85 × parent (0.68× / 0.65×) the
same transmission tax as C2. The cross is re-run under the confidence gate if C2b passes
(`conf_gate` is now a cross-stage option; both children gated by their own source's confidence).
**C3 re-run under the gate (`results/llm_society_v2_calib_b_cross_gated`, k = 300, τ = 0.5) fails
for a NEW reason.** Union child setops 0.46 / alphabet 0.73 (0.55× / 0.76× of parents) both held,
both diluted. Linear child 0.80 / 0.22 one skill at 0.95×, the other lost. Read together with C2b
(a *one*-skill child retains 0.870.93 under the same gate and budget): a single skill transmits; a
**two-skill child dilutes each skill by 2545% even from union-preserved data.** This is E4's
conservation law relocated from the sample budget to the *learning* budget a fixed adapter (r = 16)
and fixed epochs split across skills. Under it H1 (≥ 5 families at 0.6 in one agent) is predicted to
fail by construction, whatever the operators do. F8 in the fault ledger.
**C3b (before deciding anything): does capacity or budget lift the two-skill child?** Three variants
of the gated cross 6 epochs at r = 16; r = 64 (α = 128) at 3 epochs; both. Pass criterion as C3:
union child 0.85 × parent on *each* family. If one passes, the campaign adopts that child budget
(cost re-estimated). If none passes, the vertical claim cannot be tested at 0.5B with self-distilled
inheritance, and the honest options are a reduced campaign (H3H5 only, which do not need multi-skill
children) or 7B.
**C3b results (`results/llm_society_v2_calib_c3b_*`; union child accuracy and ×parent):**
| child budget | setops | alphabet | verdict |
|---|---|---|---|
| r16, 3 ep, τ 0.5 (C3 gated) | 0.46 (0.55×) | 0.73 (0.76×) | fail |
| **r16, 6 ep, τ 0.5** | **0.73 (0.87×)** | **0.79 (0.82×)** | at the gate within noise (SE 0.06 on the ratio) |
| r64, 3 ep, τ 0.5 | 0.72 (0.90×) | 0.58 (0.67×) | fail; r64 founders weaker (0.80/0.87) and less confident (35% routed to alphabet) |
| r64, 6 ep, τ 0.5 | 0.35 (0.44×) | 0.50 (0.57×) | fail overfits |
**Budget, not capacity, is the lever; rank stays 16.** The kept count explains the residual: at τ = 0.5
the union child kept 1941 of 2700 prompts, of which only ~600 are its two competent families the
mild gate passes the *max* of two parents' confidences, so ~70% of the child's data is confident
garbage. C2b's own table had τ = 0.85 as the best mean retention (0.88) at a third of the data.
**C3c** (last calibration run): the cross at τ = 0.85, 3 and 6 epochs.
**Prediction update carried into the go/no-go, whatever C3c says.** A two-skill child retains 0.85×
per skill at best; H1 as written (best agent B + 0.20, 5 of 9 families at 0.6) needs five or
six skills co-resident at 0.6 in one r = 16 adapter, which the calibration does not support at 0.5B.
The realistic bar the calibration *does* support **H1: children holding 23 families beat every
founder on overall accuracy (≥ B + 0.05) and the best agent climbs monotonically for 3
generations** is recorded now, before the campaign, as the primary vertical readout, with H1 kept
as the stretch criterion. H3, H4, H5 do not need multi-skill agents and are unchanged.
**C3c results (`results/llm_society_v2_calib_c3c_*`, τ = 0.85):** 3 epochs union child 0.72 / 0.79
(0.86× / 0.82×), kept 1041 of 2700; 6 epochs 0.69 / 0.75. The tight gate reproduces the 6-epoch
mild-gate retention at half the training, and more budget beyond that buys nothing: a two-skill child
plateaus at 0.85× / 0.8× of its parents. **Final child budget: τ = 0.85, 3 epochs, k = 300, r = 16.**
Cost re-estimate 9): inheritance answers (9 parents × 2700 prompts with log-probs) now dominate at
11 min per generation-arm; training 9 min; evaluation 3 min 25 min per generation-arm,
5 h per (seed, arm), ** 80 L40S-h** for 16 elements. Walltime 8 h.
**Stage B, C5 consensus (`results/llm_society_v2_calib_b_consensus`) — passes.** Consensus accuracy over
the nine founders at gen 0 = 0.31 (< 0.35); corr(conformity, own accuracy) = +0.29; min pairwise
distance 0.63. Conformity is not a truth proxy here, so H3 is testable by the predicted mechanism.
## 12a. Go/no-go (GG, 2026-09-07, 21:30): **NO-GO at 0.5B; plan 7B.**
Grounds: the calibration passed C1, C4, C5, and under the amended inheritance channel C2, but C3's
retention half exposed a ceiling no budget moves: a two-skill child holds each skill at 0.85× / 0.8×
of its parents, and the vertical claim needs five or six skills co-resident in one r = 16 adapter. At
0.5B the society experiment could test H3H5 but not the claim the paper's stated gap is about.
Nothing is submitted. What today produced is a **measured transmission ceiling for self-distilled LoRA
inheritance at 0.5B** three mechanistically distinct limits (near-clone founders; interference from
confident off-expertise answers, 2040%/generation, removable by a confidence gate; the multi-skill
learning-budget plateau) and it goes in the SI as the reason the 0.5B tier was not run 11, row 4).
## 13. The 7B plan (for GG review; nothing runs without a go)
**What changes at 7B, and why it is not a re-run.** The premise "one founder is the sole expert on its
family" requires a base that *cannot* do the family. Qwen2.5-7B-Instruct already scores 0.99 on easy
arith, 0.69 on easy strings, and will be high on roman / binary / setops / digits; the base bound
(≤ 0.40) will exclude most of the current nine. The hard variants exist for three families only (7B
base: lists 0.34, strings 0.67, arith 0.48). **Phase 0 is therefore task design**: 9 disjoint
families with 7B base 0.4 and specialist 0.8 multi-step, cipher, and compositional variants of
the current generators plus the C1 band and conflict gate re-run at 7B. This is a day of work before
any GPU time, and it cannot be done on the 16 GB local card (7B training needs the L40S), so every
calibration step goes through the CX3 queue (183 queued at last check).
**Cost anchors (L40S, from `llm_merge_hpc` / `llm_moe_hpc` / `llm_hard`):** 7B generation 10
prompts/s (0.5B 40); 7B LoRA SFT 15 example-passes/s (0.5B 60). Per generation-arm at N = 9,
k = 300, τ = 0.85, 3 epochs: inheritance answers 9 × 2700 / 10 40 min; evaluation 18 × 373 / 10
11 min; training 9 × (≈ 1000 × 3 / 15) 31 min ** 80 min per generation-arm**.
| scope | elements | G | per element | total L40S-h | walltime |
|---|---|---|---|---|---|
| Phase 1 calibration A + B at 7B | 2 jobs | | ~2 h each | **~4** | 4 h |
| Headline: `full` vs `no_grounding`, 3 seeds | 6 | 10 | ~13 h | **~80** | 2 × 8 h with resume, or one 16 h |
| H3 + H4: 3 arms, 3 seeds | 9 | 10 | ~13 h | **~120** | as above |
| Full grid: 4 arms, 4 seeds | 16 | 12 | ~16 h | **~260** | 3 × 8 h with resume, or one 24 h |
Checkpoint/resume already makes multi-requeue elements safe. A vLLM generation path would cut the
dominant 40-minute term by 510× but adds a dependency and a second code path; noted, not proposed.
**Gates carried over unchanged:** C1 band (base 0.40 strict, specialist 0.60; upper bound 1.0), C1b
conflict < 0.41, C1c distance 0.5, C2/C2b retention 0.85 (gated channel), C3 union 0.85× per
family on a two-founder cross, C4 AUC 0.7 for 8 of 9, C5 consensus < 0.35. **The 7B-specific
prediction that decides whether to proceed past Phase 1:** with the larger adapter margin at 7B, the
two-skill cross should clear 0.85× on *both* families at 3 epochs. If it does not, the multi-skill
plateau is not a 0.5B artefact and the vertical claim should be pursued with a different inheritance
channel (e.g. inheriting *weights*, not answers which is what `llm_merge_hpc` already showed
composes at 7B) rather than with more scale.
**Hypotheses:** H1 restored as written (≥ 5 of 9 families at 0.6, B + 0.20) that is the point of
going to 7B; H1 kept as the fallback readout; H3H6 unchanged; H2 deferred.
**Decisions for GG before Phase 0 starts:** (i) scope row from the table; (ii) whether Phase 0 task
design is worth the day, given the alternative in the prediction paragraph above; (iii) whether the
SI text for the 0.5B ceiling 11 row 4) is drafted now or after 7B.
## 14. Build log
- 2026-09-07 `families.py` (17 candidates, all self-verifying and deterministic), `society_ops.py`
(pooled survival, capped complementary mating, confidence-routed union, score-proportional single
parent), `society_v2.py` (`kind: llm_society_v2`; per-generation checkpoint + resume; founder lock
for concurrent arm-jobs; H6 source diagnostics; families-alive), `calibrate.py`
(`kind: llm_society_calib`, stages families / transmission / cross / consensus), configs
(`society_v2_calib_a/b`, `society_v2_smoke`), `hpc/llm_society_v2.pbs` (16-element seed × arm
array), `figures/plot_llm_society.py` (the §8 layout, written before unblinding), 9 new pure tests
(164 green). v1 code path untouched and still green. Smoke calibration A launched locally.