# Work order — `llm_society`: the composed society at LLM scale (C3, the paper's stated gap) **Date opened:** 2026-09-07. **Decision (GG):** run it before PNAS submission — "if it's an experiment that can be done in weeks rather than months then any reviewer would ask to see it." ## Question Does a population of LLM agents under the four composed operators — grounded evaluation, directed recombination, diversity-preserving selection, retraining (mutation) — climb and hold capability, while each ablation fails distinctly? The LLM instantiation of E11; the paper's largest stated gap (Fig. 1's "open" cell). ## Design (mirrors E11; one grounding knob, in the evaluation channel) Population of `N` agents = LoRA adapters on frozen Qwen2.5-0.5B-Instruct, initialised as light per-family specialists (round-robin over families, agent-specific seeds → initial diversity). Non-overlapping generations (Wright–Fisher style; no elitism). Per generation: 1. **Produce & score.** Every agent answers (a) a fixed *validation* pool (verifier-scored → true fitness; the selection signal for grounded arms), (b) a fresh per-generation *conformity* pool (prompts only). Consensus = modal normalised answer per prompt; agent conformity = agreement with the consensus. A fixed *test* pool (never selected on) gives the reported fitness. 2. **Select** `P = N/2` parents by `s = g·fitness + (1−g)·conformity`. Diversity ON: greedy quality-diversity (next parent maximises `s + λ·behavioural distance` to those already chosen; distance = disagreement rate on the conformity pool — verifier-free, so usable in every arm). Diversity OFF: plain top-P by `s`. 3. **Breed** (sex ON): parents paired by descending behavioural distance (complementary mates); per pair, `n_cand` Dirichlet-weighted merges are screened on the arm's own selection signal (verifier val-pool for grounded arms; conformity for `no_grounding` — the verifier never enters that arm's loop) and the fittest offspring is kept (directed sex, E10). Sex OFF: children are redistilled copies of the selected parents. 4. **Reproduce** (the inheritance channel, identical in all arms): each child is a *fresh* LoRA trained from the base on (prompt → its source model's own answer) pairs over a fresh training pool — self-consumption made literal; knowledge survives only through the data channel. SFT stochasticity + fresh pools are the mutation operator. **Arms:** `full` (g>0, sex, diversity) · `no_grounding` (g=0) · `no_sex` · `no_diversity`. **Verifier truth is used for reporting in all arms** (test-pool fitness, consensus accuracy) but enters the *loop* only where g>0. ## Metrics (per arm × seed × generation) Per agent: test fitness (overall + per family + worst family), conformity, selection score, selected flag, parentage. Population: best/mean true fitness, behavioural diversity (mean pairwise disagreement), consensus accuracy, conformity−truth gap. E11's three panels re-drawn at LLM tier. ## Falsifiers (pre-registered) 1. `full` does not exceed `no_grounding` in final best true fitness → grounded evaluation adds nothing at LLM tier; the composed-society claim fails its LLM test. 2. Ablations do not fail distinctly (no conformity−truth gap in `no_grounding`; `no_sex` matches `full` on assembled capability; `no_diversity` matches on diversity trajectory) → the complementary-contributions claim does not transfer. 3. Honest alternative outcome: consensus may stay noisy rather than harden into a confident-wrong mean (specialists are wrong *differently* off-family) → `no_grounding` fails by drift, not by conformity; report the observed signature either way. ## Cost & schedule | Stage | Scale | Compute | Where | |---|---|---|---| | Smoke | N=4, G=2, tiny pools, `full` only | ~15 min | local 16 GB | | Pilot | N=6–8, G=8–10, `full` + `no_grounding`, 1 seed | ~4–10 GPU·h | local overnight / 1 CX3 job | | Campaign | 4 arms × 3 seeds, N=8–12, G=10–12 | ~50–150 L40S·h, ~1 day wall-clock as 12 parallel jobs | CX3 (`/imperial-hpc`); rent only if queue fails | | 7B confirm (optional) | headline contrast only (`full` vs `no_grounding`, 1–2 seeds) | ~100–300 GPU·h | CX3 / rented | ## Checklist - [x] Design + falsifiers (this document) - [x] `src/llm/society.py` (pure operators + GPU loop), `kind: llm_society` - [x] Pure-function tests (consensus, conformity, QD selection, pairing) — 155 green - [x] `configs/llm/society_smoke.yaml` → smoke green (exit 0; conformity 0.54→0.75, diversity 0.69→0.41 in one self-consumption generation — the signature is live) - [x] Pilot v1 (n_inherit=240, ep2, no elitism): **arms indistinguishable** — both drop ~0.1 in the first distillation generation and plateau at the same mutation–selection floor (best ≈0.40, mean ≈0.35, consensus 0.63→0.4 in both). Diagnosis: inheritance too lossy; the ~25%/generation distillation tax swamps the selection differential. v1 archived at `results/llm_society_pilot1`. - [x] Pilot v2 (n_inherit=600, ep3, founders 600×3, elitism=1 all arms, n_parents=3, n_conf=90): **the contrast is alive.** `full` climbs 0.583 → 0.617 best true fitness and holds (a bred offspring beat every founder at gen 3); `no_grounding` decays monotonically 0.583 → 0.533 despite identical elitism — conformist selection keeps the wrong elite. Self-consumption signature present: in `no_grounding` the population agrees with itself more (conformity 0.60 → 0.73) while the consensus gets less right (0.63 → 0.49); the confident-but-wrong gap reaches +0.24 vs +0.16 in `full`. One seed, N=6: direction only. `results/llm_society`. - [ ] GG gate: review pilot curves before the campaign - [ ] `hpc/llm_society.pbs` array job (arm × seed) → campaign - [ ] Figure `paper/manuscript/make_figs.py` panel(s); fold into main.md (replaces the "open" cell) - [ ] 7B headline confirm (optional, post-campaign decision) ## Open design decisions (defaults chosen; GG may override) - Families: the three *easy* variants at 0.5B (headroom exists; hard variants risk gen-0 inheritance data being mostly wrong → uninformative universal collapse). Revisit after pilot. - g for grounded arms: 0.5 (equal weight); sweepable later. - Population constant, non-overlapping generations, no elitism (faithful WF; champions must survive through inheritance, not by fiat). - Adapter disk hygiene: delete generation t−1 adapters once generation t is trained (keep gen 0 and final); ~35 MB × N × G × arms otherwise.