- paper/pnas -> paper/manuscript (venue-neutral)
- configs/layer1 -> configs/inheritance, src/knowledge -> src/inheritance
(imported as `inheritance`), make layer1 -> make inheritance; layer2 alias dropped
- inheritance and trained-network bundles named after the manuscript figure
they feed (fig2_grounding_sweep, figS3_rebaselining, ...), or descriptively
where they feed none; configs keep their `experiment:` value so parquet
hashes are unchanged, only output.dir moves
- figure scripts, SI figure sources, notebooks, REPRODUCING.md, README and the
SI Methods/tables updated; make clean no longer deletes tracked manifests;
reproduce.sh hashes the s{seed}/ layouts too
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
659 lines
65 KiB
Markdown
659 lines
65 KiB
Markdown
# Layer 1 Execution Plan — The Lamarckian Society (analytical core)
|
||
|
||
*Created 2026-07-04. Scope: blueprint §7 build-order steps 1–4 (scaffold → Layer 1 complete, laptop-reproducible). Layer 2 is explicitly out of scope here and gated on Layer 1's scientific-validation tests passing.*
|
||
|
||
Source of truth: `lamarckian-society-technical-blueprint-v1.md`. Where it is silent I record a decision below rather than improvise silently.
|
||
|
||
---
|
||
|
||
## Design decisions to pin before coding (blueprint leaves these open)
|
||
|
||
These are the only places the spec is genuinely underdetermined. Recommendation given for each; flagged for sign-off.
|
||
|
||
1. **Config framework — recommend thin pydantic + PyYAML, not Hydra.** Blueprint says "Hydra or a thin equivalent." Hydra is heavyweight and its global-state/`os.chdir` behaviour fights the "pass `rng` explicitly, results are a pure function of resolved config" requirement. A thin loader (pydantic models for validation + a small sweep-expander) is dependency-light, aligns with the stdlib-first preference, and makes the "write resolved config beside results" contract trivial. *Decision: thin pydantic loader.*
|
||
|
||
2. **Selection fitness `f_i` (the reality-anchored score).** Blueprint: "fitness is predictive accuracy against `p*`" but gives no formula for the discrete model. *Decision:* `f_i = p*_eff_i` (truth frequency = fitness; reality-anchored by construction). Post-selection distribution `p'_i ∝ p_i^(1−α) · f_i`, matching the blueprint's stated `w_i ∝ f_i·(p_i)^(−α)` with `α=0` recovering fitness-proportional greedy. Document as a modelling choice; `select-then-sample` vs `sample-then-select` is the documented robustness switch (§2.2).
|
||
|
||
3. **`tail_mask` definition.** Two knobs exist (`tail_frac`, `tail_threshold`). *Decision:* the metric-bearing tail set is the §2.3 definition `{i : p*_i < tail_threshold}`. `tail_frac` only drives the `twocomponent` construction (fraction of items placed in the low-mass component). Document that for `zipf`, `tail_frac` is unused.
|
||
|
||
4. **`TrueDist` type.** *Decision:* a frozen dataclass `TrueDist(p_star: np.ndarray, regions: np.ndarray, tail_mask: np.ndarray)` — immutable, so `p*` cannot be mutated in place (except the deliberate re-mint path, which produces a new object).
|
||
|
||
5. **Region partition.** *Decision:* contiguous blocks; require `K % R == 0` (assert with a clear error) to keep per-region math clean for v1.
|
||
|
||
6. **Re-mint semantics (E6).** *Decision:* re-mint replaces `p_star_eff` with a *new* `TrueDist` built from current `p_t` (fresh tail_mask recomputed on `p_t`), and the *original* `TrueDist` is retained **only** for the KL-to-original metric, never for grounding. This is the irreversibility the experiment tests.
|
||
|
||
---
|
||
|
||
## Phase 0 — Scaffold & environment
|
||
|
||
- [x] Install `uv` (`curl -LsSf https://astral.sh/uv/install.sh | sh`; lands in `~/.local/bin`, no sudo).
|
||
- [x] **Reproducibility path = `uv` venv from a committed, hash-pinned `uv.lock`** (GG decision, 2026-07-04; no Apptainer/Docker for Layer 1). `pyproject.toml` (Python ≥3.11; deps: numpy, scipy, pandas, pyarrow, matplotlib, pydantic, pyyaml; dev: pytest). Commit `uv.lock`. Create `.venv` via `uv sync`.
|
||
- [x] Repo layout per §5: `src/inheritance/`, `configs/inheritance/`, `figures/`, `results/` (gitignored), `tests/`, `paper/`. Add `src/lamarckian/` package root or make `knowledge` importable (decide package name — recommend `src/lamarckian/knowledge/...` with `src`-layout).
|
||
- [x] `.gitignore` (`.venv/`, `results/`, `__pycache__/`, `*.parquet` under results but keep hashes).
|
||
- [x] Seeding util `lamarckian/utils/seeding.py`: master seed → `np.random.SeedSequence(seed).spawn(n)` → per-replicate `np.random.default_rng(child)`. No global RNG anywhere.
|
||
- [x] Config loader `lamarckian/config.py`: pydantic schema mirroring the §2.7 YAML, a `load_config(path)`, a `expand_sweeps(cfg) -> list[ResolvedConfig]`, and `write_resolved(cfg, dir)`.
|
||
- [x] Manifest util: `write_manifest(dir, results_df)` recording lib versions, master seed, `git rev-parse HEAD`, content hash of `results.parquet`.
|
||
- [x] `Makefile` targets (`env`, `test`, `layer1`, `figures`, `clean`) + `pytest` skeleton. **Gate: `make test` green on a trivial test.**
|
||
- [x] Move `lamarckian-society-technical-blueprint-v1.md` → `paper/blueprint.md` per §5 (confirm with GG first — it's referenced by name elsewhere).
|
||
|
||
## Phase 1 — Core primitives + null model + VALIDATION GATE
|
||
|
||
Implement to the normative signatures in §2.7. Order chosen so each piece is unit-testable before the next depends on it.
|
||
|
||
- [x] `knowledge/truth.py::make_true_distribution` → `TrueDist`. Support `tail ∈ {zipf, twocomponent}`. Unit tests: normalisation, region block sizes, tail_mask matches threshold, determinism from seed.
|
||
- [x] `knowledge/metrics.py`: `forward_kl` (with `eps` floor, logged), `heterozygosity`, `tail_mass`, `support_size`, plus per-region variants. Unit tests on hand-computed small vectors.
|
||
- [x] `knowledge/step.py::generation_step` — **null path first** (single teacher, `m=0`, selection `none`): `c ~ Multinomial(n, p_t)`, `p_{t+1}=c/n`. Exactly neutral Wright–Fisher.
|
||
- [x] `knowledge/lineage.py::run_lineage(cfg, seed)` → tidy per-generation DataFrame (all §2.3 metrics, global + per-region).
|
||
- [x] `tests/test_scientific_validation.py` — **the spine:**
|
||
- **Pred. 1** heterozygosity decay: mean `H_t` over replicates matches `H₀(1−1/n)^t`. Prefer testing the full deterministic mean-recursion trajectory (subsumes the fixed point), within Monte-Carlo CI.
|
||
- **Pred. 2** fixation probability = initial frequency (long runs, statistical tolerance).
|
||
- [x] **HARD GATE: do not proceed until Pred. 1–2 pass.** If drift ≠ analytic decay, the harness is wrong — fix here.
|
||
|
||
## Phase 2 — Grounding + E1 + E2 (the headline)
|
||
|
||
- [x] `knowledge/step.py::structured_multinomial(m_vector, p_star, regions, policy, rng)` — per-region immigration draws from `p*` restricted+renormalised to each region; `uniform` spreads `m` evenly, `matched` concentrates on exercised regions. Returns length-K counts.
|
||
- [x] Extend `generation_step` with grounding (pooled draw, `g = m/(n+m)`).
|
||
- [x] **Pred. 3** validation — *exact* equilibrium `H_eq = H*·m(2n+m−1)/(n+2nm+m²)`: run to stationarity (burn-in + late-generation + replicate averaging), assert `<0.1%` rel. error vs closed form across an `m` grid. Also assert the `m→0` and `m→∞` limits.
|
||
- [x] **Pred. 4** validation — tail-persistence: item of freq `p*_i` maintained iff `m·p*_i ≳ 1`; verify the survival transition location statistically.
|
||
- [x] `knowledge/experiment.py::run_experiment(cfg)` — sweep grid × `n_replicates`; long-form results + CIs; write `results.parquet` + `resolved_config.yaml` + `manifest.json`.
|
||
- [x] **E1** config + run: `m=0`, single teacher, no selection. Expect `H` geometric decay, support→1, KL diverges, tail-first loss.
|
||
- [x] **E2** config + run: sweep `g`, single teacher, uniform grounding, no selection. Locate critical `g*` (transition in **tail mass / support**, since H is smooth in m — the sharp threshold is in discrete tail survival). Report `g*` with CI. **This is the load-bearing result.**
|
||
- [x] `figures/plot_collapse_null.py`, `plot_fig2_grounding_sweep.py` — read `results.parquet` only.
|
||
|
||
## Phase 3 — E3–E6
|
||
|
||
- [x] **E3 region-matched grounding.** Fixed total `m`; `uniform` vs `matched`; one designated inherited-but-unwatered region with a rare tail. Expect uniform lets that region's tail collapse; matched holds it. Per-region metrics essential. `plot_figS5_aimed_grounding.py`.
|
||
- [x] **§2.7.1 correlated-teacher construction** — `knowledge/teachers.py`:
|
||
- `make_retention_matrix(T, K_T, rho, q, rng)` — shared-switch exchangeable Bernoulli.
|
||
- `make_correlated_teachers(...)` — retention→distributions (head kept at `p*`; tail at `p*_i` if retained else `tail_floor`; renormalise). `region_specialisation` option.
|
||
- **Pred. 5** validation: `make_retention_matrix` reproduces marginal `q`, pairwise `ρ`, and union coverage `U(K_T,ρ,q)=T[ρq+(1−ρ)(1−(1−q)^K_T)]` to 3 decimals over a `(ρ,q)` grid.
|
||
- [x] **E4 multi-teacher decorrelation.** Sweep `K_T∈{1,2,3,5}`, `ρ∈[0,1]` at fixed `q`, matched budget (`n/K_T` each). Report **both** union `U` and post-distillation surviving coverage; show their gap shrinks as `g` rises. `plot_figS8_multiparent_union.py` (coverage surface over `(K_T,ρ)`).
|
||
- [x] **E5 QD vs greedy.** `apply_selection` (`none`/`greedy`/`qd`, pinned fitness form). Sweep novelty `α`. Expect greedy→fixation (`H→0`), qd holds `H` plateau + re-introduces tails. `plot_figS12_quality_diversity.py`.
|
||
- [x] **E6 re-mint gate.** Re-mint at high vs low `H`; track KL to *original* truth. Expect collapsed re-mint locks KL high forever; gated (high-H) does not. `plot_figS3_rebaselining.py`.
|
||
|
||
## Phase 4 — Reproducibility polish (Layer 1 slice)
|
||
|
||
- [x] `configs/inheritance/E1..E6.yaml` all committed with explicit params (no magic numbers in code).
|
||
- [x] `paper/figure_manifest.md` — the §6 claim→experiment→figure rows for Layer 1.
|
||
- [x] `make layer1` runs E1–E6; `make figures` regenerates all figures from committed parquet.
|
||
- [x] Full `test_correctness.py` (shapes, normalisation, determinism) + `test_scientific_validation.py` (Pred. 1–5) green in CI.
|
||
- [x] `reproduce.sh` (`uv sync` → `make test` → `make layer1` → `make figures` → write `REPRODUCED.md` diffing committed result hashes) + `README.md` reproduce section. **No container** — the committed `uv.lock` is the reproducibility source of truth (per GG, 2026-07-04); a Dockerfile may later wrap the same lockfile for Layer 2's GPU work.
|
||
|
||
---
|
||
|
||
## Definition of done (Layer 1)
|
||
|
||
Every Layer-1 row of blueprint §6 has a committed figure produced by `make figures` from committed results; all §2.4 analytic checks (Pred. 1–5) pass; `make layer1 && make figures` reproduces from a clean `.venv`. Then — and only then — Layer 2 may begin.
|
||
|
||
## Falsifier watch (report honestly if hit)
|
||
|
||
- E2 tail mass flat in `g`, or only stabilises as `g→1` → multigenerational thesis refuted.
|
||
- E3 uniform protects as well as matched → region-matching claim dies.
|
||
- E4 no surviving-coverage benefit at matched budget → recombination claim dies.
|
||
- E5 qd ≤ greedy stationary `H` → QD does no work.
|
||
- E6 collapsed lineage recovers original-truth tails after re-mint → irreversibility overstated.
|
||
|
||
## Review — progress log
|
||
|
||
**2026-07-04 — Phases 0 & 1 complete; hard gate PASSED.**
|
||
|
||
- Reorg: docs → `paper/` (`blueprint.md`, `the-lamarckian-society-v4.md`). src-layout under `src/inheritance/`.
|
||
- A pre-existing `tests/test_scientific_validation.py` (author-supplied, 22 KB) turned out to hard-specify the package contract — implemented *to it* rather than inventing interfaces. Key contracts it locked (now honoured): package imports as `knowledge.*`; `run_lineage(cfg_dict, seed)` returns a tidy per-gen frame with a `heterozygosity` column, rows 0..T; `p_0` initialises **uniform** (`H_0=1−1/K`); `metrics.heterozygosity` and `teachers.make_retention_matrix` match the reference to 1e-12 / closed form.
|
||
- Env: `uv` 0.11.26 installed; `pyproject.toml` + `uv.lock` committed; numpy 2.5, pandas 3.0, scipy 1.18, pydantic 2.13, pytest 9.1.
|
||
- Modules written: `metrics`, `seeding`, `config` (dataclasses + `from_dict`), `truth`, `teachers`, `step`, `lineage`. Config is dataclass-based (not pydantic) — the conformance test passes a raw dict; dataclasses validate cleanly and stay stdlib-simple. **Pydantic still a dep for the Phase-2 YAML/experiment layer.**
|
||
- **Results: `make test` green — 68 passed** (48 scientific-validation, 20 correctness). Conformance tests RAN (not skipped): Pred. 1 (neutral decay), Pred. 3 (exact `H_eq`), Pred. 5 (union coverage) all pass against the real package. **The Pred. 1–2 hard gate is passed**, and grounding already conforms to the exact equilibrium.
|
||
|
||
Design decisions #1 (dataclasses now / pydantic at YAML layer), #2 (fitness `f_i=p*_i`), #3 (threshold tail_mask), #6 (re-mint discard) all implemented as planned. Region design: each region an identical 1/R-mass block (symmetric; reduces to global Zipf at R=1, matching the reference).
|
||
|
||
**2026-07-04 — Phase 2 complete (E1 + E2).**
|
||
|
||
- `experiment.py`: sweep expansion (Cartesian grid; special-cases `g→m`), paired replicate seeds (shared across grid points), output contract (`results.parquet` + `resolved_config.yaml` + `manifest.json` with lib versions + git commit + sha256). CLI `python -m inheritance.experiment <cfg>`.
|
||
- **E1 (null collapse)** — reproduces tail-first collapse: H geometric decay matches `H₀(1−1/n)ᵗ` within CI; tail items die ~10× faster than head items; support 500→1; forward-KL diverges. Figure `results/collapse_null/E1.png`.
|
||
- **E2 (headline)** — `H_sim` tracks the *exact* `H_eq` closed form across the sweep; phase boundary at **`g* ≪ 1`**: g=0.005 (m=1 real sample vs n=200) → 68% of truth H; g=0.05 → 96%. g=0 slides to ~0.10 over 500 gens. Figure `results/fig2_grounding_sweep/E2.png`. **Headline result achieved.**
|
||
- Metric subtlety found & fixed: aggregate **`tail_mass` is a drift martingale** (mean-conserved), so it's a poor collapse indicator. Added `tail_support`/`head_support`/`tail_frac_alive`/`head_frac_alive`; E1 & E2 figures now use tail-*item* survival, which is honest and monotone.
|
||
- E2 extended 300→500 generations (GG-approved) so the g=0 arm visibly approaches 0 while g>0 arms sit on plateaus.
|
||
- Makefile `layer1`/`figures` wired to E1–E2. `make test` still green (68).
|
||
|
||
**2026-07-04 — Phase 3 complete (E3–E6) + E2 analysis add-ons.**
|
||
|
||
- **E3** region-matched grounding: added `grounding.exercised` knob + per-region `tailalive_region_r`. Target region tail survival 0.49 (matched) vs 0.07 (uniform). Note: per-region *H* is mass-confounded — used tail-item survival instead.
|
||
- **E4** multi-teacher recombination: bespoke `run_coverage` runner (`kind: coverage`). Union coverage matches `U(K_T,ρ,q)` exactly. **Key finding (GG-approved): mean-mixture distillation gives NO surviving benefit (conservation law — dilution cancels the union gain); max-merge (M2N2-style) does.** E4 reports both. In CLAUDE.md.
|
||
- **E5** QD vs greedy: greedy → H≈0.01 (fixation); qd holds H 0.48–0.88 rising with α. qd ≫ greedy.
|
||
- **E6** re-mint gate: added `arm` multi-override sweep type. Re-mint while collapsed → KL-to-original diverges (lock-in) + accelerates H collapse; diversity gate (H≥0.75) blocks it → bounded; healthy re-mint harmless.
|
||
- **E2 analysis add-ons** (companion work order `tasks/workorder-E2-analysis-addons.md`, verified): new `analysis.py` (`reduce_to_stationary`, `critical_grounding` bootstrap CI) — real E2 **g*=0.048, CI [0.047,0.050]**; `metrics.tail_band_metrics` + per-band lineage logging; `tests/test_analysis.py` reproduces the work order's verified numbers exactly. E2 figure rebuilt 2×2. **Deviation:** used truth-mass-weighted tail coverage instead of raw `tail_mass` (a drift martingale).
|
||
- All six figures regenerate via `make figures`; **71 tests green**.
|
||
|
||
---
|
||
|
||
# Layer 1.5 — Architecture-general neural existence proof (RNN/VAE/MLP + synthetic/MNIST)
|
||
|
||
*Created 2026-07-04. Plan: `~/.claude/plans/we-are-going-to-cheerful-fog.md`. Re-scopes Layer 2:
|
||
build a cheap, architecture-general neural collapse proof in real trained weights on a
|
||
fully-synthetic sandbox (exact known `p*`) before the LLM rung. Locked decisions: exact-oracle
|
||
categorical token sequences; Histogram+RNN+VAE+MLP; real MNIST as secondary confirmation; LLM +
|
||
C3 vertical claim deferred.*
|
||
|
||
## Progress log
|
||
|
||
**2026-07-04 — Stages A, B, plumbing complete.**
|
||
|
||
- **Env:** installed `uv` 0.11.26 (`~/.local/bin`); `/home` was 100% full — GG approved clearing
|
||
pip/yay/browser caches (~10 GB freed). Base venv synced; 71 Layer-1 tests green.
|
||
- **Stage A (scaffold, pure NumPy):** `src/neural/` — `config.py` (frozen dataclasses reusing
|
||
`inheritance.config` GroundingCfg/RemintCfg/MetricsCfg/_sub), `synthetic.py` (mode-truth via
|
||
`make_true_distribution`; lossless identity + stochastic style token grammar), `oracle.py`
|
||
(`ExactOracle` zero-error + `measure_distribution`), `models.py` (`GenerativeModel` protocol +
|
||
`HistogramModel` bridge), `evaluate.py` (reuses `inheritance.metrics`, Layer-1 row schema),
|
||
`generation_loop.py` (`run_generative_lineage`, reuses `allocate_m`/`structured_multinomial`).
|
||
15 correctness tests green.
|
||
- **Stage B — HARD GATE PASSED:** `tests/test_neural_validation.py` — histogram lineage reproduces
|
||
Pred. 1 (neutral decay, <3% rel err), Pred. 3 (exact `H_eq`, <5%), and tracks Layer-1
|
||
`run_lineage` directly (<3%). The neural plumbing reproduces the analytic core.
|
||
- **Plumbing:** `neural/experiment.py` (`run_and_save` dispatch on `kind`, reuses `_apply_param`
|
||
g→m, paired seeds); extended `inheritance.experiment.save_artifacts` (optional `extra_libs`,
|
||
`extra_manifest`, injectable `grid`; skips missing libs — backward compatible). `configs/neural/bridge.yaml`,
|
||
Makefile `neural`/`env-neural`/`layer2` targets, `.gitignore`. (Experiments are named
|
||
descriptively — `bridge`, `collapse`, `grounding`, `architectures` — not by code.)
|
||
- **`bridge` result (17s):** neural **g\* = 0.0474, CI [0.045, 0.052]** — reproduces Layer-1 E2's
|
||
g\*=0.048 essentially exactly (g=0.005→67% of H*, g=0.05→96%). **89 tests green.**
|
||
|
||
**2026-07-04 — Stage C: torch models + collapse/grounding/architectures.**
|
||
|
||
- **Env:** torch **2.12.1+cu130** (default PyPI wheel ships CUDA 13, matches RTX A4000 driver;
|
||
no custom index needed, cp314 wheels exist). `--extra neural` = torch only; `--extra mnist` =
|
||
torchvision (later). `UV_CACHE_DIR=/tmp` during install (RAM-backed) to spare `/home`.
|
||
- **Models:** `torch_models.py` (RNNGenerator, autoregressive GRU), `torch_mlp.py` (autoregressive
|
||
MLP, causal-masked), `torch_vae.py` (sequence VAE), `train.py` (determinism flags + device/seed
|
||
helpers derived from the SeedSequence stream). `tests/test_neural_torch.py` (torch-gated): gen-0
|
||
fidelity (rnn+mlp) + dry-collapse/grounded-holds. **92 tests green.**
|
||
- **Validated regime:** K=256, n=200, zipf_s=1.3, RNN hidden=128/epochs=25. RNN gen-0 fidelity
|
||
KL(p*‖p̂)=0.008, 64/64 (or 256/256) modes recovered. MLP fidelity KL=0.011. **VAE does NOT clear
|
||
the gen-0 gate** on the Zipf-codeword task (KL≈0.8; prior-hole mismatch — sampling z~N(0,I) misses
|
||
the aggregate posterior) → excluded from `architectures` to avoid confounding collapse with underfitting.
|
||
- **`collapse` (in weights):** dry RNN lineage collapses — forward-KL rises to ~2.2 vs grounded
|
||
~1.4; grounding lifts tail survival (tailalive 0.31 dry → 0.50 at g=0.02). Sign confirmed.
|
||
- **`grounding` (neural phase boundary):** stationary H hovers 80–91% of H* and is **noisy / non-monotonic**
|
||
at 5 reps — no crisp g*. **KEY FINDING:** the neural models' smoothing inductive bias *partially
|
||
resists* H-collapse (dry H stays ~83% of H*), so **forward-KL and tail survival are the sharp
|
||
neural collapse metrics, not H** (mirrors Layer-1's "H is smooth; the threshold lives in tail
|
||
survival"). `grounding` needs (a) forward-KL as the phase metric, (b) more reps (≥10), and/or (c) a
|
||
stronger-collapse regime for a clean neural g*.
|
||
- **`architectures` (architecture-generality) — clean result:** collapse + grounding-rescue appear in ALL three
|
||
model classes (dry→grounded forward-KL: histogram 6.2→4.6, MLP 4.8→1.3, RNN 3.8→1.1; tailalive
|
||
RNN 0.41→0.64, MLP 0.07→0.20). The WF operator is architecture-general. Bonus: neural smoothing
|
||
lets RNN/MLP retain *more* tail than the exact histogram under grounding (they generalise to
|
||
unseen codewords) — an inductive-bias finding worth the write-up.
|
||
|
||
**2026-07-04 — `recombination` (load-bearing E4 replication).**
|
||
|
||
- `recombine.py` mirrors `run_coverage` but trains K_T specialist RNNs on assignments from the
|
||
exact shared-switch retention construction (K_T/rho/q clean; union matches the closed form), then
|
||
recombines the *measured* teacher distributions two ways: `mean` (naive pooling) vs `max`
|
||
(oracle-guided union / M2N2-style), each followed by size-n resampling. The neural merge is the
|
||
per-mode max over teacher distributions (oracle-guided), NOT weight-averaging of RNNs.
|
||
- **Result (8 reps):** at rho=0, **union rises 0.49→0.96** (supply matches closed form); analytic
|
||
**surviving_max rises 0.043→0.087 while surviving_mean stays flat ~0.045** — the conservation law
|
||
(averaging cancels the union gain; max-merge realises it). At rho=1 (identical teachers) union AND
|
||
max are flat — more identical teachers buy nothing. **The "merge, don't average" lesson holds in
|
||
the neural setting.** Trained-weight columns show the same signs but noisier: neural smoothing
|
||
inflates baseline survival and the deep tail barely clears n=200 resampling (compresses magnitude)
|
||
— the expected inductive-bias caveat. torch-gated test added. **93 tests green.**
|
||
|
||
**2026-07-05 — `grounding` refinement + figure (honest reframing).**
|
||
|
||
- Re-ran at **18 reps** (n_eval 15000, 30 gens, g grid refined to 9 points). Falsifier pinned in
|
||
the config *before* running.
|
||
- **forward-KL is the operative neural collapse metric — NOT H or tail-survival.** The RNN's
|
||
smoothing keeps spurious tail modes alive, so `tail_truth_mass_alive` is flat/**non-monotone**
|
||
in g (dry 0.54 > most grounded) and H stays 0.77–0.85 of H\*. Stationary **forward-KL** falls
|
||
monotonically (dry 2.08 → g=0.2: 0.75), significant at g≥0.05 (paired t→3.3; 89% of lineages
|
||
improve at g=0.2). Refines the earlier "forward-KL AND tail survival" note.
|
||
- **The sharp `g*≪1` is an exact-operator feature, softened by neural smoothing.** Median-recovery
|
||
grounding (half the KL gap closed) g≈**0.04** (bootstrap CI [0.004, 0.116]) — echoes Layer-1's
|
||
0.048 — but full (95%) recovery needs g≈0.19. Quantitative `g*≪1` is carried by the histogram
|
||
**bridge** (0.047); the RNN confirms the SIGN and softens the sharpness (blueprint §3.5 met).
|
||
- **Honest note:** the pre-registered 95%-of-H\*/tail-survival falsifier is *not* met — because
|
||
those are the wrong metrics for a smoothing model, not because grounding fails. Reported as such.
|
||
- **Robustness fix:** a fully-degenerate RNN can emit only invalid codewords → `measure_distribution`
|
||
now returns a terminal-collapse sentinel (fixation on the dominant mode) instead of crashing a long
|
||
sweep. Edge-case test added. `figures/plot_figS6_grounding_rnn.py` written (4-panel, states its own verdict),
|
||
wired into `make figures` (glob all `plot_*.py` except `plot_E[1-6]`).
|
||
|
||
**2026-07-05 — neural figures (all five).**
|
||
|
||
- `figures/plot_{bridge,collapse,grounding,architectures,recombination}.py`, each a pure function
|
||
of its committed bundle (reuse `figures/_figlib.py`), wired into `make figures` (glob all
|
||
`plot_*.py` except `plot_E[1-6]` / `_*`). **bridge**: neural histogram runner sits exactly on the
|
||
exact `H_eq` curve, g*=0.047 (HARD-GATE visual). **collapse**: dry GRU forward-KL climbs, grounded
|
||
held; H barely moves. **architectures**: grouped bars — forward-KL falls / tail survival rises with
|
||
grounding across histogram/GRU/MLP. **recombination**: union matches closed form; max-merge rises
|
||
while mean-distill stays flat (analytic + trained + rho=1 control). **grounding**: the reframed
|
||
4-panel (forward-KL phase boundary, recovery, metric-choice).
|
||
|
||
**2026-07-05 — real-MNIST external-validity tier (`mnist_collapse`).**
|
||
|
||
- New image tier plugged into the existing contract (metrics/grounding/output are data-agnostic and
|
||
reused verbatim): `mnist_data.py` (load, per-class thickness bins, mode = class×thickness bijection,
|
||
`MnistSampler`), `mnist_oracle.py` (frozen CNN + deterministic thickness = `ClassifierOracle`,
|
||
confusion matrix), `mnist_vae.py` (`ConvVAEGenerator`), `mnist_loop.py` (`run_mnist_lineage`), plus
|
||
`kind=mnist_lineage` dispatch in `experiment.py`, `configs/neural/fig2_mnist_collapse.yaml`,
|
||
`figures/plot_fig2_mnist_collapse.py`, `figures/mnist_montage.py` (eyeball diagnostic), `MnistCfg`/`OracleCfg`.
|
||
- **Gates:** CNN mode accuracy **98.5%** (30×30 confusion matrix in the manifest = noise floor);
|
||
VAE gen-0 recovers full 30/30 support (over-smooths freq, KL≈0.5, no prior hole).
|
||
- **Result (4 reps):** dry (g=0) VAE **collapses to a single mode** (KL 0.5→18, support 30→1, tail
|
||
1.0→0.06, H→0); **g=0.1 holds all 30 modes** (KL≈0.6, full tail, H≈0.9). Collapse + grounding-rescue
|
||
confirmed on real images. VAE needs ~10% grounding vs synthetic ~5% (cf. the `grounding` finding).
|
||
**99 tests green** (+5 torchvision-gated). `make mnist` / `make env-mnist` (kept out of `make neural`).
|
||
|
||
**2026-07-05 — learning kernel (Layer-1 extension) + Riis positioning.**
|
||
|
||
- Prompted by revisiting Layer 1 vs 1.5 and the Riis 2026 paper (arXiv:2604.08554). Added
|
||
`knowledge/kernel.py` (`LearningKernelCfg`: reset `u` = smoothing, temperature `τ` = sharpening,
|
||
floor `ε`), wired into `step.generation_step` (post-refit) / `StepCtx` / `DynamicsCfg` — **identity
|
||
by default, so the 68 Layer-1 scientific-validation + correctness tests are unchanged.**
|
||
- **Result:** neutral Wright–Fisher fails BOTH neural models, oppositely. VAE regime (n=6000,K=30):
|
||
neutral is inert, sharpening `τ=0.8` reproduces the collapse-to-one-mode. RNN regime (n=200,K=256):
|
||
neutral → H=0, mutation `u=0.006` reproduces the H-floor (~0.68). Uniform-mutation overshoots the
|
||
RNN's KL → its prior is truth-like, not uniform (honest caveat, future refinement).
|
||
- `configs/inheritance/kernel_{sharpen,smooth}.yaml`, `figures/plot_kernel.py` (overlays analytic arms vs
|
||
the committed neural endpoints), READMEs, `tests/test_kernel.py` (+6). Wired into `make layer1`.
|
||
- **Strategic (see CLAUDE.md finding):** concede "collapse=drift" to Riis (prior art; cite); his
|
||
mixed environment retains OLD SYNTHETIC (no real-data injection) → pessimistic, no g* that prevents
|
||
collapse. Our defensible novelty: recombination "merge-don't-average" (flagship), the learning-kernel
|
||
axis (he flags as future work), grounding-as-immigration, real-weights+MNIST breadth, and the
|
||
Lamarckian society + vertical/cumulative C3 claim (not yet run). Reposition: from "collapse is drift"
|
||
to a population-genetic CONTROL THEORY for sustaining open-ended knowledge.
|
||
|
||
**2026-07-05 — multi-locus society frame (E7/E8): raised the ceiling to enter the society.**
|
||
|
||
- Prompted by "enter the society with a robust theoretical frame." The single-locus fixed-`p*` model
|
||
can't express "exceeding" a ceiling. Generalized knowledge to a distribution over **genotypes**
|
||
(`knowledge/genotype.py`: `L` biallelic loci, `K=2^L`, additive fitness, recombination = product of
|
||
per-locus marginals). Reuses all K-mode machinery + `make_retention_matrix` (locus mastery).
|
||
- **E8 (star, `kind: society`, `knowledge/society.py`) — the vertical claim:** decorrelated parents
|
||
recombined; **sexual merge reaches the optimum (12/12, a genotype no parent had)** as parent count
|
||
grows / `ρ→0`, while best-parent (~8.7) and mean-mixture soup (~11.6) plateau. `configs/inheritance/figS9_specialist_superparent.yaml`,
|
||
`plot_figS9_specialist_superparent.py`, README. The Fisher–Muller effect for AI.
|
||
- **E7 (`kind: genotype_lineage`, `knowledge/genotype_lineage.py`) — advantage of sex:** sexual lineage
|
||
adapts faster than asexual (LD→0 vs LD spike). Honest: a speed advantage, not a permanent ratchet gap.
|
||
- **Metaphor shift (GG):** sexual reproduction with **unbounded parents**, not teacher→pupil (which caps
|
||
at the ceiling). Collapse = asexual degradation; cure = sex, no parent limit. Unifies E4+E6 under
|
||
evolution-of-sex theory; beyond Riis's single-locus n-grams. `tests/test_genotype.py` (+7).
|
||
Experiment dispatch (`kind` in {genotype_lineage, society}) + `make layer1` wired.
|
||
|
||
**2026-07-05 — sexual-transmission model made rigorous (E9/E10): landscape robustness + directed sex.**
|
||
|
||
- GG excited by the sexual metaphor; wanted it robust before the full society. Added NK landscape
|
||
(`genotype.nk_fitness`), finite n-parent `crossover`, `hill_climb` (parents = local optima).
|
||
- **E9 (`recomb_landscape`) — "why sex?":** on rugged/epistatic landscapes, blind recombination →
|
||
**outbreeding depression** (offspring below parents, worse with ruggedness + recombination rate);
|
||
the optimal recombination rate shrinks with ruggedness. Design rule: merge freely when
|
||
complementary, sparingly + selectively when entangled.
|
||
- **E10 (`directed_sex`) — AI beats biology:** random ("biological") sex craters with ruggedness
|
||
(0.66→0.51); **directed sex** (choose mates + select offspring + unbounded parents, iterated) tracks/
|
||
exceeds the best parent at every ruggedness. The distinctly-AI superpower, no biological analog.
|
||
- Complete picture: dramatic super-parent offspring when complementary (E8); outbreeding-depression
|
||
risk when entangled (E9); directed sex resolves it (E10). `configs/inheritance/{E9,E10}.yaml`,
|
||
`plot_{E9,E10}.py`, READMEs, +5 tests (117 green).
|
||
|
||
**2026-07-05 — the dynamic Lamarckian society (E11): the vertical claim / C3 realized.**
|
||
|
||
- `knowledge/dynamic_society.py`: finite population of N agents (genotypes) on a rugged NK landscape
|
||
(reality); composes grounding + directed sex + quality-diversity selection + mutation. Grounding
|
||
made load-bearing via consensus-conformity (self-consumption): selection on
|
||
`g·true_fitness + (1-g)·conformity` (GG decision). `kind: dynamic_society` dispatch.
|
||
- **4-arm ablation (12 reps), each breaks distinctly (global_opt≈0.79):** full 0.78 (climbs to optimum,
|
||
diversity maintained longest); no_sex 0.77; no_diversity/greedy 0.74; **no_grounding 0.48
|
||
(self-consumption collapse to unfit consensus).** Only the full society climbs. Integrates E1-E6 +
|
||
kernel + E7-E10 into one system: needs ALL of grounding + directed sex + diversity.
|
||
- `configs/inheritance/fig4_society_ablation.yaml`, `plot_fig4_society_ablation.py`, README, `tests/test_dynamic_society.py` (+5, 122 green).
|
||
Closes C3 analytically; the LLM rung remains the eventual empirical instantiation.
|
||
|
||
## Remaining (all optional / next)
|
||
- [ ] **NK/epistasis landscape** (sign epistasis can make recombination harmful — the honest limit of
|
||
"sex always helps"); **multi-allelic loci**. Deepens the frame.
|
||
- [ ] **Learning-kernel refinement:** truth-like smoothing prior (`prior="truth"`) + measurement floor
|
||
for a quantitative RNN match; **multi-locus / linkage** modes (class×style) as the rigorous home for
|
||
recombination. Both enrich predictive power and separate us further from Riis's single-locus n-grams.
|
||
- [ ] **The Lamarckian society experiments** (multi-agent grounding + decorrelated specialists +
|
||
recombination + QD-selection + re-mint) and the **vertical/cumulative C3 claim** — the highest-ceiling,
|
||
wholly-novel frame; not yet entered.
|
||
- [ ] **`region_matched`** grounding (R>1), **`remint`** re-mint gate (optional).
|
||
- [ ] **VAE fidelity:** fix the prior-hole mismatch (KL-annealing / free-bits / larger latent) so it
|
||
clears the gen-0 gate, then add to `architectures`. Or document as a known limitation.
|
||
- [ ] Real-MNIST secondary tier (`ClassifierOracle` + confusion matrix; `--extra mnist`).
|
||
- [ ] `figures/plot_<name>.py` (reuse `figures/_figlib.py`); wire into `make figures`.
|
||
|
||
## Discovered during work
|
||
|
||
- **E2 grounding policy vs. the analytic H_eq:** Pred. 3's closed form is derived for *plain* immigration `Multinomial(m, p*)`. Implemented as `policy="proportional"`, and every policy reduces to it at `R=1`. E2 should therefore run at `R=1` (or `proportional`) so the phase-boundary sweep tracks the exact `H_eq`; region structure is E3's concern. Decide E2's `init` (uniform vs truth) when building its config.
|
||
- `init: {uniform|truth}` added to `TruthCfg` (uniform default, mandated by the decay conformance test). E1/E2 may want `truth` start for a clean "tail collapses from the truth" story — revisit in Phase 2.
|
||
|
||
## Potential agents
|
||
|
||
*(none proposed yet)*
|
||
|
||
**2026-07-05 — LLM prototype (`llm_merge`): first real-LLM step, honest/partial.**
|
||
|
||
- `src/llm/` package: tasks+exact-match verifier, batched eval, LoRA specialise (manual SFT), peft
|
||
weight-merge (soup/ties), runner (`kind: llm_merge`). Base Qwen2.5-0.5B-Instruct on one 16GB GPU.
|
||
- **Result (seed 1):** merges are the ONLY models competent across all 3 disjoint families
|
||
(worst-family ~0.25 vs <0.16 for any single specialist) — the Fisher-Muller signature, robust.
|
||
Overall-exceeds is marginal (soup 0.64 vs best spec 0.63; ties below), and averaging dilutes peaks
|
||
(lists 0.43->0.26 = "merge don't average" in real weights). Pipeline works end-to-end; strict
|
||
overall-exceeds needs scale (bigger base/more families/seeds/dilution-resistant merge) = HPC step.
|
||
- Python 3.14 + transformers 5.13 OK; note transformers-5.x apply_chat_template returns a dict.
|
||
`make env-llm`/`make llm`; `figures/plot_llm_merge.py`, README, `tests/test_llm.py` (+3, 125 green).
|
||
|
||
**2026-07-05 — LLM merge 7B firm-up on Imperial CX3 (`llm_merge_hpc`): marginal sign → decisive.** ✅
|
||
- Ran on one L40S (46 GB) via `/imperial-hpc` runbook; 8 min walltime; Qwen2.5-7B-Instruct, 200 tests/family.
|
||
- **Both merges 0.87 overall > best specialist 0.77** (decisive +10 pts) and beat every specialist on
|
||
every family; worst-family 0.62 vs ≤0.57. Both 0.5B caveats resolved: overall-exceeds is now clean,
|
||
and dilution VANISHES (merge 0.62 > lists-spec 0.57 on lists) — dilution was a small-model artefact.
|
||
- `results/llm_merge_hpc/` (README legend, data-driven figure title). Next refinement: module-level
|
||
union-preserving recombination (MoE-expert/adapter-union = real-weight E8 max-merge), not delta-avg.
|
||
|
||
**2026-07-05 — MoE-expert / union recombination (`llm_moe`): E8's `max` vs `mean` in real weights.**
|
||
- `src/llm/moe.py`: router (oracle + training-free nearest-centroid over base embeddings) + MoE
|
||
generate + router-free per-module `max_merge`. `kind: llm_moe` reuses the cached specialists.
|
||
- **0.5B result:** routing beats fusion decisively — overall 0.74/worst 0.43 vs soup 0.64/0.26, no
|
||
dilution (recovers each specialist's own-family peak). Riders: learned router trivially perfect
|
||
(1.00, lexically-separable families) and static `max_merge` a poor union (0.46, not input-adaptive).
|
||
- `configs/llm/{moe,moe_hpc}.yaml`, `figures/plot_llm_moe.py`, README, +2 tests (127 green),
|
||
`hpc/llm_moe.pbs`.
|
||
- **7B firm-up (`llm_moe_hpc`, CX3 L40S, 9 min): the ordering FLIPS.** At 7B fusion wins —
|
||
soup 0.87 > routing 0.84 > max_merge 0.78 (0.5B had routing 0.74 > soup 0.64). Routing is capped at
|
||
the best parent per family; fusion *composes beyond* it at a capable base (soup lists 0.62 >
|
||
spec 0.57). So "merge, don't average" is a **weak-base law**, not universal — union wins under
|
||
dilution (0.5B), fusion wins under composition (7B). Refines E8. Next: fusion + offspring-selection
|
||
(directed sex). `results/llm_moe_hpc/` README + regime-aware figure.
|
||
|
||
**2026-07-05 — Directed sex (`llm_directed`): E10 in weights = breed offspring + select on verifier.**
|
||
- `src/llm/directed.py`: sample a population of Dirichlet-weighted merges, score on a held-out VAL
|
||
split, keep the best-overall + best-worst-family, report on a fresh TEST split. `kind: llm_directed`.
|
||
- **0.5B:** directed selection beats the single a-priori soup on the bred objective — directed_overall
|
||
0.69 > soup 0.64; directed_balanced worst-family 0.37 > soup 0.26. Riders: single-objective selection
|
||
trades off the other axis (overall-breed tanks lists 0.17); a global blend still trails per-input
|
||
routing (0.74). **7B (CX3 L40S, 9 min):** directed ≈ soup (0.868 ≈ 0.873) — soup already composes to
|
||
ceiling on near-saturated families (strings/arith 1.00), no fitter offspring to breed.
|
||
- **Through-line:** recombination refinements pay off ∝ how suboptimal the default soup is — big at
|
||
0.5B, nil at 7B. Honest limit: 7B families near-saturated; a harder benchmark is the fair next test.
|
||
- `configs/llm/{directed,directed_hpc}.yaml`, `figures/plot_llm_directed.py`, READMEs, `hpc/llm_directed.pbs`,
|
||
Makefile `llm` target, +3 tests (130 green). `results/llm_directed{,_hpc}/`.
|
||
|
||
**2026-07-05 — HARD benchmark: the 7B "fusion wins / no headroom" nulls were SATURATION artefacts.** ✅
|
||
- Easy families saturated 7B (strings/arith 1.00), confounding the moe/directed 7B nulls. Built a hard
|
||
task variant (`hard: true`: multi-step lists, Caesar ciphers, multi-step/larger arith) threaded through
|
||
make_tasks/train_specialist/runners; hard specialists cache as `spec_*_hard`. Ran both at 7B on CX3
|
||
(one L40S, 24 min, unsaturated).
|
||
- **Both nulls flip back to the 0.5B ordering:** union/routing 0.500 > fusion 0.40 (soup dilutes the
|
||
strings specialist 0.665→0.300, below even the best single specialist 0.425); directed selection 0.492
|
||
> soup 0.392 (+10 pts). **The operative variable is HEADROOM, not base-size** — "merge, don't average"
|
||
and "directed sex" hold whenever there's room to lose to dilution (weak base OR hard tasks); fusion only
|
||
wins where easy tasks let a strong base compose to ceiling. Vindicates E8 max>mean at 7B.
|
||
- `configs/llm/{moe_hard,moe_hard_hpc,directed_hard_hpc}.yaml`, `hpc/llm_hard.pbs`, hard READMEs+figures,
|
||
+1 test (131 green). `results/llm_{moe,directed}_hard_hpc/`.
|
||
|
||
**2026-08-11 — PNAS submission campaign opened.** GG approved PNAS as target after the post-hold
|
||
re-assessment (fresh lit scan: speciation/sex-framing/mating-systems/headroom all still unclaimed;
|
||
new concessions First-Extinction 2509.20101 + qt-trait 2407.17493 + verifier-injection 2510.16657;
|
||
E13 exposed to richer-symmetry objection 2606.23607). Full plan: `tasks/workorder-pnas-submission.md`.
|
||
Phases: (1) E13 hardening (scale-aware alignment + emergent-divergence condition), (2) arXiv preprint,
|
||
(3) llm_speciation + multi-seed LLM arc, (4) PNAS-format manuscript (5 consolidated figures, dual
|
||
audience), (5) submission mechanics (Zenodo DOI, cover letter, editor/reviewer suggestions).
|
||
|
||
**2026-09-07 — `llm_society` opened: the composed society at LLM scale (C3), pre-submission.** GG
|
||
decision: a weeks-scale experiment closing the paper's largest stated gap must be in the submission
|
||
("any reviewer would ask to see it"); rent compute if CX3 queues fail. Full design + falsifiers +
|
||
schedule: `tasks/workorder-llm-society.md`. E11 re-instantiated in LoRA agents: grounding knob in
|
||
the evaluation channel (g·verifier + (1−g)·conformity), self-consumption inheritance (children
|
||
distilled from their source's own answers), directed sex (complementary pairing + Dirichlet
|
||
offspring screened on the arm's own signal), QD selection. `src/llm/society.py` (+4 pure tests,
|
||
155 green), `kind: llm_society`, configs `society_smoke.yaml` / `society.yaml`. Stages: smoke
|
||
(local, ~15 min) → pilot full vs no_grounding (GG gate) → 4-arm × 3-seed CX3 campaign → figure +
|
||
manuscript fold-in.
|
||
|
||
**2026-09-07 — v1 `llm_society` campaign landed (4 seeds) and is NEGATIVE; v2 pre-registered.**
|
||
Best-agent overall at gen 9, 3-seed means: `no_sex` 0.558 ≥ `no_diversity` 0.539 ≥ `full` 0.506 ≫
|
||
`no_grounding` 0.436 (worst arm in every seed from gen 2). Conformity−truth gap does not separate the
|
||
arms. Read through the framework the null was structurally guaranteed (near-clone founders over 3
|
||
families; 2³ competence states; linear blending at 0.5B = the dilution regime; parents truncated
|
||
before breeding, unlike E11's survival-over-pool; `n_test`=40 → SE 0.079) — details and fixes in
|
||
`tasks/prereg-llm-society-v2.md` §1, lesson in `tasks/lessons.md`. Nothing enters the manuscript;
|
||
Fig. 1A's "stated gap" stands. **v2** (`kind: llm_society_v2`): L=12 families / one founder each,
|
||
confidence-routed union inheritance, pooled survival, checkpoint+resume, 240 test items, g=0.85,
|
||
G=12; six numerical hypotheses H1–H6; calibration gates C1–C5 must pass before submission (GG
|
||
reviews). GG decisions: 0.5B; `sex_linear` dropped (H2 deferred); no family vetoes.
|
||
- [x] families / operators / v2 loop / calibration runner / configs / PBS array / figure script / 9 tests (164 green)
|
||
- [x] smoke (4 arms, figure + stats script) → calibration A pass 1 (6/17 in band) → pass 2 (9 in band; C1b needed the gate re-derived 0.35→0.41 from the grid) → GG chose L=9
|
||
- [x] calibration B: C2 FAILED as pre-registered (retention ≤0.81 at k≤150; interference, not the observation floor) → C2b: k=300 + confidence gate τ=0.5 gives mean retention 0.87 (PASS); C3 operator half passes (union holds both families, linear loses one), retention half re-run gated; C5 passes (consensus 0.31)
|
||
- [x] campaign configs set: L=9, k_inherit=300, conf_gate=0.5, epochs=3, n_test 27/family, g=0.85, G=12; PBS 16 elements × 8 h
|
||
- [x] gated cross (C3): two-skill child plateaus at ~0.85×/0.8× of parents at any budget (3 vs 6 epochs; r64 hurts); tight gate τ=0.85 gives the 6-epoch retention at 3 epochs
|
||
- [x] **GG go/no-go (21:30): NO-GO at 0.5B** — the vertical claim needs 5–6 co-resident skills the r=16 adapter cannot hold; today = a measured transmission ceiling (SI material). 7B plan drafted: prereg §13
|
||
- [ ] GG: 7B scope (headline ~80 / H3+H4 ~120 / full ~260 L40S-h), Phase-0 task design go, SI text timing
|
||
- [ ] SI: the 0.5B calibration ceiling as the reason the tier was not run (prereg §11 row 4) — three limits, numbers from results/llm_society_v2_calib_*
|
||
- [ ] REPRODUCING.md: rows for the v1 campaign (4 seeds), v2 smoke, and the 9 calibration bundles
|
||
- [ ] stage code on CX3, `qsub hpc/llm_society_v2.pbs` (16 elements); local hedge = `society_v2_s1.yaml`
|
||
- [ ] `figures/stats_llm_society.py` (per-seed paired contrasts H1/H3/H4, AUC for H5, supplied-vs-retained for H6)
|
||
- [ ] fold the outcome per prereg §11
|
||
|
||
**2026-09-08 — v3 `llm_compose` run (3 seeds): H1 PASS, H2–H5 null; design could not show the claim.**
|
||
Composition at gen 0 is real and replicated (surplus +0.087/+0.033/+0.093; union-exceedance ~0.12; also
|
||
on MATH-500). Decay hypotheses uninterpretable: the lineages barely drifted (q_math 0.54 → 0.50–0.57)
|
||
and, more fundamentally, a fixed skill set has its ceiling at gen 0 — GG: "are models learning NEW
|
||
skills at EACH generation? … that was not the problem being addressed." v3 was Weismannian (fresh LoRA
|
||
each generation) and retention-only. Two of my errors: C3 unchecked (code specialist 0.075 on MBPP →
|
||
q_code noise), and three premature reads of a single-seed trajectory.
|
||
**v4 `llm_curriculum` — continual learning in a population** (`prereg-llm-society-v4.md`): Lamarckian
|
||
channel (`continue_lora_training`), Latin-square curriculum (complementarity 1.0 → 0.0 by construction,
|
||
H6 predicts the *shape*), arms isolated/society/society_dry/seed_bank (GG's ancestor-merge idea —
|
||
temporal vs spatial complementarity, direction uncommitted), single-shot SoTA baselines at matched budget
|
||
as the falsifier. GG decisions: 3×9×9, replay fixed-total, baselines get the same directed selection.
|
||
- [x] G0 PASS (accumulation 0.74 → 0.95); G2 FAIL ×2 (v2 families don't interfere; one pair at +0.65);
|
||
G3 negative (merging costs −0.01…−0.13 with nothing to repair); base = 0.094 → one family lifts all to 0.417
|
||
- [x] **v5 curriculum**: 11 real-dataset families, 5 answer shapes, per-family verifiers, disjoint splits
|
||
(`curriculum_data.py`, +6 tests, 56 green); selection rule fixed in prereg §8a
|
||
- [ ] stage A calibration running (`curriculum-v5-calib`): base + 11 specialists × 11 families
|
||
- [ ] stage B: zero-replay forgetting probe on the survivors (mean drop ≥ 0.15, not single-family)
|
||
- [ ] GG go/no-go → seed 1 local + CX3 array (seeds 2–3), then baselines
|
||
|
||
**2026-09-08 — v5 curriculum campaign done (3 seeds); all hypotheses fail; the useful finding is a
|
||
split in the theory.** Real-dataset curriculum (6 families, 5 answer formats, per-family verifiers)
|
||
replaced the procedural set. Results: not-merging wins (0.80), merging-with-own-ancestor middling
|
||
(0.66), merging-with-a-peer collapses (0.27), single-shot merging unstable (0.125–0.764). Cause:
|
||
two families are answer-format destroyers that propagate through merges and compound because
|
||
offspring continue the lineage. Scope limits: merging was obligate (no veto) and there is NO
|
||
selection between lineages — a gene-flow experiment, not a selection one.
|
||
**Key new measurement:** same-skill adapters (seed/data draw only) are near-orthogonal in weight
|
||
space (cos +0.006), disagree on 24% of prompts, and **merging them beats the best parent by +0.087,
|
||
exactly at the either-right ceiling**. So decorrelation-in-what-you-know is harmful while
|
||
decorrelation-in-how-you-encode-it is beneficial — the framework's single rho conflates them.
|
||
- [ ] v6: three arms (no-merge · complementary · parallel) + veto + population selection (~30 GPU-h)
|
||
- [ ] decide how the E9-risk result and the two-variations split enter the manuscript (beside Fig. 5A)
|
||
|
||
**2026-09-08 (later) — four mechanism probes; two of my explanations retracted.** (1) Same-skill
|
||
adapters: 85% of a LoRA's change is run-specific noise; merging two beats the better parent by +0.087,
|
||
at the either-right ceiling. (2) Denoising before crossing adds +0.025 on both skills at once
|
||
(inbred-lines signature). (3) A single merge of clean adapters is PROTECTIVE (0.825 vs 0.550 best
|
||
parent) — retracts "destructive skill propagates through merges". (4) Five chained convex merges lose
|
||
nothing, while signal-preserving additive weights collapse (1.02 vs 0.52 retention) — retracts
|
||
"geometric signal dilution"; the real constraint is bounding drift from the base. (5) Scaling probe
|
||
(GG's control): base 0.000, and the adapter works down to 1/8 then dies — 1/16 = 0.450, 1/32 = 0.000.
|
||
So the chain's apparent retention was ANSWER FORMAT supplied by the dominant partner, not the skill.
|
||
Consistent with Fig. 3C-D: functional conflict predicts merge damage, weight geometry does not.
|
||
- [ ] test the remaining candidate: continued training ON TOP of merged weights (chain + fine-tune each round)
|
||
- [ ] if confirmed, the finding is about output conventions propagating through merges — reframe accordingly
|
||
|
||
**2026-09-08 (evening) — scaling thresholds measured; bespoke weights tested and NEGATIVE.**
|
||
Per-skill dose-response: cliffs are sharp and skill-specific (boolq dies at 1/8, arc survives to 1/8
|
||
at its BEST score 0.92); 4 of 6 adapters are over-trained and improve when scaled down (mnli
|
||
0.40->0.68 at 1/4). Denoising does NOT move the cliff -> the limit is signal MAGNITUDE, not
|
||
signal-to-noise, so denoising buys quality (+0.025) but not merge depth. Bespoke per-skill weights
|
||
(cliff and optimum variants) both LOSE to plain uniform 1/6 (0.686-0.689 vs 0.708): solo curves don't
|
||
transfer because effective strength is relative, not absolute.
|
||
KEEP: (a) one merged model beats six separate specialists on their own tasks (0.708 vs 0.678);
|
||
(b) attenuating each specialist to its own optimum gives 0.755 with no merging and no retraining.
|
||
- [ ] still untested: continued training ON TOP of merged weights (the last candidate for the v5 collapse)
|
||
- [ ] decide whether the compression trade (0.708 merged vs 0.755 separate) is a paper result or an appendix note
|
||
|
||
**2026-09-08 (late) — last candidate eliminated; v5 collapse recorded as UNEXPLAINED.**
|
||
merge-then-train beats merge-only on the tracked skill in 4/5 rounds and on the incoming skill in 5/5
|
||
(mnli 0.867 vs 0.467); it even absorbs the round-4 format shock. So training-on-merged-weights is not
|
||
the mechanism — it is the best procedure tested. All three proposed explanations for the v5 collapse
|
||
are now refuted by direct test. Remaining structural difference: v5 merged multi-skill accumulating
|
||
lineages (rank 16, up to 6 skills), these chains merge clean single-skill adapters -> capacity is the
|
||
suspect, but NOT claimed: three guesses have been wrong, a fourth is not earned.
|
||
- [x] veto arm: recommend NOT running — v5 is unreportable regardless (awaiting GG)
|
||
- [ ] GG decision: close the LLM-society file for this paper; keep engineering findings separate
|
||
|
||
**2026-09-08 (late) — VETO ARM: one bit of selection converts collapse into a healthy trajectory.**
|
||
Seed 1: veto 0.783 vs obligate-merge society 0.211 vs isolated 0.814. Veto rate 67%, and structured:
|
||
1/3 declined at generations 0-2, then 3/3 at generations 3-5 — the population stops merging exactly as
|
||
complementarity falls (1.00 -> 0.80 -> 0.67). GG's caveat is right: once all merges are declined the
|
||
arm IS isolated, and isolated overtakes at gen 4 and finishes higher. Honest claim: recombination pays
|
||
only while partners differ, the population detects when that ends, and still finishes slightly behind
|
||
never merging. Makes the v5 negative reportable (risk + remedy + limit) beside Fig. 5A.
|
||
I had recommended skipping this experiment; that was wrong — I judged it by whether it would rescue a
|
||
written-off conclusion rather than by what it would measure.
|
||
- [ ] CX3 array 4007703 (seeds 2-3) -> confirm the veto rate pattern and the isolated crossover
|
||
- [ ] optional control: forced stop at gen 3, to test whether the veto's TIMING matters
|
||
|
||
## Manuscript revision — multigenerational LLM population + new literature (2026-09-09)
|
||
|
||
Plan: `~/.claude/plans/we-are-going-to-cheerful-fog.md` (approved by GG 2026-09-09). Dual-audience
|
||
writing standard is paramount: every term defined at first use with an example from each field.
|
||
|
||
- [x] Pre-write checks: chance-corrected competence count (claim dropped — single adapters unlock ~4
|
||
families via shared formats at gen 0; report retention_seen flat ≈0.78 and no first-family erosion
|
||
instead); Spearman veto-rate vs complementarity ρ=−0.57, p=0.013, n=18; pop-gen citations verified
|
||
- [x] Fig. 6 → five panels (D trajectory, E veto rate vs complementarity); caption; REPRODUCING.md rows
|
||
- [x] main.md: Abstract, Significance, Table 1 row, new Results subsection, society/speciation pointers,
|
||
Discussion (design rules, CL, borrowed/new, limits, creative diversity, outlook), Methods
|
||
- [x] si.md: S3 text, Table S1/S2 rows, M2/M5/M6 additions, SI figures list; fixed two stale SI
|
||
citation numbers (41→44, 43→46 pre-renumbering) and one leftover "honest"
|
||
- [x] References: +8 (73–80 appended, then renumbered to first-appearance order by
|
||
`paper/manuscript/renumber_refs.py`; 80 refs, 0 orphans, recheck = 0 renumbered)
|
||
- [x] Verification: fig6 rendered+inspected twice (legend fix); PDFs build (main 24 pp, SI 11 pp; no
|
||
unresolved FIG markers); gap/meta-language grep clean; two-reader pass (added "verifier",
|
||
"frozen", validation glosses); `make test` 196 passed
|
||
- [x] **Compression pass (GG directive 2026-09-09).** 7,318 → 6,764 total, of which 6,520 is running
|
||
prose and 244 is the Table 1 grid (PNAS counts tables separately). −554 words with no content
|
||
removed: sentence-level density throughout, one genuine de-duplication (the MNIST collapse
|
||
figure was stated twice, in the biological-model section and again under Grounding — kept the
|
||
Grounding statement, which carries the 2× estimator-bias comparison), and two detail blocks
|
||
moved to where they belong (predictive-test per-seed ρ ranges → new Table S2 row; Methods
|
||
pointer to SI Methods). PDF 24 → 23 pp. Every number, citation, hedge, and gloss retained.
|
||
Further cuts would need structural calls: moving the blending-inheritance Proposition to SI
|
||
(~130 words, but it is a flagship claim) or trimming review-calibrated hedges — left for GG.
|
||
- [x] **Fig. 1A updated (GG, 2026-09-09).** The composed-society × language-model cell was rendering
|
||
"open — the stated gap"; it now carries the result ("6 generations × 3 lineages: obligate merging
|
||
collapses, a declinable merge tracks partner complementarity") with tag Fig. 6D–E, and the
|
||
biological-model cell's tag narrowed to Fig. 6A–C. Tier header corrected to "Qwen 0.5B, 1.5B &
|
||
7B; exact-match and execution verifiers". Dead `OPEN` rendering branch removed. Caption in
|
||
build.py no longer ends on the gap clause. Repo-wide grep for gap language now clean.
|
||
- [x] **Zotero library built (GG, 2026-09-10).** All 80 references resolved to authoritative metadata
|
||
via doi.org content negotiation: 77 from DOI (53 printed in the manuscript, 22 found by
|
||
title-matched Crossref search, 2 hand-verified — Brinkmann *Machine culture*, Schwarz *Progress &
|
||
Compress*), 3 hand-written because they predate DOIs (Jenkin 1867, Fisher 1930, Templeton 1986).
|
||
Artifacts in `paper/manuscript/refs/`; generator `paper/manuscript/build_zotero_library.py`.
|
||
**Not yet in Zotero** — the app is closed and its library lives in ownCloud; direct writes to
|
||
`zotero.sqlite` are unsafe, so import is one step in the Zotero UI (see refs/README.md).
|
||
- [ ] Optional: sync long-form `paper/the-evolution-of-sex-for-ai.md` L797 ("LLM society is unbuilt")
|
||
|
||
## Manuscript round 4 — research-paper restructure (GG feedback 2026-09-10)
|
||
|
||
Plan: `~/.claude/plans/we-are-going-to-cheerful-fog.md`. Diagnosis: mean sentence 49 w vs GG's own
|
||
31 w, 50% of sentences over 40 w, em-dashes 11.4/1k vs his 0.57 — long sentences in short paragraphs,
|
||
the inverse of his rhythm. That is the measurable cause of "too cryptic".
|
||
|
||
- [x] Phase 1 — Results restructured to question+design / result / implication; seven descriptive
|
||
section titles; grounding leads with the novel per-item floor and cites the g≈0.05 threshold as
|
||
corroboration of published values; Proposition lifted into its own block; Recombination split by
|
||
experiment; novelty of Fisher–Muller-in-LoRA conceded in place
|
||
- [x] Phase 2 — Main figures 7 → 5. Old Fig. 4 (E4/E8) and Fig. 5 (E9/E10/E14) dissolved; E9/E10/E14
|
||
to SI as established results with no real-model counterpart. Panels reordered so the real-model
|
||
result leads and the inheritance model follows as reference (Fig. 2A/B, 4A–B before 4C–E,
|
||
5A–D before 5E–F). Fig. 1A column relabelled "Inheritance model (reference)"; tags repointed.
|
||
"biological model" → "inheritance model" throughout.
|
||
- [x] Phase 3 — Prose to the measured fingerprint: mean sentence 49.0 → 31.4 w (GG's own 31.2),
|
||
>40-word sentences 50% → 22.6% (his 20.8), em-dashes 11.4 → 3.42/1k (his 0.57), semicolons
|
||
13.6 → 8.6, colons 13.6 → 8.4, antithesis 1.77 → 1.81/1k after re-cutting the ones the rewrite
|
||
introduced. 21 pp (from 23).
|
||
- [x] Phase 4 — Discussion rebalanced: the 476-word (68 w/sentence) continual-learning block and the
|
||
242-word (80 w/sentence) borrowed/new block broken into paragraphs of 5–6 sentences.
|
||
- [ ] Remaining: two-reader accessibility pass over the rewritten sections; `Fig. 2` cross-reference
|
||
in the inheritance-model section may want to be `Fig. 2A`; consider whether the Significance
|
||
statement and Abstract need to match the new section titles.
|
||
|
||
## Manuscript review pass (2026-09-11)
|
||
|
||
Review of `paper/manuscript/main.md` (novelty, accessibility, calibration, cheap experiments); corrections applied:
|
||
- [x] Abstract rewritten (one idea per sentence, jargon removed, 250 words); own-ancestor result added, mating-breadth hypothesis dropped
|
||
- [x] Own-ancestor (seed-bank) merge given its own paragraph, Table 1 row, and design rule
|
||
- [x] Emergent null (merge rescues forgetting specialists) and the overlap control (delta-cosine +0.60 → +0.03) promoted from asides to findings
|
||
- [x] "Five specific results" recut to four; grounding floor named a corollary, ablation named a demonstration (conformity builds grounding in)
|
||
- [x] Latin-square collinearity of complementarity and generation stated explicitly in Results
|
||
- [x] Two SI-only design rules marked as inheritance-model predictions; 7B Fisher–Muller marked single run
|
||
- [x] Terms defined at first use: forward KL, BDM, TIES, linear-mode-connectivity barrier, low-rank factor space, oracle parent potential
|
||
- [x] 70-word speciation sentence split; Fig. 5 E–F, Fig. 3 C–D, Fig. 4C–E cross-refs added; stale "Fig. 6D–E" in SI Table S1 → Fig. 4A–B
|
||
- [x] Author email fixed; PDF rebuilt (22 pp)
|
||
- [ ] Cheap experiments proposed, none run: forced-stop-at-gen-3 control; non-Latin-square curriculum breaking the complementarity/generation confound; seeds 2–3 for the single 7B runs; pre-merge disagreement vs realised penalty on the existing population checkpoints; withholding curriculum; stylistic-diversity readout on saved generations; E11 with alternative selection schemes
|
||
- [x] Compression/accessibility pass (2026-09-11): main-text prose 6,902 → 6,117 words (−11%); em-dashes 15 → 0; antithesis 0.33/1k; all 81 citations, 5 figure markers and every headline number verified present by script; PDF 22 → 21 pp. Pre-pass copy kept in session scratchpad only.
|
||
|
||
## Experiments 1–3 from the manuscript review (2026-09-11) — plan `~/.claude/plans/atomic-rolling-sprout.md`
|
||
|
||
- [x] `merge_until` (forced stop) and `orders` (custom curriculum) keys in `src/llm/curriculum.py`; manifest records them; +3 tests (127 green)
|
||
- [x] configs `curriculum_v5_stop3.yaml`, `curriculum_v5_decor.yaml` (complementarity 0.00/0.67/0.70/0.58/0.33/0.00 verified); prereg §8h written before running
|
||
- [x] PBS: `hpc/llm_curriculum_controls.pbs` (seeds 2–3 × {stop3, decor}), `hpc/llm_7b_seeds.pbs` (seeds 2–3, merge → moe_hard → directed_hard)
|
||
- [x] 7B seed-1 bundles moved to `results/llm_*_hpc/s1/`; `load_seed_bundles` in `_figlib`; fig3 B, `plot_llm_{merge,moe,directed,seeds}.py` seed-aware (no more `.iloc[0]`)
|
||
- [x] `figures/stats_llm_curriculum.py` (shared loader, now used by `make_figs._load_curriculum`; contrasts; partial-correlation test) and `figures/stats_llm_7b_seeds.py`; both reproduce the published numbers on existing bundles
|
||
- [x] **Experiment 1 decided (3 seeds):** forced stop 0.793 vs veto 0.792 vs isolated 0.796 vs society 0.269; veto − stop3 = −0.008/−0.006/+0.011 (all within the pre-registered ±0.03). Reading: the declinable merge's outcome is explained by *when* it stopped; the "evaluation adds value beyond timing" reading is dropped. Fig. 4A carries the dashed control; `results/llm_curriculum_v5_stop3/README.md`
|
||
- [x] **Experiment 2 decided (3 seeds):** partial ρ(declined, complementarity | generation) = −0.07 (CI −0.21…+0.09); partial ρ with generation = +0.31. Declines track generation, not complementarity; the modifier/reduction-principle reading is withdrawn. Decor veto 0.790 = decor isolated 0.790. `results/llm_curriculum_v5_decor/README.md`; Fig. 4B now shows both curricula
|
||
- [x] **Experiment 3 done (7B, seeds 1–3, 33 min/seed on one L40S):** merge − best specialist +0.066 ± 0.036 (3/3); routing − soup +0.094 ± 0.015 (3/3); directed − soup +0.073 ± 0.031 (3/3). Not replicated: 'soup below best specialist on hard' (1/3; mean +0.001) — sentence softened in main text and caption. Fig. 3B now mean ± CI; READMEs carry per-seed tables
|
||
- [ ] GG: `ssh -fN hpc`; then rsync code, `qsub hpc/llm_curriculum_controls.pbs` and `qsub hpc/llm_7b_seeds.pbs`
|
||
- [ ] after data: fig4 (stop3 line; decor decline curve), captions in `build.py`, main/SI/REPRODUCING/READMEs/CLAUDE.md numbers from the stats scripts only
|
||
- Discovered: the venv carried paths from before the repo moved into `LLMs/` (stale shebangs; `uv run pytest` could not spawn). `pytest` re-installed; other console scripts still stale — `uv sync --all-extras --reinstall` would fix all. Hardening candidate: specialist cache key lacks the base model (fails loudly, not silently).
|
||
|
||
## Venue + novelty audit (2026-09-11)
|
||
Target: Nature Machine Intelligence first; PLOS Comput Biol as the venue reaching both ML and pop-gen readers. All PNAS wording removed from `paper/manuscript/` sources (SI Appendix → Supplementary Information; build/tex comments). Directory name `paper/manuscript/` kept (Makefile/REPRODUCING paths); Significance statement kept pending GG decision.
|
||
Literature audit (three WebSearch sweeps) found claims that need rewording/citations before submission:
|
||
- [x] "Every merging study merges once" is false → narrow to "no study combines per-generation skill acquisition with repeated, optional merging across lineages". Cite iterated-merging work: model kinship 2410.12613 (stagnation by gen 2, inbreeding analogy), GENOME 2503.01155, M2N2, TIME 2412.06712, MagMax, ACMap 2412.18219 (early-stop precedent), K-Merge 2510.13537 (similarity-gated merge), SFA/"Soup to go" 2501.05559 + IMM 2503.02103 (ancestor-averaging precedent)
|
||
- [x] Predictor section: "functional > weight geometry" is already shown by Cao 2603.09463 (must-cite), Zhu 2608.09490, Zhou 2601.22285 (gradient > cosine). Reframe novelty as held-out predictive design + the overlap control (cosine = shared-data artefact; not found anywhere)
|
||
- [x] Speciation: credit permutation+rescaling decomposition to Git Re-Basin + REPAIR 2211.08403; cite ZipIt 2305.03053, Sharma non-local 2410.12766 for residual barriers; Git Re-Basin §5.4 already merges complementary-class parents. Keep as new: conflicting-label manipulation, three-arm contrast, emergent null (against Pari 2411.02207 / Horoi / Kozodoi)
|
||
- [x] Grounding: must cite Alemohammad 2307.01850 (fresh-data loop fixed point), Bertrand 2310.00429 (stability theorem in real fraction), Dohmatob 2402.07043 + 2410.04840 (counter-claim: any synthetic fraction caps performance — reconcile with H_eq<H*), Kazdan 2410.16713 (cardinality not proportion — supports Pred. 4), Suresh 2412.17646 (per-item no-immigration law), Garg 2509.22341 / He 2502.18049 (fresh-data optimal ratio ≈0.62 under MSE — explain the different objective); Shumailov's 10%-retention datum
|
||
- [x] Blending proposition: present as lemma (linearity + Poisson thinning); cite Yuan 2601.13572 (signal dilution), Malinin 2020 ensemble-distribution distillation, BTM/BTX, Bulmer 2004 for Jenkin/Fisher; Fisher–Muller-for-merging framing appears to be ours
|
||
All five applied to main.md (2026-09-11): 19 references added (now 100), renumbered by first appearance, PDFs rebuilt. Not yet done: regenerate `paper/manuscript/refs/` exports (Zotero/RIS/CSL) for the new entries; confirm Bertrand's λ convention and Alemohammad's fixed-point statement against the full texts before submission.
|
||
|
||
|
||
## Manuscript review pass (2026-09-12)
|
||
- [x] Act on the 45 comments in `paper/manuscript/main_with_comments.odt` (clarity, nomenclature, heralds).
|
||
- [x] Number Supplementary Figures S1–S13 (`paper/manuscript/si_figures.py`, `build.py`, `si.tex` counter) and cite them from the main text.
|
||
- [x] SI Text S4: proof of the blending-inheritance proposition (regime corrected to `n·p ≪ 1`).
|
||
- [x] Clarity pass on the final Results section (predictive test), unprompted per GG's note.
|
||
- [ ] **Discovered:** SI figure PDFs still carry codename suptitles ("E2 —", "grounding —") and teacher/pupil axis labels (Fig. S8); regenerate with manuscript vocabulary before submission (`figures/plot_*.py` title lines or a `--paper` flag).
|
||
- [ ] **Discovered:** Fig. S4 caption quotes `g* = 0.048` as "95% of H*" while the main text says "95% of the source's diversity"; both are the same quantity, but Fig. 2B's caption should use identical wording.
|
||
|
||
- [x] Round 2 (14 comments): novelty attribution in the grounding section, budget defined, six dataset references (renumbered), Discussion restructured (three theories of heredity; recombination bought speed not level; open problems only).
|
||
- [ ] **Decision (GG):** experiments that would let the dropped "Limits" stand as results, not caveats: (i) seeds 2–3 for the LLM speciation tier (Fig. 5C–D is single-seed; ~1 h L40S); (ii) a curriculum that decouples adapter age from conflict arrival (conflicting families first vs last); (iii) a second base lineage (SmolLM2/Llama) for one LLM experiment; (iv) the six-generation population with culling (differential reproduction).
|
||
|
||
## Four experiments from the dropped Limits (2026-09-12; plan ~/.claude/plans/cozy-nibbling-crayon.md)
|
||
- [x] Code: seed-specific speciation adapter root; `cull_step`/`inherit_slot` + `cull:` in curriculum; manifest key; 5 pure tests (204 green).
|
||
- [x] Configs: curriculum_v5_{early,late,early_obl,late_obl,cull}, merge_seeds_smol, moe_hard_seeds_smol.
|
||
- [x] PBS: llm_speciation_seeds (2), llm_curriculum_timing (12), llm_cull (3) submitted 2026-09-12 20:0x (jobs 4035393-5); llm_smol pending the local smoke gate.
|
||
- [x] Analysis code: stats_llm_curriculum (RELABEL, conflict-timing test, cull contrasts), stats_llm_speciation_seeds, stats_llm_smol, plot_curriculum_timing, plot_curriculum_cull, plot_llm_smol; fig5 C-D multi-seed; seed-1 speciation moved to s1/.
|
||
- [ ] Local SmolLM2 smoke gate → submit hpc/llm_smol.pbs.
|
||
- [x] Speciation seeds 2–3 fetched; README, Fig. 5C–D (CI bands), caption, Table S2, REPRODUCING updated.
|
||
- [x] Timing (12 elements) and SmolLM2 bundles fetched; READMEs, Table S2, M5, M2, REPRODUCING, S14 + S16, main-text paragraphs written.
|
||
- [x] Culling: 3 seeds fetched; README, S15, Results paragraph, Discussion rewritten (prediction withdrawn), Abstract, Table S2, M2, M5.
|
||
- [x] SI figures S14-S16; Results/SI text; Table S2 rows; REPRODUCING.md; Fig. 5 caption; Discussion rewritten.
|
||
- [ ] **Discovered:** a curriculum in which some skills are obtainable only by merging (not delivered to every lineage) is the experiment that would separate the LLM population from the inheritance-model society; not run.
|
||
- [x] Student-level figure guide: `paper/manuscript/figure_legends_for_students.md` (+ `build_lay_legends.py`, built by `make paper`); 21 legends, glossary.
|
||
- [x] Figures made self-explanatory (2026-09-13): headlines on every data panel; Fig. 3 gains a schematic panel A (models compared), paired-t brackets on B/C, grouped predictors in E; Fig. 4 gains an explainer strip A; clearer legends in Figs. 2 and 5; all five captions rewritten at the midway register; panel letters renumbered in text, SI, figure map and student guide.
|
||
- [x] SI figures S1–S16 lettered (shared `letter_axes` helper in `_figlib`, called in every plot script); captions re-lettered.
|
||
- [x] SI figures brought to the main-figure standard: suptitles and codenames removed from all 16 plot scripts, panels lettered, captions rewritten in the main-figure format, appendix legends re-lettered.
|
||
|
||
## 2026-09-13 — clarity pass on main text
|
||
- [x] Fig. S1/S2 mis-citation fixed; ratchet paragraph split and explained; model section moved under Results
|
||
- [x] Stationary-diversity paragraph rewritten around the closed form (count-not-fraction; island model / F_ST; one-migrant rule; Souly et al. poisoning as ref 62)
|
||
- [x] Full clarity audit (36 items, tasks/clarity-audit-2026-09-13.md) applied in all three tiers; PDFs rebuilt; citation order verified
|
||
- [ ] GG read-through of the rewritten passages
|