Manuscript revision and pending experiment work, snapshot before restructuring
Clarity pass over the main text (36-item audit), Discussion rewrite and cut, acknowledgements, Souly et al. as ref 62, lettered SI panels, model section moved under Results; plus the untracked curriculum/society/compose/smol configs, runners, figures, stats and tests that the SI already cites. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
This commit is contained in:
parent
e4804adabc
commit
84124de143
450 changed files with 52813 additions and 1202 deletions
245
tasks/todo.md
245
tasks/todo.md
|
|
@ -412,3 +412,248 @@ offspring screened on the arm's own signal), QD selection. `src/llm/society.py`
|
|||
155 green), `kind: llm_society`, configs `society_smoke.yaml` / `society.yaml`. Stages: smoke
|
||||
(local, ~15 min) → pilot full vs no_grounding (GG gate) → 4-arm × 3-seed CX3 campaign → figure +
|
||||
manuscript fold-in.
|
||||
|
||||
**2026-09-07 — v1 `llm_society` campaign landed (4 seeds) and is NEGATIVE; v2 pre-registered.**
|
||||
Best-agent overall at gen 9, 3-seed means: `no_sex` 0.558 ≥ `no_diversity` 0.539 ≥ `full` 0.506 ≫
|
||||
`no_grounding` 0.436 (worst arm in every seed from gen 2). Conformity−truth gap does not separate the
|
||||
arms. Read through the framework the null was structurally guaranteed (near-clone founders over 3
|
||||
families; 2³ competence states; linear blending at 0.5B = the dilution regime; parents truncated
|
||||
before breeding, unlike E11's survival-over-pool; `n_test`=40 → SE 0.079) — details and fixes in
|
||||
`tasks/prereg-llm-society-v2.md` §1, lesson in `tasks/lessons.md`. Nothing enters the manuscript;
|
||||
Fig. 1A's "stated gap" stands. **v2** (`kind: llm_society_v2`): L=12 families / one founder each,
|
||||
confidence-routed union inheritance, pooled survival, checkpoint+resume, 240 test items, g=0.85,
|
||||
G=12; six numerical hypotheses H1–H6; calibration gates C1–C5 must pass before submission (GG
|
||||
reviews). GG decisions: 0.5B; `sex_linear` dropped (H2 deferred); no family vetoes.
|
||||
- [x] families / operators / v2 loop / calibration runner / configs / PBS array / figure script / 9 tests (164 green)
|
||||
- [x] smoke (4 arms, figure + stats script) → calibration A pass 1 (6/17 in band) → pass 2 (9 in band; C1b needed the gate re-derived 0.35→0.41 from the grid) → GG chose L=9
|
||||
- [x] calibration B: C2 FAILED as pre-registered (retention ≤0.81 at k≤150; interference, not the observation floor) → C2b: k=300 + confidence gate τ=0.5 gives mean retention 0.87 (PASS); C3 operator half passes (union holds both families, linear loses one), retention half re-run gated; C5 passes (consensus 0.31)
|
||||
- [x] campaign configs set: L=9, k_inherit=300, conf_gate=0.5, epochs=3, n_test 27/family, g=0.85, G=12; PBS 16 elements × 8 h
|
||||
- [x] gated cross (C3): two-skill child plateaus at ~0.85×/0.8× of parents at any budget (3 vs 6 epochs; r64 hurts); tight gate τ=0.85 gives the 6-epoch retention at 3 epochs
|
||||
- [x] **GG go/no-go (21:30): NO-GO at 0.5B** — the vertical claim needs 5–6 co-resident skills the r=16 adapter cannot hold; today = a measured transmission ceiling (SI material). 7B plan drafted: prereg §13
|
||||
- [ ] GG: 7B scope (headline ~80 / H3+H4 ~120 / full ~260 L40S-h), Phase-0 task design go, SI text timing
|
||||
- [ ] SI: the 0.5B calibration ceiling as the reason the tier was not run (prereg §11 row 4) — three limits, numbers from results/llm_society_v2_calib_*
|
||||
- [ ] REPRODUCING.md: rows for the v1 campaign (4 seeds), v2 smoke, and the 9 calibration bundles
|
||||
- [ ] stage code on CX3, `qsub hpc/llm_society_v2.pbs` (16 elements); local hedge = `society_v2_s1.yaml`
|
||||
- [ ] `figures/stats_llm_society.py` (per-seed paired contrasts H1/H3/H4, AUC for H5, supplied-vs-retained for H6)
|
||||
- [ ] fold the outcome per prereg §11
|
||||
|
||||
**2026-09-08 — v3 `llm_compose` run (3 seeds): H1 PASS, H2–H5 null; design could not show the claim.**
|
||||
Composition at gen 0 is real and replicated (surplus +0.087/+0.033/+0.093; union-exceedance ~0.12; also
|
||||
on MATH-500). Decay hypotheses uninterpretable: the lineages barely drifted (q_math 0.54 → 0.50–0.57)
|
||||
and, more fundamentally, a fixed skill set has its ceiling at gen 0 — GG: "are models learning NEW
|
||||
skills at EACH generation? … that was not the problem being addressed." v3 was Weismannian (fresh LoRA
|
||||
each generation) and retention-only. Two of my errors: C3 unchecked (code specialist 0.075 on MBPP →
|
||||
q_code noise), and three premature reads of a single-seed trajectory.
|
||||
**v4 `llm_curriculum` — continual learning in a population** (`prereg-llm-society-v4.md`): Lamarckian
|
||||
channel (`continue_lora_training`), Latin-square curriculum (complementarity 1.0 → 0.0 by construction,
|
||||
H6 predicts the *shape*), arms isolated/society/society_dry/seed_bank (GG's ancestor-merge idea —
|
||||
temporal vs spatial complementarity, direction uncommitted), single-shot SoTA baselines at matched budget
|
||||
as the falsifier. GG decisions: 3×9×9, replay fixed-total, baselines get the same directed selection.
|
||||
- [x] G0 PASS (accumulation 0.74 → 0.95); G2 FAIL ×2 (v2 families don't interfere; one pair at +0.65);
|
||||
G3 negative (merging costs −0.01…−0.13 with nothing to repair); base = 0.094 → one family lifts all to 0.417
|
||||
- [x] **v5 curriculum**: 11 real-dataset families, 5 answer shapes, per-family verifiers, disjoint splits
|
||||
(`curriculum_data.py`, +6 tests, 56 green); selection rule fixed in prereg §8a
|
||||
- [ ] stage A calibration running (`curriculum-v5-calib`): base + 11 specialists × 11 families
|
||||
- [ ] stage B: zero-replay forgetting probe on the survivors (mean drop ≥ 0.15, not single-family)
|
||||
- [ ] GG go/no-go → seed 1 local + CX3 array (seeds 2–3), then baselines
|
||||
|
||||
**2026-09-08 — v5 curriculum campaign done (3 seeds); all hypotheses fail; the useful finding is a
|
||||
split in the theory.** Real-dataset curriculum (6 families, 5 answer formats, per-family verifiers)
|
||||
replaced the procedural set. Results: not-merging wins (0.80), merging-with-own-ancestor middling
|
||||
(0.66), merging-with-a-peer collapses (0.27), single-shot merging unstable (0.125–0.764). Cause:
|
||||
two families are answer-format destroyers that propagate through merges and compound because
|
||||
offspring continue the lineage. Scope limits: merging was obligate (no veto) and there is NO
|
||||
selection between lineages — a gene-flow experiment, not a selection one.
|
||||
**Key new measurement:** same-skill adapters (seed/data draw only) are near-orthogonal in weight
|
||||
space (cos +0.006), disagree on 24% of prompts, and **merging them beats the best parent by +0.087,
|
||||
exactly at the either-right ceiling**. So decorrelation-in-what-you-know is harmful while
|
||||
decorrelation-in-how-you-encode-it is beneficial — the framework's single rho conflates them.
|
||||
- [ ] v6: three arms (no-merge · complementary · parallel) + veto + population selection (~30 GPU-h)
|
||||
- [ ] decide how the E9-risk result and the two-variations split enter the manuscript (beside Fig. 5A)
|
||||
|
||||
**2026-09-08 (later) — four mechanism probes; two of my explanations retracted.** (1) Same-skill
|
||||
adapters: 85% of a LoRA's change is run-specific noise; merging two beats the better parent by +0.087,
|
||||
at the either-right ceiling. (2) Denoising before crossing adds +0.025 on both skills at once
|
||||
(inbred-lines signature). (3) A single merge of clean adapters is PROTECTIVE (0.825 vs 0.550 best
|
||||
parent) — retracts "destructive skill propagates through merges". (4) Five chained convex merges lose
|
||||
nothing, while signal-preserving additive weights collapse (1.02 vs 0.52 retention) — retracts
|
||||
"geometric signal dilution"; the real constraint is bounding drift from the base. (5) Scaling probe
|
||||
(GG's control): base 0.000, and the adapter works down to 1/8 then dies — 1/16 = 0.450, 1/32 = 0.000.
|
||||
So the chain's apparent retention was ANSWER FORMAT supplied by the dominant partner, not the skill.
|
||||
Consistent with Fig. 3C-D: functional conflict predicts merge damage, weight geometry does not.
|
||||
- [ ] test the remaining candidate: continued training ON TOP of merged weights (chain + fine-tune each round)
|
||||
- [ ] if confirmed, the finding is about output conventions propagating through merges — reframe accordingly
|
||||
|
||||
**2026-09-08 (evening) — scaling thresholds measured; bespoke weights tested and NEGATIVE.**
|
||||
Per-skill dose-response: cliffs are sharp and skill-specific (boolq dies at 1/8, arc survives to 1/8
|
||||
at its BEST score 0.92); 4 of 6 adapters are over-trained and improve when scaled down (mnli
|
||||
0.40->0.68 at 1/4). Denoising does NOT move the cliff -> the limit is signal MAGNITUDE, not
|
||||
signal-to-noise, so denoising buys quality (+0.025) but not merge depth. Bespoke per-skill weights
|
||||
(cliff and optimum variants) both LOSE to plain uniform 1/6 (0.686-0.689 vs 0.708): solo curves don't
|
||||
transfer because effective strength is relative, not absolute.
|
||||
KEEP: (a) one merged model beats six separate specialists on their own tasks (0.708 vs 0.678);
|
||||
(b) attenuating each specialist to its own optimum gives 0.755 with no merging and no retraining.
|
||||
- [ ] still untested: continued training ON TOP of merged weights (the last candidate for the v5 collapse)
|
||||
- [ ] decide whether the compression trade (0.708 merged vs 0.755 separate) is a paper result or an appendix note
|
||||
|
||||
**2026-09-08 (late) — last candidate eliminated; v5 collapse recorded as UNEXPLAINED.**
|
||||
merge-then-train beats merge-only on the tracked skill in 4/5 rounds and on the incoming skill in 5/5
|
||||
(mnli 0.867 vs 0.467); it even absorbs the round-4 format shock. So training-on-merged-weights is not
|
||||
the mechanism — it is the best procedure tested. All three proposed explanations for the v5 collapse
|
||||
are now refuted by direct test. Remaining structural difference: v5 merged multi-skill accumulating
|
||||
lineages (rank 16, up to 6 skills), these chains merge clean single-skill adapters -> capacity is the
|
||||
suspect, but NOT claimed: three guesses have been wrong, a fourth is not earned.
|
||||
- [x] veto arm: recommend NOT running — v5 is unreportable regardless (awaiting GG)
|
||||
- [ ] GG decision: close the LLM-society file for this paper; keep engineering findings separate
|
||||
|
||||
**2026-09-08 (late) — VETO ARM: one bit of selection converts collapse into a healthy trajectory.**
|
||||
Seed 1: veto 0.783 vs obligate-merge society 0.211 vs isolated 0.814. Veto rate 67%, and structured:
|
||||
1/3 declined at generations 0-2, then 3/3 at generations 3-5 — the population stops merging exactly as
|
||||
complementarity falls (1.00 -> 0.80 -> 0.67). GG's caveat is right: once all merges are declined the
|
||||
arm IS isolated, and isolated overtakes at gen 4 and finishes higher. Honest claim: recombination pays
|
||||
only while partners differ, the population detects when that ends, and still finishes slightly behind
|
||||
never merging. Makes the v5 negative reportable (risk + remedy + limit) beside Fig. 5A.
|
||||
I had recommended skipping this experiment; that was wrong — I judged it by whether it would rescue a
|
||||
written-off conclusion rather than by what it would measure.
|
||||
- [ ] CX3 array 4007703 (seeds 2-3) -> confirm the veto rate pattern and the isolated crossover
|
||||
- [ ] optional control: forced stop at gen 3, to test whether the veto's TIMING matters
|
||||
|
||||
## Manuscript revision — multigenerational LLM population + new literature (2026-09-09)
|
||||
|
||||
Plan: `~/.claude/plans/we-are-going-to-cheerful-fog.md` (approved by GG 2026-09-09). Dual-audience
|
||||
writing standard is paramount: every term defined at first use with an example from each field.
|
||||
|
||||
- [x] Pre-write checks: chance-corrected competence count (claim dropped — single adapters unlock ~4
|
||||
families via shared formats at gen 0; report retention_seen flat ≈0.78 and no first-family erosion
|
||||
instead); Spearman veto-rate vs complementarity ρ=−0.57, p=0.013, n=18; pop-gen citations verified
|
||||
- [x] Fig. 6 → five panels (D trajectory, E veto rate vs complementarity); caption; REPRODUCING.md rows
|
||||
- [x] main.md: Abstract, Significance, Table 1 row, new Results subsection, society/speciation pointers,
|
||||
Discussion (design rules, CL, borrowed/new, limits, creative diversity, outlook), Methods
|
||||
- [x] si.md: S3 text, Table S1/S2 rows, M2/M5/M6 additions, SI figures list; fixed two stale SI
|
||||
citation numbers (41→44, 43→46 pre-renumbering) and one leftover "honest"
|
||||
- [x] References: +8 (73–80 appended, then renumbered to first-appearance order by
|
||||
`paper/pnas/renumber_refs.py`; 80 refs, 0 orphans, recheck = 0 renumbered)
|
||||
- [x] Verification: fig6 rendered+inspected twice (legend fix); PDFs build (main 24 pp, SI 11 pp; no
|
||||
unresolved FIG markers); gap/meta-language grep clean; two-reader pass (added "verifier",
|
||||
"frozen", validation glosses); `make test` 196 passed
|
||||
- [x] **Compression pass (GG directive 2026-09-09).** 7,318 → 6,764 total, of which 6,520 is running
|
||||
prose and 244 is the Table 1 grid (PNAS counts tables separately). −554 words with no content
|
||||
removed: sentence-level density throughout, one genuine de-duplication (the MNIST collapse
|
||||
figure was stated twice, in the biological-model section and again under Grounding — kept the
|
||||
Grounding statement, which carries the 2× estimator-bias comparison), and two detail blocks
|
||||
moved to where they belong (predictive-test per-seed ρ ranges → new Table S2 row; Methods
|
||||
pointer to SI Methods). PDF 24 → 23 pp. Every number, citation, hedge, and gloss retained.
|
||||
Further cuts would need structural calls: moving the blending-inheritance Proposition to SI
|
||||
(~130 words, but it is a flagship claim) or trimming review-calibrated hedges — left for GG.
|
||||
- [x] **Fig. 1A updated (GG, 2026-09-09).** The composed-society × language-model cell was rendering
|
||||
"open — the stated gap"; it now carries the result ("6 generations × 3 lineages: obligate merging
|
||||
collapses, a declinable merge tracks partner complementarity") with tag Fig. 6D–E, and the
|
||||
biological-model cell's tag narrowed to Fig. 6A–C. Tier header corrected to "Qwen 0.5B, 1.5B &
|
||||
7B; exact-match and execution verifiers". Dead `OPEN` rendering branch removed. Caption in
|
||||
build.py no longer ends on the gap clause. Repo-wide grep for gap language now clean.
|
||||
- [x] **Zotero library built (GG, 2026-09-10).** All 80 references resolved to authoritative metadata
|
||||
via doi.org content negotiation: 77 from DOI (53 printed in the manuscript, 22 found by
|
||||
title-matched Crossref search, 2 hand-verified — Brinkmann *Machine culture*, Schwarz *Progress &
|
||||
Compress*), 3 hand-written because they predate DOIs (Jenkin 1867, Fisher 1930, Templeton 1986).
|
||||
Artifacts in `paper/pnas/refs/`; generator `paper/pnas/build_zotero_library.py`.
|
||||
**Not yet in Zotero** — the app is closed and its library lives in ownCloud; direct writes to
|
||||
`zotero.sqlite` are unsafe, so import is one step in the Zotero UI (see refs/README.md).
|
||||
- [ ] Optional: sync long-form `paper/the-evolution-of-sex-for-ai.md` L797 ("LLM society is unbuilt")
|
||||
|
||||
## Manuscript round 4 — research-paper restructure (GG feedback 2026-09-10)
|
||||
|
||||
Plan: `~/.claude/plans/we-are-going-to-cheerful-fog.md`. Diagnosis: mean sentence 49 w vs GG's own
|
||||
31 w, 50% of sentences over 40 w, em-dashes 11.4/1k vs his 0.57 — long sentences in short paragraphs,
|
||||
the inverse of his rhythm. That is the measurable cause of "too cryptic".
|
||||
|
||||
- [x] Phase 1 — Results restructured to question+design / result / implication; seven descriptive
|
||||
section titles; grounding leads with the novel per-item floor and cites the g≈0.05 threshold as
|
||||
corroboration of published values; Proposition lifted into its own block; Recombination split by
|
||||
experiment; novelty of Fisher–Muller-in-LoRA conceded in place
|
||||
- [x] Phase 2 — Main figures 7 → 5. Old Fig. 4 (E4/E8) and Fig. 5 (E9/E10/E14) dissolved; E9/E10/E14
|
||||
to SI as established results with no real-model counterpart. Panels reordered so the real-model
|
||||
result leads and the inheritance model follows as reference (Fig. 2A/B, 4A–B before 4C–E,
|
||||
5A–D before 5E–F). Fig. 1A column relabelled "Inheritance model (reference)"; tags repointed.
|
||||
"biological model" → "inheritance model" throughout.
|
||||
- [x] Phase 3 — Prose to the measured fingerprint: mean sentence 49.0 → 31.4 w (GG's own 31.2),
|
||||
>40-word sentences 50% → 22.6% (his 20.8), em-dashes 11.4 → 3.42/1k (his 0.57), semicolons
|
||||
13.6 → 8.6, colons 13.6 → 8.4, antithesis 1.77 → 1.81/1k after re-cutting the ones the rewrite
|
||||
introduced. 21 pp (from 23).
|
||||
- [x] Phase 4 — Discussion rebalanced: the 476-word (68 w/sentence) continual-learning block and the
|
||||
242-word (80 w/sentence) borrowed/new block broken into paragraphs of 5–6 sentences.
|
||||
- [ ] Remaining: two-reader accessibility pass over the rewritten sections; `Fig. 2` cross-reference
|
||||
in the inheritance-model section may want to be `Fig. 2A`; consider whether the Significance
|
||||
statement and Abstract need to match the new section titles.
|
||||
|
||||
## Manuscript review pass (2026-09-11)
|
||||
|
||||
Review of `paper/pnas/main.md` (novelty, accessibility, calibration, cheap experiments); corrections applied:
|
||||
- [x] Abstract rewritten (one idea per sentence, jargon removed, 250 words); own-ancestor result added, mating-breadth hypothesis dropped
|
||||
- [x] Own-ancestor (seed-bank) merge given its own paragraph, Table 1 row, and design rule
|
||||
- [x] Emergent null (merge rescues forgetting specialists) and the overlap control (delta-cosine +0.60 → +0.03) promoted from asides to findings
|
||||
- [x] "Five specific results" recut to four; grounding floor named a corollary, ablation named a demonstration (conformity builds grounding in)
|
||||
- [x] Latin-square collinearity of complementarity and generation stated explicitly in Results
|
||||
- [x] Two SI-only design rules marked as inheritance-model predictions; 7B Fisher–Muller marked single run
|
||||
- [x] Terms defined at first use: forward KL, BDM, TIES, linear-mode-connectivity barrier, low-rank factor space, oracle parent potential
|
||||
- [x] 70-word speciation sentence split; Fig. 5 E–F, Fig. 3 C–D, Fig. 4C–E cross-refs added; stale "Fig. 6D–E" in SI Table S1 → Fig. 4A–B
|
||||
- [x] Author email fixed; PDF rebuilt (22 pp)
|
||||
- [ ] Cheap experiments proposed, none run: forced-stop-at-gen-3 control; non-Latin-square curriculum breaking the complementarity/generation confound; seeds 2–3 for the single 7B runs; pre-merge disagreement vs realised penalty on the existing population checkpoints; withholding curriculum; stylistic-diversity readout on saved generations; E11 with alternative selection schemes
|
||||
- [x] Compression/accessibility pass (2026-09-11): main-text prose 6,902 → 6,117 words (−11%); em-dashes 15 → 0; antithesis 0.33/1k; all 81 citations, 5 figure markers and every headline number verified present by script; PDF 22 → 21 pp. Pre-pass copy kept in session scratchpad only.
|
||||
|
||||
## Experiments 1–3 from the manuscript review (2026-09-11) — plan `~/.claude/plans/atomic-rolling-sprout.md`
|
||||
|
||||
- [x] `merge_until` (forced stop) and `orders` (custom curriculum) keys in `src/llm/curriculum.py`; manifest records them; +3 tests (127 green)
|
||||
- [x] configs `curriculum_v5_stop3.yaml`, `curriculum_v5_decor.yaml` (complementarity 0.00/0.67/0.70/0.58/0.33/0.00 verified); prereg §8h written before running
|
||||
- [x] PBS: `hpc/llm_curriculum_controls.pbs` (seeds 2–3 × {stop3, decor}), `hpc/llm_7b_seeds.pbs` (seeds 2–3, merge → moe_hard → directed_hard)
|
||||
- [x] 7B seed-1 bundles moved to `results/llm_*_hpc/s1/`; `load_seed_bundles` in `_figlib`; fig3 B, `plot_llm_{merge,moe,directed,seeds}.py` seed-aware (no more `.iloc[0]`)
|
||||
- [x] `figures/stats_llm_curriculum.py` (shared loader, now used by `make_figs._load_curriculum`; contrasts; partial-correlation test) and `figures/stats_llm_7b_seeds.py`; both reproduce the published numbers on existing bundles
|
||||
- [x] **Experiment 1 decided (3 seeds):** forced stop 0.793 vs veto 0.792 vs isolated 0.796 vs society 0.269; veto − stop3 = −0.008/−0.006/+0.011 (all within the pre-registered ±0.03). Reading: the declinable merge's outcome is explained by *when* it stopped; the "evaluation adds value beyond timing" reading is dropped. Fig. 4A carries the dashed control; `results/llm_curriculum_v5_stop3/README.md`
|
||||
- [x] **Experiment 2 decided (3 seeds):** partial ρ(declined, complementarity | generation) = −0.07 (CI −0.21…+0.09); partial ρ with generation = +0.31. Declines track generation, not complementarity; the modifier/reduction-principle reading is withdrawn. Decor veto 0.790 = decor isolated 0.790. `results/llm_curriculum_v5_decor/README.md`; Fig. 4B now shows both curricula
|
||||
- [x] **Experiment 3 done (7B, seeds 1–3, 33 min/seed on one L40S):** merge − best specialist +0.066 ± 0.036 (3/3); routing − soup +0.094 ± 0.015 (3/3); directed − soup +0.073 ± 0.031 (3/3). Not replicated: 'soup below best specialist on hard' (1/3; mean +0.001) — sentence softened in main text and caption. Fig. 3B now mean ± CI; READMEs carry per-seed tables
|
||||
- [ ] GG: `ssh -fN hpc`; then rsync code, `qsub hpc/llm_curriculum_controls.pbs` and `qsub hpc/llm_7b_seeds.pbs`
|
||||
- [ ] after data: fig4 (stop3 line; decor decline curve), captions in `build.py`, main/SI/REPRODUCING/READMEs/CLAUDE.md numbers from the stats scripts only
|
||||
- Discovered: the venv carried paths from before the repo moved into `LLMs/` (stale shebangs; `uv run pytest` could not spawn). `pytest` re-installed; other console scripts still stale — `uv sync --all-extras --reinstall` would fix all. Hardening candidate: specialist cache key lacks the base model (fails loudly, not silently).
|
||||
|
||||
## Venue + novelty audit (2026-09-11)
|
||||
Target: Nature Machine Intelligence first; PLOS Comput Biol as the venue reaching both ML and pop-gen readers. All PNAS wording removed from `paper/pnas/` sources (SI Appendix → Supplementary Information; build/tex comments). Directory name `paper/pnas/` kept (Makefile/REPRODUCING paths); Significance statement kept pending GG decision.
|
||||
Literature audit (three WebSearch sweeps) found claims that need rewording/citations before submission:
|
||||
- [x] "Every merging study merges once" is false → narrow to "no study combines per-generation skill acquisition with repeated, optional merging across lineages". Cite iterated-merging work: model kinship 2410.12613 (stagnation by gen 2, inbreeding analogy), GENOME 2503.01155, M2N2, TIME 2412.06712, MagMax, ACMap 2412.18219 (early-stop precedent), K-Merge 2510.13537 (similarity-gated merge), SFA/"Soup to go" 2501.05559 + IMM 2503.02103 (ancestor-averaging precedent)
|
||||
- [x] Predictor section: "functional > weight geometry" is already shown by Cao 2603.09463 (must-cite), Zhu 2608.09490, Zhou 2601.22285 (gradient > cosine). Reframe novelty as held-out predictive design + the overlap control (cosine = shared-data artefact; not found anywhere)
|
||||
- [x] Speciation: credit permutation+rescaling decomposition to Git Re-Basin + REPAIR 2211.08403; cite ZipIt 2305.03053, Sharma non-local 2410.12766 for residual barriers; Git Re-Basin §5.4 already merges complementary-class parents. Keep as new: conflicting-label manipulation, three-arm contrast, emergent null (against Pari 2411.02207 / Horoi / Kozodoi)
|
||||
- [x] Grounding: must cite Alemohammad 2307.01850 (fresh-data loop fixed point), Bertrand 2310.00429 (stability theorem in real fraction), Dohmatob 2402.07043 + 2410.04840 (counter-claim: any synthetic fraction caps performance — reconcile with H_eq<H*), Kazdan 2410.16713 (cardinality not proportion — supports Pred. 4), Suresh 2412.17646 (per-item no-immigration law), Garg 2509.22341 / He 2502.18049 (fresh-data optimal ratio ≈0.62 under MSE — explain the different objective); Shumailov's 10%-retention datum
|
||||
- [x] Blending proposition: present as lemma (linearity + Poisson thinning); cite Yuan 2601.13572 (signal dilution), Malinin 2020 ensemble-distribution distillation, BTM/BTX, Bulmer 2004 for Jenkin/Fisher; Fisher–Muller-for-merging framing appears to be ours
|
||||
All five applied to main.md (2026-09-11): 19 references added (now 100), renumbered by first appearance, PDFs rebuilt. Not yet done: regenerate `paper/pnas/refs/` exports (Zotero/RIS/CSL) for the new entries; confirm Bertrand's λ convention and Alemohammad's fixed-point statement against the full texts before submission.
|
||||
|
||||
|
||||
## Manuscript review pass (2026-09-12)
|
||||
- [x] Act on the 45 comments in `paper/pnas/main_with_comments.odt` (clarity, nomenclature, heralds).
|
||||
- [x] Number Supplementary Figures S1–S13 (`paper/pnas/si_figures.py`, `build.py`, `si.tex` counter) and cite them from the main text.
|
||||
- [x] SI Text S4: proof of the blending-inheritance proposition (regime corrected to `n·p ≪ 1`).
|
||||
- [x] Clarity pass on the final Results section (predictive test), unprompted per GG's note.
|
||||
- [ ] **Discovered:** SI figure PDFs still carry codename suptitles ("E2 —", "grounding —") and teacher/pupil axis labels (Fig. S8); regenerate with manuscript vocabulary before submission (`figures/plot_*.py` title lines or a `--paper` flag).
|
||||
- [ ] **Discovered:** Fig. S4 caption quotes `g* = 0.048` as "95% of H*" while the main text says "95% of the source's diversity"; both are the same quantity, but Fig. 2B's caption should use identical wording.
|
||||
|
||||
- [x] Round 2 (14 comments): novelty attribution in the grounding section, budget defined, six dataset references (renumbered), Discussion restructured (three theories of heredity; recombination bought speed not level; open problems only).
|
||||
- [ ] **Decision (GG):** experiments that would let the dropped "Limits" stand as results, not caveats: (i) seeds 2–3 for the LLM speciation tier (Fig. 5C–D is single-seed; ~1 h L40S); (ii) a curriculum that decouples adapter age from conflict arrival (conflicting families first vs last); (iii) a second base lineage (SmolLM2/Llama) for one LLM experiment; (iv) the six-generation population with culling (differential reproduction).
|
||||
|
||||
## Four experiments from the dropped Limits (2026-09-12; plan ~/.claude/plans/cozy-nibbling-crayon.md)
|
||||
- [x] Code: seed-specific speciation adapter root; `cull_step`/`inherit_slot` + `cull:` in curriculum; manifest key; 5 pure tests (204 green).
|
||||
- [x] Configs: curriculum_v5_{early,late,early_obl,late_obl,cull}, merge_seeds_smol, moe_hard_seeds_smol.
|
||||
- [x] PBS: llm_speciation_seeds (2), llm_curriculum_timing (12), llm_cull (3) submitted 2026-09-12 20:0x (jobs 4035393-5); llm_smol pending the local smoke gate.
|
||||
- [x] Analysis code: stats_llm_curriculum (RELABEL, conflict-timing test, cull contrasts), stats_llm_speciation_seeds, stats_llm_smol, plot_curriculum_timing, plot_curriculum_cull, plot_llm_smol; fig5 C-D multi-seed; seed-1 speciation moved to s1/.
|
||||
- [ ] Local SmolLM2 smoke gate → submit hpc/llm_smol.pbs.
|
||||
- [x] Speciation seeds 2–3 fetched; README, Fig. 5C–D (CI bands), caption, Table S2, REPRODUCING updated.
|
||||
- [x] Timing (12 elements) and SmolLM2 bundles fetched; READMEs, Table S2, M5, M2, REPRODUCING, S14 + S16, main-text paragraphs written.
|
||||
- [x] Culling: 3 seeds fetched; README, S15, Results paragraph, Discussion rewritten (prediction withdrawn), Abstract, Table S2, M2, M5.
|
||||
- [x] SI figures S14-S16; Results/SI text; Table S2 rows; REPRODUCING.md; Fig. 5 caption; Discussion rewritten.
|
||||
- [ ] **Discovered:** a curriculum in which some skills are obtainable only by merging (not delivered to every lineage) is the experiment that would separate the LLM population from the inheritance-model society; not run.
|
||||
- [x] Student-level figure guide: `paper/pnas/figure_legends_for_students.md` (+ `build_lay_legends.py`, built by `make paper`); 21 legends, glossary.
|
||||
- [x] Figures made self-explanatory (2026-09-13): headlines on every data panel; Fig. 3 gains a schematic panel A (models compared), paired-t brackets on B/C, grouped predictors in E; Fig. 4 gains an explainer strip A; clearer legends in Figs. 2 and 5; all five captions rewritten at the midway register; panel letters renumbered in text, SI, figure map and student guide.
|
||||
- [x] SI figures S1–S16 lettered (shared `letter_axes` helper in `_figlib`, called in every plot script); captions re-lettered.
|
||||
- [x] SI figures brought to the main-figure standard: suptitles and codenames removed from all 16 plot scripts, panels lettered, captions rewritten in the main-figure format, appendix legends re-lettered.
|
||||
|
||||
## 2026-09-13 — clarity pass on main text
|
||||
- [x] Fig. S1/S2 mis-citation fixed; ratchet paragraph split and explained; model section moved under Results
|
||||
- [x] Stationary-diversity paragraph rewritten around the closed form (count-not-fraction; island model / F_ST; one-migrant rule; Souly et al. poisoning as ref 62)
|
||||
- [x] Full clarity audit (36 items, tasks/clarity-audit-2026-09-13.md) applied in all three tiers; PDFs rebuilt; citation order verified
|
||||
- [ ] GG read-through of the rewritten passages
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue