# Clarity audit of paper/manuscript/main.md (2026-09-13) Standard: an interpretive sentence must state the concrete formula, number or mechanism it refers to; figure citations must match what the figure plots; no herald sentences; no process ghosts. Three parallel audits (Intro+model; Results 1–4; Results 5–6 + Discussion). Line numbers refer to main.md at the time of the audit. Nothing below has been applied yet. ## Tier 1 — factual or self-contradictory (fix regardless) 1. L131–132 "no amount of merging can recover it (Fig. S3)". Fig. S3 is re-baselining (E6); it has no merging arm. It shows a population that adopts its own collapsed output as reference never regains lost items. Recite it for that. 2. L595–598 Discussion: replay fractions "sit where the inheritance model's operational threshold lies". Contradicts the count-not-fraction closed form. Rewrite: they bracket 5% at n=200, and each delivers tens to thousands of replayed examples per step, past the ten copies that hold 95%. 3. L617–620 "past a threshold, become irreversible". Reintroduces the retracted threshold. The irreversibility condition is concrete: every copy gone from every parent and source (Fig. S3). 4. L276–277 "(Fig. 3B–C)" cited for a 2×2 (two sizes × easy/hard) comparison; 3B is 0.5B easy merge-vs-specialist, 3C is 7B hard routing-vs-average. Cite precisely; point to Table S2 for the rest. 5. L121–122 autoencoder "collapses faster (Fig. 2)". The comparison with drift is Fig. S2A–B; Fig. 2A shows the collapse. Cite both. 6. L107–110 three closed forms named, one written. The union formula `T[ρq + (1−ρ)(1−(1−q)^K_T)]` appears nowhere in the paper, yet Table 1 labels that row "closed form". Write it (here or in the merging section). 7. L541–544 the in-sample rank correlation of the conflict predictor is never given (only its CI and the held-out 0.35–0.40). Insert the value from Table S2. 8. L263 "at any rarity (Fig. S8)" — supported only at the tested rarities. ## Tier 2 — result named but not stated / revelation lands vague 9. L246–256 Jenkin/blending null: "exact description" asserted; the conservation law announced without saying what is conserved (expected rare-item mass q·p in the child, independent of K). 10. L277–283 headroom definition is near-tautological ("routing wins when routing would score higher"). State what sets it: 7B-easy at ceiling (1.00 on two families) → nothing to recover; 7B-hard average 0.41 vs routing 0.50 in every seed; 0.5B same gap on easy tasks. 11. L235–237 "twice the fraction ... the measured price of the estimator bias" → 10% vs 5%, because sharpening loses rare modes faster than sampling alone. 12. L636–637 "the ablation shows what removing it does" → 0.48 vs 0.78, confident and wrong. 13. L607–612 "equilibrium theory / failure theory / prediction" labels → give the three contents (m ≈ 1/p per step, 2m/(2m+1) kept; self-replay = g=0; functional disagreement predicts where weight distance does not, with shared-data control). 14. L492–494 "what moves the cliff" never stated → share of shared prompts under contradictory conventions (Fig. 5B). 15. L447–452 six-generation ceiling never named → what one adapter holds (0.80 for a single model taught the whole syllabus). 16. L577–579 and L633–635 Discussion "headroom"/"composed" carry the result → give the margins (routing +0.09, screened offspring +0.07 on hard tasks, every seed; parity at ceiling). 17. L513–515 "paid this floor" — floor undefined in main text → obligate-merge arm from generation 3. 18. L148 Table 1 "(the headroom rule)" used before defined (L281). 19. L147 Table 1 "Consequence-level only" is internal shorthand. 20. L318–321 conformity term never motivated → consensus = learning from own outputs when there is no verifier; g=0 rewards agreement with itself. ## Tier 3 — heralds, process ghosts, jargon 21. L73–76 three programme sentences announcing the paper. 22. L82 "Population genetics prices each decision." 23. L83–86 Fig. 1B "society in time rather than in space" — say what the shift buys. 24. L221–225 "What these runs add is the comparison ... exposes two departures". 25. L254–256 herald before the proposition. 26. L496 "The pre-registered emergent test constrains the claim most." 27. L331–333 "partly built in ... could alter the picture" (review-response ghost) → separability point. 28. L560–563 and L656–659 "does not validate a specifically population-genetic mechanism" said twice; keep one, as a statement about the subject. 29. L468–471 semicolon chain, "count-to-effect link"; the quadratic snowball is Orr (43), not Fig. 5E–F. 30. L292–294 clonal interference in shorthand → spell out: two variants in different individuals never meet in an asexual descendant. 31. L346–348 "the Lamarckian channel biology forbids and engineering permits" → one clause of anchor. 32. L273–275 nonlinearity caveat leaves out why the analogy holds (1/N scaling of an update, 67). 33. L103–106 Wright–Fisher used before defined; "learner, not organism" antithesis. 34. L93–94, L98–99, L111–112 housekeeping in reader-facing text ("I use the word throughout", "allele frequency of the dictionary", "standing tests in the codebase"). 35. L372–413 fourteen numeric pairs in one paragraph; split at "Three things rise with generation". 36. Fig. 3D legend colour order vs text order (conflict, overlap, duration) — check they match. ## Proposed rewrites (from the audits; GG's voice to be checked before applying) 1. "Once every copy of a rare item is gone from all parents and all sources nothing can rebuild it, and the inheritance model shows the trap in its commonest form: a population that adopts its own collapsed output as its new reference never recovers the items it had lost, whatever real data it is fed afterwards (Fig. S3), so remedies must act while copies still survive somewhere." 2. "The replay fractions the field settled on empirically (about 1% for instruction tuning, 89; 5% to 25% in continual pretraining, 90) bracket the 5% found here at 200 samples per generation, and the closed form says why they scatter: at typical batch sizes each delivers tens to thousands of replayed examples per step, well past the ten copies per generation that hold 95% of a source's diversity, so the number that matters is the count of replayed examples of each skill, not the fraction." 3. "Evaluation that averages over capabilities hides exactly the losses drift predicts first, the rare ones, and a rare capability is recoverable only while some parent or source still holds a copy (Fig. S3). Monitoring the tail, the accuracy on the rarest items rather than the mean, is therefore the leading indicator." 4. Cite "(Fig. 3C for 7B on hard tasks; the 0.5B and 7B-easy comparisons in Supplementary Information, Table S2)" — verify the SI location. 5. "(Fig. 2A; the comparison with drift in Fig. S2)". 6. "(the heterozygosity decay `E[H_t] = H_0(1 − 1/n)^t`; the stationary diversity under real data, written in the next subsection; and the expected number of rare items held by at least one of `K_T` parents, `T[ρq + (1−ρ)(1−(1−q)^K_T)]`, used in the merging section)". 7. Insert ρ and CI from Table S2. 9. "Averaging two models does to a rare capability exactly what Jenkin said blending would do to a rare variant: a child fit to the mean of `K` parents sees the item `K` times more often in the mixture and at `1/K` of its mass when it does, and for a rare item the two cancel. Blending inheritance is therefore the null model of merging, and the proposition below states the cancellation exactly." Proposition lead: "In the inheritance model the dilution is a conservation law: the expected mass of a rare item in the child is `q·p` whatever the number of parents, so averaging over more parents neither helps nor harms a rare item's survival." 10. "Routing wins by exactly the amount averaging loses to dilution, and two things set that loss. A strong base on easy tasks has none: after averaging, the 7B model scores at ceiling (1.00 on two of three families), and routing has nothing to recover. Hard tasks restore it: at 7B the average falls to the level of the best single specialist (0.41), because each specialist's own skill is diluted, and routing among the intact specialists scores 0.50, ahead in every seed. A weak base (0.5B) shows the same gap on easy tasks. The variable is headroom, the distance between the average and the ceiling, and neither model size nor task difficulty alone." 11. "The autoencoder needed about 10% real data where the inheritance model needed 5%; the difference is what its sharpening bias costs, since a learner that concentrates mass on common modes loses rare ones faster than sampling alone would." 12. "...and removing it is the one ablation that fails outright: a population selected on agreement with its own consensus instead of on the verifier settles at 0.48 against 0.78 for the full society (Fig. 4D–F), confident and wrong." 13. "For continual learning the results give three things. The replay ratio has a formula: `m ≈ 1/p` examples per step of the rarest skill one refuses to lose, and 2m/(2m+1) of the diversity is kept. Replaying a network's own output (33, 91–93) is grounding with `g = 0` and collapses on the timescale of Fig. 2, one hop being too short to see it. And a pre-merge test (functional disagreement on shared probes) predicts interference where weight distance does not, with a control for shared training data that the regression (86) and distance (87, 88) studies lack." 14. "That no single model can answer one prompt two ways is a matter of information, not of training (SI Text S1, Proposition S2). What the population view adds is where the cliff sits: it moves with the share of shared prompts under contradictory conventions (Fig. 5B), and in a population that share grows whenever lineages adopt conventions independently." 15. "Under a curriculum that delivers every skill to every lineage the ceiling is what one adapter can hold (0.80 for a single model taught the whole syllabus), and sex and selection each reach it sooner without raising it." 16. "*Route or screen rather than average whenever the average falls short of the best parent on any task*: on hard tasks routing beat averaging by 0.09 in every seed and screened offspring by 0.07 (Fig. 3C), whereas on tasks the 7B base already answered at ceiling the plain average matched them and nothing was lost." / "Blending dilutes whichever parent's skill is rarest, so routing and offspring screening pay only where the plain average falls short of the best parent, and were inert where it did not (Fig. 3B–C)." 17. "...and this is the cost the obligate-merge arm of the six-generation population paid from generation 3 onward, when its partners began carrying opposite conventions for the same prompts (Fig. 4B)." 18. Table cell: "Fig. 3B–C: merging beats blending whenever the weight-average scores well below the best parent, and blending suffices when it does not". 19. Table cell: "The irreversibility is reproduced (Fig. S3); the mutational mechanism of the ratchet is not modelled, see (30)". 20. "...*grounded evaluation*: an agent is scored partly against reality and partly against the population's own consensus (`g`·true-fitness + (1−g)·conformity). The consensus term stands for what a population does when it has no verifier, which is to learn from its own outputs, so `g = 0` is a population that rewards agreement with itself." 21. "Drift is only the entry point, because population genetics is above all a theory of what keeps a population from decaying (immigration, recombination, selection, population structure) and of where each of those fails, and every one of them has a counterpart an operator of a model population can switch on: real data entering each generation, merging, verifier-anchored selection and the choice of which models merge with which." 22. "...transposed from a single network to a population whose members inherit from one another, and each of them has a population-genetic answer with a number attached (how many real samples, how far the average sits below the best parent, how much the parents disagree on shared inputs)." 23. "Fig. 1B draws the change of viewpoint the transfer rests on. Models are usually pictured as a society in space, contemporaries exchanging messages. The couplings that matter here (training on model output, merging, real data entering each generation) run between generations, and a society coupled in time is what population genetics describes." 24. "Compared against the exact model, the trained networks depart in two ways, both consequences of the estimator bias measured above. The threshold softens: ..." 26. "The sharpest test is whether isolation emerges with no conflicting signal anywhere, as a true Bateson–Dobzhansky–Muller incompatibility would: children were diverged..." 27. "The arm without grounding fails by construction, since a rule that scores agreement will converge on agreement; what the ablation adds is that the other two removals fail in different ways, so recombination and diversity are not substitutes for grounding or for each other. Magnitudes depend on the mutation, restart and selection schemes, which were not varied." 28. Keep one, in the Discussion: "Three of the framework's refinements failed (confidence weighting, the modifier reading of declines, selection turning speed into level), and the results are consistent with any account in which rare items are lost by sampling and conflicting conventions cannot share weights. The population-genetic reading earned its place by supplying the nulls and the overlap control, not by being the only mechanism left standing." 29. "In the inheritance model (Fig. 5E–F) hybrid fitness stays at the parents' level while the lineages remain compatible and then falls below the ancestor, sooner the more incompatibilities the genomes carry. Orr showed the number of such incompatibilities grows with the square of divergence (43). Whether a growing number of conflicts produces a fall in performance in a trained network is the question the simulation cannot answer." 30. "...the *Fisher–Muller effect* (37, 38). In an asexual population two useful variants that arise in different individuals can never meet in one descendant; the lineages carrying them compete, and one is lost. Recombination puts both into one offspring, which is why sexual populations adapt faster." 31. "...so what the parent learned in its lifetime passes to the child, the inheritance of acquired characters that Lamarck proposed and biology rejected, and that a weight file makes trivial." 32. "(a network is nonlinear in its weights, so averaging weights does not average outputs; but an update held by one of `N` parents is still scaled by `1/N` in the average (67), which is the dilution the proposition describes)". 33. "The resampling step is the Wright–Fisher process, population genetics' canonical model of neutral evolution, in which each generation is a random sample of size `n` from the last. In this *inheritance model* the Wright–Fisher population is the sample a child is trained on and its individuals are the `n + m` samples; it is a model of a learner." 34. "An item is the counterpart of an allele; a *capability* is what an item stands for." / "(allele frequency, in Table 1)" / cut or "(Methods)".