diff --git a/paper/pnas/body.tex b/paper/pnas/body.tex index 9540d5f..b9762ef 100644 --- a/paper/pnas/body.tex +++ b/paper/pnas/body.tex @@ -20,7 +20,7 @@ An operator of a model population faces recurring decisions for which there is n \begin{figure*}[p]\centering % fig1 \includegraphics[width=\textwidth]{figs/fig1.pdf} -\caption{The experimental programme. Each population-genetic abstraction (Table 1) is tested at up to three tiers, ordered left to right by increasing realism: an exact Wright--Fisher simulator over knowledge distributions (closed forms; bitwise-reproducible), trained neural networks measured against exact oracles (recurrent, feedforward, and variational-autoencoder generators on a synthetic mode universe, and a convolutional VAE on MNIST), and language models (LoRA specialists on Qwen bases at 0.5B and 7B, scored by an exact-match verifier). Colour separates the two categories: the population-genetic theory tier in blue, the two AI-model tiers in oranges. The same abstractions are carried across all three. Rows are the framework's mechanisms, each defined at the left margin; filled cells name the experiments run at each tier. Each claim is tested at the cheapest tier that can falsify it, and a costlier tier is entered only where it adds a discriminating test rather than a replication: grounding at language-model scale is established in prior work (21, 30) and is not re-run; epistasis and the society skip the middle tier, whose distinctive value (exact oracles) does not bear on those operator-level questions; and the society at language-model scale is the integrative experiment this paper specifies but does not run --- its stated gap.}\label{fig1} +\caption{The experimental programme. Each population-genetic abstraction (Table 1) is tested at up to three tiers, ordered left to right by increasing realism: an exact Wright--Fisher simulator over knowledge distributions (closed forms; bitwise-reproducible), trained neural networks measured against exact oracles (recurrent, feedforward, and variational-autoencoder generators on a synthetic mode universe, and a convolutional VAE on MNIST), and language models (LoRA specialists on Qwen bases at 0.5B and 7B, scored by an exact-match verifier). Colour separates the two categories: the population-genetic theory tier in blue, the two AI-model tiers in oranges. The same abstractions are carried across all three. Rows are the framework's mechanisms, each defined at the left margin; filled cells name the experiments run at each tier, and each carries, in its corner, the figure or table where that result is reported, so this figure doubles as a map of the paper. Each claim is tested at the cheapest tier that can falsify it, and a costlier tier is entered only where it adds a discriminating test rather than a replication: grounding at language-model scale is established in prior work (21, 30) and is not re-run; epistasis and the society skip the middle tier, whose distinctive value (exact oracles) does not bear on those operator-level questions; and the society at language-model scale is the integrative experiment this paper specifies but does not run --- its stated gap.}\label{fig1} \end{figure*} \section*{The minimal model, and where its exactness ends} @@ -63,7 +63,7 @@ The largest returns from the transfer concern merging. \emph{Blending inheritanc The generative payoff is the \emph{Fisher--Muller effect} (35, 36), the classical account of why sex speeds adaptation: in an asexual population, beneficial variants arising in different individuals can only compete until all but one lineage is lost, whereas recombination assembles them in one offspring, producing a \emph{genotype} (an individual's combination of variants, one at each \emph{locus}, or position) fitter than any parent. In the multi-locus model, sexual merging of decorrelated specialists climbs to the global optimum, a genotype no parent held, while the best single parent and the blended average both plateau below (Fig. 4). In real language models the signature replicates under seed replication: merges of three LoRA (37) specialists beat every parent overall (decisively at 7B: 0.87 vs 0.77), and on the sharper worst-family metric the merged models are the only ones competent everywhere, in every seed (Fig. 3A). -Sex has risks and, for AI, an unfair advantage, both quantified on Kauffman's NK fitness landscapes (38), the standard model of \emph{epistasis}, biology's term for interaction between genes: the fitness contribution of a variant depends on which variants occupy the other loci, much as a component's value in an ML system depends on the components around it. Each of the landscape's \texttt{N} sites interacts with \texttt{K} others (the model's eponymous parameters), and raising that interaction count tunes the landscape from smooth and additive to rugged and many-peaked (Fig. 5). When skills are entangled, blind recombination produces offspring \emph{below} their parents, worsening with ruggedness, and the optimal recombination rate shrinks as entanglement grows. Biology knows this failure as \emph{outbreeding depression}, the reason conservation practice warns against crossing locally adapted populations: in the textbook case, an ibex herd in the Tatra Mountains restocked with animals from Turkey and Sinai produced fertile hybrids that bore their young in the coldest month of winter, and the herd died out (39). But an engineered population can do what biology cannot: recombine unbounded parents, choose complementary mates, and \emph{screen many candidate offspring against a verifier before keeping one}. This directed sex converts the outbreeding catastrophe into a reliable gain in the model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks, including one seed where the blend failed catastrophically and selection was immune (Fig. 3A). Finally, population \emph{structure} is itself a knob: sweeping the mate-pool breadth from monogamous (repeated local pairings) to promiscuous (\emph{panmixia}: any model may merge with any other) against ruggedness, wide mixing maximises the population mean while monotonically destroying diversity, and the best \emph{champion} shifts from wide breadth on smooth landscapes to intermediate breadth on rugged ones (Fig. 5C), the mating-system phenomenon known to structured-population search, mapped onto merging populations. +Sex has risks and, for AI, an unfair advantage, both quantified on Kauffman's NK fitness landscapes (38), the standard model of \emph{epistasis}, biology's term for interaction between genes: the fitness contribution of a variant depends on which variants occupy the other loci, much as a component's value in an ML system depends on the components around it. Each of the landscape's \texttt{N} sites interacts with \texttt{K} others (the model's eponymous parameters), and raising that interaction count tunes the landscape from smooth and additive to rugged and many-peaked (Fig. 5). When skills are entangled, blind recombination produces offspring \emph{below} their parents, worsening with ruggedness, and the optimal recombination rate shrinks as entanglement grows. Biology knows this failure as \emph{outbreeding depression}, the reason conservation practice warns against crossing locally adapted populations: in the textbook case, an ibex herd in the Tatra Mountains restocked with animals from Turkey and Sinai produced fertile hybrids that bore their young in the coldest month of winter, and the herd died out (39). But an engineered population can do what biology cannot: recombine unbounded parents, choose complementary mates, and \emph{screen many candidate offspring against a verifier before keeping one}. This directed sex converts the outbreeding catastrophe into a reliable gain in the model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks, including one seed where the blend failed catastrophically and selection was immune (SI Appendix, Table S2). Finally, population \emph{structure} is itself a knob: sweeping the mate-pool breadth from monogamous (repeated local pairings) to promiscuous (\emph{panmixia}: any model may merge with any other) against ruggedness, wide mixing maximises the population mean while monotonically destroying diversity, and the best \emph{champion} shifts from wide breadth on smooth landscapes to intermediate breadth on rugged ones (Fig. 5C), the mating-system phenomenon known to structured-population search, mapped onto merging populations. \begin{figure*}[p]\centering % fig3 \includegraphics[width=\textwidth]{figs/fig3.pdf} diff --git a/paper/pnas/build.py b/paper/pnas/build.py index 18b1b5a..315ae1a 100644 --- a/paper/pnas/build.py +++ b/paper/pnas/build.py @@ -31,7 +31,8 @@ FIGURES: dict[str, tuple[list[str], str]] = { "exact-match verifier). Colour separates the two categories: the population-genetic theory " "tier in blue, the two AI-model tiers in oranges. The same abstractions are carried across " "all three. Rows are the framework's mechanisms, each defined at the left margin; filled " - "cells name the experiments run at each tier. Each claim is tested at the cheapest tier that " + "cells name the experiments run at each tier, and each carries, in its corner, the figure " + "or table where that result is reported, so this figure doubles as a map of the paper. Each claim is tested at the cheapest tier that " "can falsify it, and a costlier tier is entered only where it adds a discriminating test " "rather than a replication: grounding at language-model scale is established in prior work " "(21, 30) and is not re-run; epistasis and the society skip the middle tier, whose " diff --git a/paper/pnas/figs/fig1.pdf b/paper/pnas/figs/fig1.pdf index 21d8914..403c9ca 100644 Binary files a/paper/pnas/figs/fig1.pdf and b/paper/pnas/figs/fig1.pdf differ diff --git a/paper/pnas/main.md b/paper/pnas/main.md index 4e5e21d..f11769a 100644 --- a/paper/pnas/main.md +++ b/paper/pnas/main.md @@ -237,7 +237,8 @@ parents, choose complementary mates, and *screen many candidate offspring agains keeping one*. This directed sex converts the outbreeding catastrophe into a reliable gain in the model (tracking or exceeding the best parent at every ruggedness) and replicates as a sign in language models: bred-and-screened merges beat the a-priori blend in every seed on headroom tasks, including -one seed where the blend failed catastrophically and selection was immune (Fig. 3A). Finally, +one seed where the blend failed catastrophically and selection was immune (SI Appendix, +Table S2). Finally, population *structure* is itself a knob: sweeping the mate-pool breadth from monogamous (repeated local pairings) to promiscuous (*panmixia*: any model may merge with any other) against ruggedness, wide mixing maximises the population mean while diff --git a/paper/pnas/main.pdf b/paper/pnas/main.pdf index 91834f8..73502f0 100644 Binary files a/paper/pnas/main.pdf and b/paper/pnas/main.pdf differ diff --git a/paper/pnas/make_figs.py b/paper/pnas/make_figs.py index b6f9cd0..f7df53a 100644 --- a/paper/pnas/make_figs.py +++ b/paper/pnas/make_figs.py @@ -82,6 +82,12 @@ def fig1(): "convention conflict $\\to$ hybrid\nbreakdown; duration null; pre-merge\npredictive test (13 cond. $\\times$ 3 seeds)"]), ] + TAGS = [("Fig. 2A", "Fig. 2B", None), + ("Fig. 4", "Table S2", "Fig. 3A\u2013B"), + ("Fig. 5", None, "Table S2"), + ("Fig. 6", None, None), + ("Fig. 7A\u2013B", "Fig. 7C\u2013D", "Figs. 7E\u2013F, 3C\u2013D")] + fig, ax = plt.subplots(figsize=(11.4, 5.3)) ax.set_axis_off() ax.set_xlim(0, 1) @@ -103,6 +109,7 @@ def fig1(): style="italic", color="white", alpha=0.92) for i2, (label, definition, cells) in enumerate(ROWS): + tags = TAGS[i2] y1 = row_top - i2 * row_h y0 = y1 - row_h + 2 * gap yc = (y0 + y1) / 2 @@ -133,7 +140,11 @@ def fig1(): else: ax.add_patch(FancyBboxPatch((x + gap, y0), cw - 2 * gap, y1 - y0, boxstyle="round,pad=0.004", fc=face, ec=edge, lw=0.9)) - ax.text(x + cw / 2, yc, cell, ha="center", va="center", fontsize=6.4, linespacing=1.35) + ax.text(x + cw / 2, yc + 0.008, cell, ha="center", va="center", + fontsize=6.4, linespacing=1.35) + if tags[j2]: # where the result lives (the ToC role) + ax.text(x + cw - gap - 0.005, y0 + 0.006, tags[j2], ha="right", va="bottom", + fontsize=5.6, style="italic", color=edge) save(fig, "fig1") # ---------------------------------------------------------------- fig 2: grounding + MNIST