Code and data associated with "The evolution of sex for artificial intelligence - A population-genetic framework for multigenerational model populations". Gilestro, 2026
Find a file
Giorgio Gilestro 6f8cef1ac5 main: keep only what reproduces the manuscript; everything else lives on dev
Removed from main (all preserved on the dev branch): the arXiv build and
its sources, design documents (blueprint, results summary, review responses,
essay drafts), tasks/ and CLAUDE.md, the cover letter and reference tooling,
two unused manuscript figures, and every experiment that feeds no figure or
number in the paper: the collapse null, the sexual-vs-asexual lineage, the
NK speciation variant, the 0.5B single-seed LLM prototypes, the compose and
society experiments with their calibration and pilot runs, and their
configs, runners, tests, figure scripts and PBS jobs. Their result bundles
are moved to results/_archive/ (ignored) so the parquets stay on disk.

Also: plot_llm_speciation reads the s{seed}/ layout; the mating-breadth
plot writes under its bundle name; Makefile targets reduced to the kept
experiments; REPRODUCING.md and README point to dev for the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 17:07:23 +01:00
configs main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
figures main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
hpc main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
notebooks Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
paper/manuscript main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
results main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
src main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
tests main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
.gitignore main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
.python-version Reproducibility pass: figure map, one-command reproduce.sh, notebooks, Makefile gaps 2026-09-07 15:50:50 +01:00
Makefile main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
pyproject.toml Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
README.md Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
reproduce.sh Restructure: descriptive tier and experiment names, paper/manuscript 2026-09-13 17:00:40 +01:00
REPRODUCING.md main: keep only what reproduces the manuscript; everything else lives on dev 2026-09-13 17:07:23 +01:00
uv.lock Reproducibility pass: figure map, one-command reproduce.sh, notebooks, Makefile gaps 2026-09-07 15:50:50 +01:00

The evolution of sex for artificial intelligence

A population-genetic framework for multigenerational model populations. Knowledge transmission between generations of learning agents is modelled literally as a WrightFisher process, not by analogy: a model's knowledge is a distribution p_t over K discrete items, a fixed true distribution p* has a rare tail, and each generational step is "sample from the parent (drift) + mix in fresh real samples (grounding/immigration) + refit". Model collapse is the loss of rare alleles under drift — and the remedies population genetics knows for drift (immigration, recombination, selection, population structure) become engineering levers for model populations.

The framework is developed at three tiers of increasing realism:

Tier What it is Hardware
Inheritance model WrightFisher simulator over knowledge distributions; closed forms, bitwise reproducible laptop
Trained networks RNN / MLP / VAE on a synthetic mode universe with an exact oracle; convolutional VAE on MNIST one GPU
Language models LoRA specialists on Qwen2.5-Instruct (0.5B / 7B) with an exact-match verifier one GPU / L40S

Reproduce

Start here: REPRODUCING.md — the authoritative map from every manuscript figure panel back to the artifact, config, and seed that produced it, plus the determinism policy and artifact-hash verification.

curl -LsSf https://astral.sh/uv/install.sh | sh   # one-time, if needed

./reproduce.sh          # env -> tests -> inheritance-model tier at committed seeds -> figures
./reproduce.sh --with-gpu   # ... and the trained-network + language-model tiers

Or tier by tier:

make env             # build .venv from the committed, hash-pinned uv.lock
make test            # correctness + closed-form scientific validation (the spine of trust)
make inheritance     # the inheritance model, every experiment at its committed seed
make figures         # per-experiment figures, from committed parquets (no re-simulation)
make paper-figures   # the manuscript's Fig. 1-5 + rebuild the PDF bodies

make help is not defined, but every target carries a ## description — grep '##' Makefile.

Notebooks

make env-notebooks && jupyter lab notebooks/
  • 01_biological_model.ipynb — builds the WrightFisher model from scratch, checks it against the closed forms (geometric diversity decay, the immigrationdrift equilibrium), and derives the grounding threshold and its per-item observation floor. Runs on a laptop in under a minute.
  • 02_paper_figures.ipynb — verifies artifact hashes, then regenerates and displays every manuscript figure from the committed artifacts.

Layout

src/inheritance/ inheritance-model tier (imported as `inheritance`)
src/neural/      trained-network tier
src/llm/         language-model tier
configs/         one YAML per experiment: inheritance/ neural/ llm/ (each declares its master seed)
figures/         plot_*.py — per-experiment diagnostics, read results.parquet only
paper/manuscript/ the manuscript: main.md, si.md, make_figs.py (Fig. 1-5), si_figures.py, build.py
notebooks/       executable walkthroughs
hpc/             PBS job scripts for the 7B tier (Imperial CX3)
tests/           correctness + test_scientific_validation.py (the closed forms as assertions)
results/         run artifacts: results.parquet (gitignored) + resolved_config.yaml + manifest.json;
                 bundles are named after the manuscript figure they feed (fig2_*, figS4_*)

Development history, design documents, pre-registrations and exploratory experiments that did not reach the manuscript live on the dev branch; main holds only what reproduces the paper.

The engineering contract

  • Reproducibility is a requirement, not a preference. The environment is a uv venv built from a committed, hash-pinned uv.lock; the biological-model tier is bitwise reproducible from a single master seed, and the GPU tiers are statistically reproducible with per-seed points reported.
  • One master seed per config, with all sub-randomness derived via SeedSequence.spawn. No code touches global RNG state; a run is a pure function of its resolved config.
  • No magic numbers in code. Every parameter lives in a YAML resolved at run time, and the resolved config is written next to the results.
  • Every run writes the same triple: results.parquet + resolved_config.yaml + manifest.json (seed, git commit, library versions, row count, SHA-256 of the results).
  • Every figure is a pure function of a committed artifact — figure scripts never re-simulate.
  • The scientific-validation tests are the spine of trust. They assert that the simulator reproduces the closed forms to within 0.5%. If they fail, the science is wrong, not just the code.