Code and data associated with "The evolution of sex for artificial intelligence - A population-genetic framework for multigenerational model populations". Gilestro, 2026
Find a file
Giorgio Gilestro 585264d0b4 llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive
Re-ran the specialist-merge experiment at a capable base (Qwen2.5-7B-Instruct,
200 tests/family) on one L40S GPU of Imperial's CX3 HPC (8 min walltime). The
two caveats the 0.5B prototype left marginal are now resolved:

- "exceeds every parent overall" is clean: both merges 0.87 vs best specialist
  0.77 (+10 pts), and above every specialist on every family.
- dilution vanishes: at 0.5B averaging diluted the lists-specialist
  (0.43->0.26); at 7B the merge beats it (0.62>0.57). Dilution was a
  small-model artefact -- a capable base composes rather than dilutes, which
  softens E4's "merge, don't average" once the parents are strong.

The figure title is now data-driven (reports ">" for 7B, "~" for 0.5B).
Adds the hpc/ smoke job script and the llm_merge walltime trim. Results synced
to results/llm_merge_hpc/ (parquet gitignored per the reproducibility contract).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 17:32:51 +01:00
configs hpc: PBS job scripts for Imperial CX3 (probe + scaled LLM merge run) 2026-07-05 16:02:09 +01:00
figures llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive 2026-07-05 17:32:51 +01:00
hpc llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive 2026-07-05 17:32:51 +01:00
paper paper: reframe the perspective paper around sexual reproduction (v4 -> v5) 2026-07-05 13:53:10 +01:00
results llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive 2026-07-05 17:32:51 +01:00
src llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00
tasks llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive 2026-07-05 17:32:51 +01:00
tests llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00
.gitignore neural: real-MNIST external-validity tier (collapse + grounding) 2026-07-05 09:19:36 +01:00
CLAUDE.md llm_merge_hpc: the 7B firm-up makes the Fisher-Muller sign decisive 2026-07-05 17:32:51 +01:00
Makefile llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00
pyproject.toml llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00
README.md paper: reframe the perspective paper around sexual reproduction (v4 -> v5) 2026-07-05 13:53:10 +01:00
uv.lock Layer 1.5: architecture-general neural existence proof 2026-07-04 21:02:49 +01:00

The Lamarckian Society — Layer 1 (analytical core)

A parametric population-genetics model of knowledge transmission across generations of learning agents. Knowledge transmission is modelled literally as a WrightFisher process (not by analogy): a model's knowledge is a distribution p_t over K discrete items; a fixed true distribution p* has a rare tail; each generational step is "sample from the parent (drift) + mix in fresh real samples (grounding/immigration) + refit." Model collapse is the loss of rare alleles under drift.

See paper/blueprint.md (the normative build spec), paper/the-lamarckian-society-v5.md (the perspective paper), and paper/results-summary.md (a summary of all results).

Reproduce

Environment is a uv venv built from the committed, hash-pinned uv.lock — that lockfile is the single source of truth for "it runs" (Layer 1 is pure NumPy/SciPy and bitwise-reproducible from a seed; no container needed).

# one-time: install uv (https://astral.sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh

uv sync                 # build .venv from uv.lock
make test               # correctness + scientific-validation tests (the spine of trust)
make layer1             # run experiments E1E6
make figures            # regenerate figures from committed results

Layout

src/knowledge/   Layer 1 package (imported as `knowledge`)
configs/layer1/  one YAML per experiment (E1..E6)
figures/         plot_EX.py — read results.parquet only
tests/           test_correctness.py + test_scientific_validation.py (analytic checks)
paper/           blueprint.md, perspective paper, figure_manifest.md
results/         written artifacts (gitignored; hashes tracked in manifest.json)