Code and data associated with "The evolution of sex for artificial intelligence - A population-genetic framework for multigenerational model populations". Gilestro, 2026
Find a file
Giorgio Gilestro 8da0dac007 llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale
Adds the union-preserving recombination operator that llm_merge lacked (E8's max,
not mean): keep each specialist LoRA intact and SELECT the right one per prompt
(MoE router: oracle, or training-free nearest-centroid over base embeddings) or
per module (max_merge = winner-take-all by delta norm). src/llm/moe.py, kind
llm_moe, reuses the cached specialists.

Result — a clean regime boundary for "merge, don't average":
- 0.5B: union wins. Routing 0.74 / worst-family 0.43 > soup 0.64 / 0.26, with no
  dilution (recovers each specialist's own-family peak). E8's max > mean in real
  weights, because at a weak base averaging dilutes.
- 7B (Imperial CX3, L40S, 9 min): the ordering INVERTS. Fusion wins — soup 0.87 >
  routing 0.84 > max_merge 0.78. Routing is capped at the best parent per family;
  fusion blends and, given a capable base, COMPOSES beyond any parent (soup lists
  0.62 > spec 0.57). Selection can't synthesise better than its best component;
  averaging-that-composes can.

So "merge, don't average" (E4/E8) is a weak-parent / small-model law, not
universal: union wins under dilution, fusion wins under composition. Refines E8
(its additive-landscape max>mean assumed no compositional headroom). The operator
to want is fusion-that-composes + offspring selection = the directed-sex ideal
(E10) — the natural next experiment.

Honest riders: the learned router is trivially perfect (lexically-distinct
families), and router-free max_merge is the weakest union (not input-adaptive).
+2 router unit tests (127 green). Results in results/llm_moe{,_hpc}/ (parquet
gitignored per the reproducibility contract).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 17:53:47 +01:00
configs llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
figures llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
hpc llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
paper paper: reframe the perspective paper around sexual reproduction (v4 -> v5) 2026-07-05 13:53:10 +01:00
results llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
src llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
tasks llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
tests llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
.gitignore neural: real-MNIST external-validity tier (collapse + grounding) 2026-07-05 09:19:36 +01:00
CLAUDE.md llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
Makefile llm_moe: the union operator (route/max-merge) vs fusion — and the regime flips at scale 2026-07-05 17:53:47 +01:00
pyproject.toml llm: first real-LLM prototype — recombining specialist LLMs (C2/C4) 2026-07-05 15:48:02 +01:00
README.md paper: reframe the perspective paper around sexual reproduction (v4 -> v5) 2026-07-05 13:53:10 +01:00
uv.lock Layer 1.5: architecture-general neural existence proof 2026-07-04 21:02:49 +01:00

The Lamarckian Society — Layer 1 (analytical core)

A parametric population-genetics model of knowledge transmission across generations of learning agents. Knowledge transmission is modelled literally as a WrightFisher process (not by analogy): a model's knowledge is a distribution p_t over K discrete items; a fixed true distribution p* has a rare tail; each generational step is "sample from the parent (drift) + mix in fresh real samples (grounding/immigration) + refit." Model collapse is the loss of rare alleles under drift.

See paper/blueprint.md (the normative build spec), paper/the-lamarckian-society-v5.md (the perspective paper), and paper/results-summary.md (a summary of all results).

Reproduce

Environment is a uv venv built from the committed, hash-pinned uv.lock — that lockfile is the single source of truth for "it runs" (Layer 1 is pure NumPy/SciPy and bitwise-reproducible from a seed; no container needed).

# one-time: install uv (https://astral.sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh

uv sync                 # build .venv from uv.lock
make test               # correctness + scientific-validation tests (the spine of trust)
make layer1             # run experiments E1E6
make figures            # regenerate figures from committed results

Layout

src/knowledge/   Layer 1 package (imported as `knowledge`)
configs/layer1/  one YAML per experiment (E1..E6)
figures/         plot_EX.py — read results.parquet only
tests/           test_correctness.py + test_scientific_validation.py (analytic checks)
paper/           blueprint.md, perspective paper, figure_manifest.md
results/         written artifacts (gitignored; hashes tracked in manifest.json)