MachineSex/REPRODUCING.md
Giorgio Gilestro 6f8cef1ac5 main: keep only what reproduces the manuscript; everything else lives on dev
Removed from main (all preserved on the dev branch): the arXiv build and
its sources, design documents (blueprint, results summary, review responses,
essay drafts), tasks/ and CLAUDE.md, the cover letter and reference tooling,
two unused manuscript figures, and every experiment that feeds no figure or
number in the paper: the collapse null, the sexual-vs-asexual lineage, the
NK speciation variant, the 0.5B single-seed LLM prototypes, the compose and
society experiments with their calibration and pilot runs, and their
configs, runners, tests, figure scripts and PBS jobs. Their result bundles
are moved to results/_archive/ (ignored) so the parquets stay on disk.

Also: plot_llm_speciation reads the s{seed}/ layout; the mating-breadth
plot writes under its bundle name; Makefile targets reduced to the kept
experiments; REPRODUCING.md and README point to dev for the rest.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 17:07:23 +01:00

213 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Reproducing every number and figure in the paper
This document is the authoritative map from the manuscript back to the code, configs, and seeds
that produced it. Every figure panel, every headline number, and the environment they were computed
in are listed below. If something in the paper is not traceable through this document, that is a
bug — please open an issue.
Manuscript: `paper/manuscript/main.md` (built to `paper/manuscript/main.pdf`).
---
## 1. The three tiers, and what each costs to reproduce
| Tier | What it is | Hardware | Determinism |
|---|---|---|---|
| Inheritance model | WrightFisher simulator over knowledge distributions (pure NumPy/SciPy) | Any laptop, no GPU | **Bitwise** from the master seed |
| Trained networks | RNN / MLP / VAE on a synthetic mode universe; convolutional VAE on MNIST | One consumer GPU (16 GB) | Statistical (GPU non-determinism documented in §5) |
| Language models | LoRA specialists on Qwen2.5-Instruct 0.5B / 7B | 0.5B: one 16 GB GPU · 7B: one L40S (46 GB) | Statistical; per-seed points reported |
The inheritance-model tier carries every quantitative claim in the paper and reproduces exactly on a
laptop in minutes. The two AI tiers are confirmatory (signs, not magnitudes) and need a GPU.
## 2. Environment
The environment is a `uv` venv built from the committed, hash-pinned `uv.lock`. That lockfile — not
a container, not a requirements file — is the single source of truth for "it runs".
```bash
curl -LsSf https://astral.sh/uv/install.sh | sh # one-time, if you don't have uv
make env # biological-model tier (pure NumPy/SciPy)
make env-neural # + torch, for the trained-network tier
make env-mnist # + torchvision, for the real-MNIST tier
make env-llm # + transformers/peft, for the language-model tier
make env-notebooks # + Jupyter, for the walkthrough notebooks
```
## 3. One command
```bash
./reproduce.sh
```
runs the environment build, the test suite (including the closed-form scientific-validation tests),
the entire biological-model tier at its committed seeds, every figure, and writes `REPRODUCED.md`
with the resulting artifact hashes for comparison against the committed manifests. It deliberately
stops at the GPU tiers; pass `--with-gpu` to include them if you have the hardware.
Tier by tier, by hand:
```bash
make test # correctness + closed-form scientific validation
make inheritance # the inheritance model, every experiment at its committed seed
make neural # trained networks (needs a GPU)
make mnist # real-MNIST tier + the Fig. 2A montage asset (needs torchvision)
make llm # the 0.5B multi-seed language-model runs behind Fig. 3B (needs a GPU)
make llm-epistasis # the controlled predictive test behind Fig. 3D-E
make llm-speciation # the LLM speciation tier behind Fig. 5C-D
make figures # per-experiment figures, from committed parquets (no re-simulation)
make paper-figures # the manuscript figures Fig. 1-5 + rebuild the PDF body
```
## 4. The figure map
Every manuscript panel, the artifact it is plotted from, the config that produced that artifact, and
that config's declared seed. All panels are drawn by `paper/manuscript/make_figs.py` (function per figure);
`make paper-figures` regenerates all of them. Figures are a **pure function of committed artifacts**
no panel re-simulates anything, with the single documented exception of the Fig. 2A montage asset.
| Panel | Drawn by | From artifact | Produced by config | Seed(s) |
|---|---|---|---|---|
| Fig. 1A, 1B | `fig1a()`, `fig1b()` | — (schematics; icons in `paper/manuscript/figs/icons/`) | — | — |
| Fig. 2A | `fig2()` | `results/fig2_mnist_collapse/mnist_montage.png` | `configs/neural/fig2_mnist_collapse.yaml` → asset from `figures/mnist_montage.py` | 20260705 |
| Fig. 2B *(reference)* | `fig2()` | `results/fig2_grounding_sweep/` | `configs/inheritance/fig2_grounding_sweep.yaml` | 20260704 |
| Fig. 3A | `fig3()` | — (schematic of the compared models; icons in `paper/manuscript/figs/icons/`) | — | — |
| Fig. 3B | `fig3()` | `results/llm_merge_seeds/` | `configs/llm/merge_seeds.yaml` | 1, 2, 3, 4, 5 |
| Fig. 3C | `fig3()` | `results/llm_moe_hard_hpc/s{1,2,3}/` | `configs/llm/moe_hard_hpc.yaml` (7B, HPC; seeds 23 via `hpc/llm_7b_seeds.pbs`); statistics `figures/stats_llm_7b_seeds.py` | 1, 2, 3 |
| Fig. 3D, 3D | `fig3()` | `results/llm_epistasis/` + `results/llm_epistasis_compat/` | `configs/llm/epistasis.yaml`, `configs/llm/epistasis_compat.yaml` | 1, 2, 3 |
| Fig. 4A | `fig4()` | — (schematic; the syllabus grid is read from `configs/llm/curriculum_v5_s1.yaml` family order, complementarity from the veto bundle) | — | — |
| Fig. 4B | `fig4()` | `results/llm_curriculum_v5/`, `results/llm_curriculum_v5_veto/`, `results/llm_curriculum_v5_stop3/` | `configs/llm/curriculum_v5_s1.yaml` (seeds 23 via `hpc/llm_curriculum.pbs`), `configs/llm/curriculum_v5_veto.yaml` (seeds 23 via `hpc/llm_veto.pbs`), `configs/llm/curriculum_v5_stop3.yaml` (seeds 23 via `hpc/llm_curriculum_controls.pbs`) | 13 (training seeds; fixed test splits) |
| Fig. 4C | `fig4()` | `results/llm_curriculum_v5_veto/`, `results/llm_curriculum_v5_decor/` | `configs/llm/curriculum_v5_veto.yaml`, `configs/llm/curriculum_v5_decor.yaml` (seeds 23 via `hpc/llm_curriculum_controls.pbs`); statistics `figures/stats_llm_curriculum.py` | 13 |
| Fig. S14 | `figures/plot_curriculum_timing.py` | `results/llm_curriculum_v5_{early,late,early_obl,late_obl}/s{1,2,3}/` | `configs/llm/curriculum_v5_{early,late}[_obl].yaml` via `hpc/llm_curriculum_timing.pbs`; statistics `figures/stats_llm_curriculum.py` | 1, 2, 3 |
| Fig. S15 | `figures/plot_curriculum_cull.py` | `results/llm_curriculum_v5_cull/s{1,2,3}/` | `configs/llm/curriculum_v5_cull.yaml` via `hpc/llm_cull.pbs`; statistics `figures/stats_llm_curriculum.py` | 1, 2, 3 |
| Fig. S16 | `figures/plot_llm_smol.py` | `results/llm_{merge_seeds,moe_hard_seeds}_smol/` (+ the Qwen bundles) | `configs/llm/{merge_seeds,moe_hard_seeds}_smol.yaml` via `hpc/llm_smol.pbs`; statistics `figures/stats_llm_smol.py` | 15; 13 |
| Fig. 4D, 4D, 4E *(reference)* | `fig4()` | `results/fig4_society_ablation/` | `configs/inheritance/fig4_society_ablation.yaml` | 20260705 |
| Fig. 5A | `fig5()` | `results/speciation_real/` | `configs/neural/speciation_real.yaml` | 13 |
| Fig. 5B | `fig5()` | `results/speciation_real_cliff/` | `configs/neural/speciation_real_cliff.yaml` | 13 |
| Fig. 5C, 5D | `fig5()` | `results/llm_speciation/s{1,2,3}/` | `configs/llm/speciation.yaml` (seeds 23 via `hpc/llm_speciation_seeds.pbs`); statistics `figures/stats_llm_speciation_seeds.py` | 1, 2, 3 |
| Fig. 5E, 5F *(reference)* | `fig5()` | `results/fig5_speciation_bdm/` | `configs/inheritance/fig5_speciation_bdm.yaml` | 12 |
Panels marked *(reference)* are inheritance-model results included to set the expectation the
real-model panels are read against, not as findings in their own right.
**Inheritance-model results reported in SI only.** These have no real-model counterpart in the paper
and reproduce known results, so they are cited as reference values and plotted only in SI, from their
per-experiment figures: blending vs union (`figS8_multiparent_union`), FisherMuller super-parent
(`figS9_specialist_superparent`), outbreeding depression (`figS10_rugged_landscapes`), directed
recombination (`figS11_directed_recombination`), and mate-pool breadth (`figS13_mating_breadth`),
each drawn by `figures/plot_<bundle>.py` from `results/<bundle>/`.
**Bundle names.** Inheritance-model and trained-network bundles are named after the manuscript
figure they feed at submission (`fig2_*`, `figS4_*`); bundles that feed no figure carry a
descriptive name. The name is a label fixed at submission, so a later renumbering of the figures
does not rename the bundles.
**Replication.** Every panel is replicated: inheritance-model panels over 12100 internal
replicates; Fig. 3B over five training seeds; Fig. 3C, 3DE, 4AB and 5CD over three (Fig. 5CD
was single-seed until 2026-09-12; seeds 23 ran via `hpc/llm_speciation_seeds.pbs`).
### Results reported in the text but not plotted in the manuscript
| Result | Artifact | Config | Seed |
|---|---|---|---|
| Region-matched grounding (Fig. S5) | `results/figS5_aimed_grounding/` | `configs/inheritance/figS5_aimed_grounding.yaml` | 20260704 |
| Quality-diversity vs greedy (Fig. S12) | `results/figS12_quality_diversity/` | `configs/inheritance/figS12_quality_diversity.yaml` | 20260704 |
| Re-baselining / irreversibility (Fig. S3) | `results/figS3_rebaselining/` | `configs/inheritance/figS3_rebaselining.yaml` | 20260704 |
| Learning kernel (Fig. S2) | `results/figS2_kernel_sharpen/`, `results/figS2_kernel_smooth/` | `configs/inheritance/figS2_kernel_{sharpen,smooth}.yaml` | 20260705 |
| Histogram bridge gate | `results/bridge/` | `configs/neural/bridge.yaml` | 20260704 |
| Neural collapse / grounding / architectures / recombination | `results/{collapse,figS6_grounding_rnn,figS1_architectures,recombination}/` | `configs/neural/*.yaml` | 20260704 |
| Emergent-isolation null | `results/speciation_real_emergent/` | `configs/neural/speciation_real_emergent.yaml` | 813 |
| Budget-controlled speciation (add design) | `results/llm_speciation_add/` | `configs/llm/speciation_add.yaml` | 1, 2, 3 |
| 0.5B multi-seed runs | `results/llm_{merge_seeds,moe_hard_seeds,directed_hard_seeds}/` | `configs/llm/{merge_seeds,moe_hard_seeds,directed_hard_seeds}.yaml` | 15; 13; 13 |
| 7B firm-ups | `results/llm_*_hpc/` | `configs/llm/*_hpc.yaml` (run via `hpc/*.pbs`) | 1 |
### Per-experiment (exploratory) figures
`figures/plot_*.py` regenerate a diagnostic figure **inside each results bundle**
(`results/<name>/<name>.pdf`), named after the experiment, not after a manuscript figure. They are
the working views, not the manuscript's; the table above is the authority on what appears in the
paper. `figures/stats_llm_epistasis.py` prints the robust statistics quoted in the predictive-test
section (clustered bootstrap, paired contrasts, leave-one-condition-out, outcome-reference
sensitivity).
## 5. Seeds and determinism
**Policy.** One master seed per config. All sub-randomness is derived from it via
`numpy.random.SeedSequence.spawn` (`src/inheritance/seeding.py`); no code touches global RNG state, and
every `rng` is passed explicitly. A run is a pure function of its resolved config.
**Biological-model tier: bitwise reproducible.** Re-running a config on the same lockfile
reproduces its `results.parquet` byte-for-byte; the `results_sha256` in each `manifest.json` is the
check.
**GPU tiers: statistically reproducible.** cuDNN kernel selection and reduction order make bitwise
equality unattainable across machines. Per-seed points are reported rather than seed-averaged
summaries alone, and the multi-seed protocols fix the evaluation sets and vary only the training
seed. Expect sign agreement and magnitudes within noise, not identical digits.
**Seed provenance.** `20260704`/`20260705`/`20260709` are date-stamped master seeds chosen at the
time each experiment was written and never re-drawn. Small integer seeds (`1`, `12`, `13`, `813`)
are likewise fixed at authoring time. No seed in this repository was selected after seeing results.
## 6. Verifying artifact integrity
Every run writes three files next to its results:
- `results.parquet` — the long-form data (the only thing figures read)
- `resolved_config.yaml` — the config **after** sweep expansion, i.e. exactly what ran
- `manifest.json` — master seed, git commit, Python and library versions, row count, and
`results_sha256` (the content hash of the parquet)
To verify a bundle you have regenerated matches the one behind the paper:
```bash
python - <<'EOF'
import hashlib, json, pathlib
for m in sorted(pathlib.Path("results").glob("*/manifest.json")):
man = json.loads(m.read_text())
pq = m.parent / "results.parquet"
if not pq.exists():
print(f"{m.parent.name:28s} (no parquet — run its config first)"); continue
got = hashlib.sha256(pq.read_bytes()).hexdigest()
ok = "OK " if got == man.get("results_sha256") else "DIFF"
print(f"{ok} {m.parent.name:28s} seed={man.get('master_seed')}")
EOF
```
`DIFF` on a biological-model bundle means a genuine discrepancy worth investigating. `DIFF` on a GPU
tier is expected (see §5) — compare the figures and the reported statistics instead.
## 7. HPC (the 7B tier)
The 7B runs were executed on Imperial College's CX3 cluster (PBS Pro, one L40S 46 GB per job). Job
scripts are in `hpc/`; each is self-contained and documents its own submission line. They stage the
same `uv.lock` environment, so the only difference from a local run is the GPU.
```bash
qsub hpc/llm_merge.pbs # 7B merge firm-up
qsub hpc/llm_hard.pbs # hard-benchmark moe + directed at 7B
qsub hpc/llm_7b_seeds.pbs # seeds 2-3 of the three 7B experiments (array)
```
## 8. Notebooks
`notebooks/` contains executable walkthroughs (`make env-notebooks`, then `jupyter lab`):
| Notebook | What it does | Needs |
|---|---|---|
| `01_biological_model.ipynb` | Builds the WrightFisher model from scratch, checks it against the three closed forms, and derives the grounding threshold interactively | Laptop |
| `02_paper_figures.ipynb` | Regenerates every manuscript figure from the committed artifacts and displays them inline, panel by panel | Laptop (artifacts must be present) |
`make notebooks` executes both end-to-end, which is itself a reproduction check.
## 9. Known gaps
- `results/**/results.parquet` is currently **gitignored** (only manifests, hashes, and resolved
configs are tracked). A fresh clone therefore has to re-run the experiments before figures can be
regenerated. The archived deposit (Zenodo DOI, on publication) includes the parquets so that the
paper's "regenerates from committed artifacts without re-simulation" holds from the archive.
- `figures/mnist_montage.py` re-runs a short dry lineage to draw its montage rather than reading a
parquet; it is an eyeball diagnostic whose quantitative counterpart is `results/fig2_mnist_collapse/`.
- The composed society at language-model scale is an open experiment at the time of writing; its
configs, pilots and pre-registration are on the `dev` branch.