The decisive experiment from the external review. 39 LoRA parent pairs (0.5B, 3 seeds) on three axes decorrelated by construction: conflict (contradictory conventions on shared prompts, private budgets fixed), compat (same prompts, SAME convention — overlap without conflict), and duration (weight divergence, zero conflict). Six pre-merge predictors; primary outcome = merge penalty (parent potential − merged achieved). League table (Spearman vs penalty, n=39): functional measures predict (dis_raw +0.460, epi_conf +0.446, p<0.005); geometry collapses (delta_cos +0.03, delta_l2 +0.17 n.s.); gradient alignment weak (−0.35); performance ~0. The first grid's apparent geometry win (+0.60) was an overlap/volume artifact — the compat control axis (added for exactly this) exposed and killed it: same overlap and data volume, zero penalty. Honest riders in the README: confidence weighting does not beat raw disagreement as a rank predictor (pre-registered internal prediction not confirmed; it does double the conflict/compat level contrast), and |rho|~0.45 is bounded by 0.5B merge-outcome noise (7B is the firm-up). Also: micro-batched gradient accumulation (OOM fix on the shared 16GB GPU), exact r-space LoRA-delta geometry (brute-force-verified test, 151 green), systemd-run runbook lesson (tmux dies with the SSH session scope on this box). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v |
||
|---|---|---|
| configs | ||
| figures | ||
| hpc | ||
| paper | ||
| results | ||
| src | ||
| tasks | ||
| tests | ||
| .gitignore | ||
| CLAUDE.md | ||
| Makefile | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
The Lamarckian Society — Layer 1 (analytical core)
A parametric population-genetics model of knowledge transmission across generations of
learning agents. Knowledge transmission is modelled literally as a Wright–Fisher
process (not by analogy): a model's knowledge is a distribution p_t over K discrete
items; a fixed true distribution p* has a rare tail; each generational step is
"sample from the parent (drift) + mix in fresh real samples (grounding/immigration) +
refit." Model collapse is the loss of rare alleles under drift.
See paper/blueprint.md (the normative build spec),
paper/the-lamarckian-society-v5.md (the perspective paper), and
paper/results-summary.md (a summary of all results).
Reproduce
Environment is a uv venv built from the committed, hash-pinned uv.lock — that
lockfile is the single source of truth for "it runs" (Layer 1 is pure NumPy/SciPy and
bitwise-reproducible from a seed; no container needed).
# one-time: install uv (https://astral.sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync # build .venv from uv.lock
make test # correctness + scientific-validation tests (the spine of trust)
make layer1 # run experiments E1–E6
make figures # regenerate figures from committed results
Layout
src/knowledge/ Layer 1 package (imported as `knowledge`)
configs/layer1/ one YAML per experiment (E1..E6)
figures/ plot_EX.py — read results.parquet only
tests/ test_correctness.py + test_scientific_validation.py (analytic checks)
paper/ blueprint.md, perspective paper, figure_manifest.md
results/ written artifacts (gitignored; hashes tracked in manifest.json)