Per GG's directive: (1) the model-societies premise is no longer asserted — the Introduction opens with the verified evidence base (3M-model ecosystem with phylogenetic lineage-mapping literature, >98%-synthetic alignment pipelines, machine-generated web share, the human-data ceiling, mainstream merging tooling, agent economies; refs 31-44, all identifiers verified by the literature scan). (2) The findings are contextualised in CONTINUAL LEARNING, where they land hardest: a new Introduction block maps the CL canon onto the operators — replay <-> grounding, with the field's measured replay fractions (1%/5%/25%) sitting on our theorized g*~0.05; pseudo-rehearsal/generative replay as precisely our ungrounded null; parameter isolation; CLS consolidation; merging-for-CL vs cross-lineage recombination; tail-first forgetting <-> tail-allele extinction; CF-vs-collapse mechanism distinction kept explicit — plus a Discussion block with five CL impact points (replay- ratio theory testable against published sweeps; a failure theory for generative replay; pre-merge interference prediction with a mechanism; a consolidate-vs-modular decision rule; tail monitoring, engaging the latent-vs-extinct objection). The scan verified the bridge is open: no prior work carries pop-gen formalism into CL. (3) Downplaying replaced by convergence framing: the diagnosis was reached independently and is corroborated by parallel arrivals (Riis; Benati; Yoon; and Crutchfield & Whalen 2012, pre-deep-learning) — cited for priority of publication, the full arc owned as one framework. References 30 -> 65; Significance carries the CL frame; 20-pp rebuild; 151 tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v |
||
|---|---|---|
| configs | ||
| figures | ||
| hpc | ||
| paper | ||
| results | ||
| src | ||
| tasks | ||
| tests | ||
| .gitignore | ||
| CLAUDE.md | ||
| Makefile | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
The Lamarckian Society — Layer 1 (analytical core)
A parametric population-genetics model of knowledge transmission across generations of
learning agents. Knowledge transmission is modelled literally as a Wright–Fisher
process (not by analogy): a model's knowledge is a distribution p_t over K discrete
items; a fixed true distribution p* has a rare tail; each generational step is
"sample from the parent (drift) + mix in fresh real samples (grounding/immigration) +
refit." Model collapse is the loss of rare alleles under drift.
See paper/blueprint.md (the normative build spec),
paper/the-lamarckian-society-v5.md (the perspective paper), and
paper/results-summary.md (a summary of all results).
Reproduce
Environment is a uv venv built from the committed, hash-pinned uv.lock — that
lockfile is the single source of truth for "it runs" (Layer 1 is pure NumPy/SciPy and
bitwise-reproducible from a seed; no container needed).
# one-time: install uv (https://astral.sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync # build .venv from uv.lock
make test # correctness + scientific-validation tests (the spine of trust)
make layer1 # run experiments E1–E6
make figures # regenerate figures from committed results
Layout
src/knowledge/ Layer 1 package (imported as `knowledge`)
configs/layer1/ one YAML per experiment (E1..E6)
figures/ plot_EX.py — read results.parquet only
tests/ test_correctness.py + test_scientific_validation.py (analytic checks)
paper/ blueprint.md, perspective paper, figure_manifest.md
results/ written artifacts (gitignored; hashes tracked in manifest.json)