Adds the union-preserving recombination operator that llm_merge lacked (E8's max,
not mean): keep each specialist LoRA intact and SELECT the right one per prompt
(MoE router: oracle, or training-free nearest-centroid over base embeddings) or
per module (max_merge = winner-take-all by delta norm). src/llm/moe.py, kind
llm_moe, reuses the cached specialists.
Result — a clean regime boundary for "merge, don't average":
- 0.5B: union wins. Routing 0.74 / worst-family 0.43 > soup 0.64 / 0.26, with no
dilution (recovers each specialist's own-family peak). E8's max > mean in real
weights, because at a weak base averaging dilutes.
- 7B (Imperial CX3, L40S, 9 min): the ordering INVERTS. Fusion wins — soup 0.87 >
routing 0.84 > max_merge 0.78. Routing is capped at the best parent per family;
fusion blends and, given a capable base, COMPOSES beyond any parent (soup lists
0.62 > spec 0.57). Selection can't synthesise better than its best component;
averaging-that-composes can.
So "merge, don't average" (E4/E8) is a weak-parent / small-model law, not
universal: union wins under dilution, fusion wins under composition. Refines E8
(its additive-landscape max>mean assumed no compositional headroom). The operator
to want is fusion-that-composes + offspring selection = the directed-sex ideal
(E10) — the natural next experiment.
Honest riders: the learned router is trivially perfect (lexically-distinct
families), and router-free max_merge is the weakest union (not input-adaptive).
+2 router unit tests (127 green). Results in results/llm_moe{,_hpc}/ (parquet
gitignored per the reproducibility contract).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|---|---|---|
| configs | ||
| figures | ||
| hpc | ||
| paper | ||
| results | ||
| src | ||
| tasks | ||
| tests | ||
| .gitignore | ||
| CLAUDE.md | ||
| Makefile | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
The Lamarckian Society — Layer 1 (analytical core)
A parametric population-genetics model of knowledge transmission across generations of
learning agents. Knowledge transmission is modelled literally as a Wright–Fisher
process (not by analogy): a model's knowledge is a distribution p_t over K discrete
items; a fixed true distribution p* has a rare tail; each generational step is
"sample from the parent (drift) + mix in fresh real samples (grounding/immigration) +
refit." Model collapse is the loss of rare alleles under drift.
See paper/blueprint.md (the normative build spec),
paper/the-lamarckian-society-v5.md (the perspective paper), and
paper/results-summary.md (a summary of all results).
Reproduce
Environment is a uv venv built from the committed, hash-pinned uv.lock — that
lockfile is the single source of truth for "it runs" (Layer 1 is pure NumPy/SciPy and
bitwise-reproducible from a seed; no container needed).
# one-time: install uv (https://astral.sh/uv)
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync # build .venv from uv.lock
make test # correctness + scientific-validation tests (the spine of trust)
make layer1 # run experiments E1–E6
make figures # regenerate figures from committed results
Layout
src/knowledge/ Layer 1 package (imported as `knowledge`)
configs/layer1/ one YAML per experiment (E1..E6)
figures/ plot_EX.py — read results.parquet only
tests/ test_correctness.py + test_scientific_validation.py (analytic checks)
paper/ blueprint.md, perspective paper, figure_manifest.md
results/ written artifacts (gitignored; hashes tracked in manifest.json)