One self-contained README.md per results/ figure folder (Layer 1 E1-E6 and Layer 1.5 bridge/collapse/grounding/architectures/recombination): plain-language claim, setup, a compact symbol glossary, a panel-by-panel walkthrough, and the takeaway + falsifier. Auto-renders when browsing the folder; carries the honest caveats (grounding's ruler reframing, the excluded VAE). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
33 lines
2.5 KiB
Markdown
33 lines
2.5 KiB
Markdown
# E4 — Recombination supplies the rare tail; only a union-preserving *merge* realises it
|
||
|
||
**Claim tested:** if several specialist models each remember a different slice of the rare tail, can
|
||
combining them reconstruct the whole tail? And does *how* you combine them matter?
|
||
|
||
**Setup (Layer 1, pure math).** `K = 500` items. We build `K_T` teacher models that each retain the
|
||
rare tail with probability `q`, and we control how **correlated** their retained tails are with a
|
||
single knob `ρ` (rho): `ρ = 0` = fully complementary teachers, `ρ = 1` = identical teachers. Swept:
|
||
`K_T ∈ {1,2,3,5}`, `ρ ∈ {0, 0.25, 0.5, 0.75, 1}`, grounding `g ∈ {0, 0.02, 0.05}`, 200 repeats.
|
||
|
||
### Symbols
|
||
- **`K_T`** number of teacher models; **`ρ`** how correlated their retained tails are (0 = diverse, 1 = clones).
|
||
- **union coverage** — fraction of the tail covered by *at least one* teacher (the raw supply).
|
||
- **surviving coverage** — fraction that actually survives in the pupil after it retrains on the combination.
|
||
- **mean-distill** — pupil trained on the pooled/averaged teacher outputs. **max-merge** — keep, per item, the strongest teacher (a union-preserving merge, à la M2N2).
|
||
|
||
### The three panels
|
||
1. **Supply.** Union tail coverage vs `ρ`, one curve per `K_T`; solid lines are the exact closed form
|
||
`U(K_T, ρ, q)`. More teachers and more diversity (lower `ρ`) supply more of the tail — and the
|
||
simulation matches the formula exactly.
|
||
2. **Realisation.** Surviving coverage vs `ρ`. **Solid = max-merge rises** with more/diverse teachers;
|
||
**dashed = mean-distill stays flat.** Averaging dilutes each teacher's rare items back below the
|
||
survival threshold — the gain is supplied but not realised.
|
||
3. **The benefit needs the right operator (`ρ = 0`).** Surviving coverage vs `K_T` under both
|
||
operators. Max-merge climbs with teacher count; mean-distill is flat — a **conservation law**:
|
||
averaging's `1/K_T` dilution exactly cancels the union gain.
|
||
|
||
### Takeaway — "merge, don't average"
|
||
Complementary specialists *contain* enough to rebuild the tail, but **naive multi-teacher distillation
|
||
(averaging) throws it away**; you must combine them with a union-preserving merge. This is load-bearing
|
||
for the paper's recombination claim and is re-tested in real neural weights in `results/recombination/`.
|
||
**Falsifier (not triggered):** if mean-distill had also risen with `K_T`, or max never beat mean, the
|
||
recombination story would collapse into "just average your models."
|