One self-contained README.md per results/ figure folder (Layer 1 E1-E6 and Layer 1.5 bridge/collapse/grounding/architectures/recombination): plain-language claim, setup, a compact symbol glossary, a panel-by-panel walkthrough, and the takeaway + falsifier. Auto-renders when browsing the folder; carries the honest caveats (grounding's ruler reframing, the excluded VAE). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2.5 KiB
E4 — Recombination supplies the rare tail; only a union-preserving merge realises it
Claim tested: if several specialist models each remember a different slice of the rare tail, can combining them reconstruct the whole tail? And does how you combine them matter?
Setup (Layer 1, pure math). K = 500 items. We build K_T teacher models that each retain the
rare tail with probability q, and we control how correlated their retained tails are with a
single knob ρ (rho): ρ = 0 = fully complementary teachers, ρ = 1 = identical teachers. Swept:
K_T ∈ {1,2,3,5}, ρ ∈ {0, 0.25, 0.5, 0.75, 1}, grounding g ∈ {0, 0.02, 0.05}, 200 repeats.
Symbols
K_Tnumber of teacher models;ρhow correlated their retained tails are (0 = diverse, 1 = clones).- union coverage — fraction of the tail covered by at least one teacher (the raw supply).
- surviving coverage — fraction that actually survives in the pupil after it retrains on the combination.
- mean-distill — pupil trained on the pooled/averaged teacher outputs. max-merge — keep, per item, the strongest teacher (a union-preserving merge, à la M2N2).
The three panels
- Supply. Union tail coverage vs
ρ, one curve perK_T; solid lines are the exact closed formU(K_T, ρ, q). More teachers and more diversity (lowerρ) supply more of the tail — and the simulation matches the formula exactly. - Realisation. Surviving coverage vs
ρ. Solid = max-merge rises with more/diverse teachers; dashed = mean-distill stays flat. Averaging dilutes each teacher's rare items back below the survival threshold — the gain is supplied but not realised. - The benefit needs the right operator (
ρ = 0). Surviving coverage vsK_Tunder both operators. Max-merge climbs with teacher count; mean-distill is flat — a conservation law: averaging's1/K_Tdilution exactly cancels the union gain.
Takeaway — "merge, don't average"
Complementary specialists contain enough to rebuild the tail, but naive multi-teacher distillation
(averaging) throws it away; you must combine them with a union-preserving merge. This is load-bearing
for the paper's recombination claim and is re-tested in real neural weights in results/recombination/.
Falsifier (not triggered): if mean-distill had also risen with K_T, or max never beat mean, the
recombination story would collapse into "just average your models."