# E4 — Recombination supplies the rare tail; only a union-preserving *merge* realises it **Claim tested:** if several specialist models each remember a different slice of the rare tail, can combining them reconstruct the whole tail? And does *how* you combine them matter? **Setup (Layer 1, pure math).** `K = 500` items. We build `K_T` teacher models that each retain the rare tail with probability `q`, and we control how **correlated** their retained tails are with a single knob `ρ` (rho): `ρ = 0` = fully complementary teachers, `ρ = 1` = identical teachers. Swept: `K_T ∈ {1,2,3,5}`, `ρ ∈ {0, 0.25, 0.5, 0.75, 1}`, grounding `g ∈ {0, 0.02, 0.05}`, 200 repeats. ### Symbols - **`K_T`** number of teacher models; **`ρ`** how correlated their retained tails are (0 = diverse, 1 = clones). - **union coverage** — fraction of the tail covered by *at least one* teacher (the raw supply). - **surviving coverage** — fraction that actually survives in the pupil after it retrains on the combination. - **mean-distill** — pupil trained on the pooled/averaged teacher outputs. **max-merge** — keep, per item, the strongest teacher (a union-preserving merge, à la M2N2). ### The three panels 1. **Supply.** Union tail coverage vs `ρ`, one curve per `K_T`; solid lines are the exact closed form `U(K_T, ρ, q)`. More teachers and more diversity (lower `ρ`) supply more of the tail — and the simulation matches the formula exactly. 2. **Realisation.** Surviving coverage vs `ρ`. **Solid = max-merge rises** with more/diverse teachers; **dashed = mean-distill stays flat.** Averaging dilutes each teacher's rare items back below the survival threshold — the gain is supplied but not realised. 3. **The benefit needs the right operator (`ρ = 0`).** Surviving coverage vs `K_T` under both operators. Max-merge climbs with teacher count; mean-distill is flat — a **conservation law**: averaging's `1/K_T` dilution exactly cancels the union gain. ### Takeaway — "merge, don't average" Complementary specialists *contain* enough to rebuild the tail, but **naive multi-teacher distillation (averaging) throws it away**; you must combine them with a union-preserving merge. This is load-bearing for the paper's recombination claim and is re-tested in real neural weights in `results/recombination/`. **Falsifier (not triggered):** if mean-distill had also risen with `K_T`, or max never beat mean, the recombination story would collapse into "just average your models."