MachineSex/results/E4
Giorgio Gilestro 3b9f4f7893 docs: accessible figure legends (README.md) for all figures
One self-contained README.md per results/ figure folder (Layer 1 E1-E6
and Layer 1.5 bridge/collapse/grounding/architectures/recombination):
plain-language claim, setup, a compact symbol glossary, a panel-by-panel
walkthrough, and the takeaway + falsifier. Auto-renders when browsing the
folder; carries the honest caveats (grounding's ruler reframing, the
excluded VAE).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 08:43:04 +01:00
..
E4.pdf Layer 1 complete: E3-E6 + E2 analysis add-ons 2026-07-04 18:54:42 +02:00
E4.png Layer 1 complete: E3-E6 + E2 analysis add-ons 2026-07-04 18:54:42 +02:00
manifest.json Layer 1 complete: E3-E6 + E2 analysis add-ons 2026-07-04 18:54:42 +02:00
README.md docs: accessible figure legends (README.md) for all figures 2026-07-05 08:43:04 +01:00
resolved_config.yaml Layer 1 complete: E3-E6 + E2 analysis add-ons 2026-07-04 18:54:42 +02:00

E4 — Recombination supplies the rare tail; only a union-preserving merge realises it

Claim tested: if several specialist models each remember a different slice of the rare tail, can combining them reconstruct the whole tail? And does how you combine them matter?

Setup (Layer 1, pure math). K = 500 items. We build K_T teacher models that each retain the rare tail with probability q, and we control how correlated their retained tails are with a single knob ρ (rho): ρ = 0 = fully complementary teachers, ρ = 1 = identical teachers. Swept: K_T ∈ {1,2,3,5}, ρ ∈ {0, 0.25, 0.5, 0.75, 1}, grounding g ∈ {0, 0.02, 0.05}, 200 repeats.

Symbols

  • K_T number of teacher models; ρ how correlated their retained tails are (0 = diverse, 1 = clones).
  • union coverage — fraction of the tail covered by at least one teacher (the raw supply).
  • surviving coverage — fraction that actually survives in the pupil after it retrains on the combination.
  • mean-distill — pupil trained on the pooled/averaged teacher outputs. max-merge — keep, per item, the strongest teacher (a union-preserving merge, à la M2N2).

The three panels

  1. Supply. Union tail coverage vs ρ, one curve per K_T; solid lines are the exact closed form U(K_T, ρ, q). More teachers and more diversity (lower ρ) supply more of the tail — and the simulation matches the formula exactly.
  2. Realisation. Surviving coverage vs ρ. Solid = max-merge rises with more/diverse teachers; dashed = mean-distill stays flat. Averaging dilutes each teacher's rare items back below the survival threshold — the gain is supplied but not realised.
  3. The benefit needs the right operator (ρ = 0). Surviving coverage vs K_T under both operators. Max-merge climbs with teacher count; mean-distill is flat — a conservation law: averaging's 1/K_T dilution exactly cancels the union gain.

Takeaway — "merge, don't average"

Complementary specialists contain enough to rebuild the tail, but naive multi-teacher distillation (averaging) throws it away; you must combine them with a union-preserving merge. This is load-bearing for the paper's recombination claim and is re-tested in real neural weights in results/recombination/. Falsifier (not triggered): if mean-distill had also risen with K_T, or max never beat mean, the recombination story would collapse into "just average your models."