MachineSex/results/E4
Giorgio Gilestro 84124de143 Manuscript revision and pending experiment work, snapshot before restructuring
Clarity pass over the main text (36-item audit), Discussion rewrite and cut,
acknowledgements, Souly et al. as ref 62, lettered SI panels, model section
moved under Results; plus the untracked curriculum/society/compose/smol
configs, runners, figures, stats and tests that the SI already cites.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 16:54:09 +01:00
..
E4.pdf Manuscript revision and pending experiment work, snapshot before restructuring 2026-09-13 16:54:09 +01:00
E4.png Manuscript revision and pending experiment work, snapshot before restructuring 2026-09-13 16:54:09 +01:00
manifest.json Layer 1 complete: E3-E6 + E2 analysis add-ons 2026-07-04 18:54:42 +02:00
README.md docs: accessible figure legends (README.md) for all figures 2026-07-05 08:43:04 +01:00
resolved_config.yaml Layer 1 complete: E3-E6 + E2 analysis add-ons 2026-07-04 18:54:42 +02:00

E4 — Recombination supplies the rare tail; only a union-preserving merge realises it

Claim tested: if several specialist models each remember a different slice of the rare tail, can combining them reconstruct the whole tail? And does how you combine them matter?

Setup (Layer 1, pure math). K = 500 items. We build K_T teacher models that each retain the rare tail with probability q, and we control how correlated their retained tails are with a single knob ρ (rho): ρ = 0 = fully complementary teachers, ρ = 1 = identical teachers. Swept: K_T ∈ {1,2,3,5}, ρ ∈ {0, 0.25, 0.5, 0.75, 1}, grounding g ∈ {0, 0.02, 0.05}, 200 repeats.

Symbols

  • K_T number of teacher models; ρ how correlated their retained tails are (0 = diverse, 1 = clones).
  • union coverage — fraction of the tail covered by at least one teacher (the raw supply).
  • surviving coverage — fraction that actually survives in the pupil after it retrains on the combination.
  • mean-distill — pupil trained on the pooled/averaged teacher outputs. max-merge — keep, per item, the strongest teacher (a union-preserving merge, à la M2N2).

The three panels

  1. Supply. Union tail coverage vs ρ, one curve per K_T; solid lines are the exact closed form U(K_T, ρ, q). More teachers and more diversity (lower ρ) supply more of the tail — and the simulation matches the formula exactly.
  2. Realisation. Surviving coverage vs ρ. Solid = max-merge rises with more/diverse teachers; dashed = mean-distill stays flat. Averaging dilutes each teacher's rare items back below the survival threshold — the gain is supplied but not realised.
  3. The benefit needs the right operator (ρ = 0). Surviving coverage vs K_T under both operators. Max-merge climbs with teacher count; mean-distill is flat — a conservation law: averaging's 1/K_T dilution exactly cancels the union gain.

Takeaway — "merge, don't average"

Complementary specialists contain enough to rebuild the tail, but naive multi-teacher distillation (averaging) throws it away; you must combine them with a union-preserving merge. This is load-bearing for the paper's recombination claim and is re-tested in real neural weights in results/recombination/. Falsifier (not triggered): if mean-distill had also risen with K_T, or max never beat mean, the recombination story would collapse into "just average your models."