Confirms model collapse and its arrest by grounding on REAL images, not just the synthetic sandbox. A conv VAE (the canonical generative-collapse model) is retrained each generation on its own generated digits, with a fraction g of fresh real MNIST mixed in. Modes = digit class x stroke- thickness bin (K=30, Zipf, ~18 tail modes); the oracle is a frozen CNN + deterministic thickness at 98.5% mode accuracy (30x30 confusion matrix recorded in the manifest as the measurement-noise floor). Result (4 reps): dry (g=0) collapses to a single mode -- forward-KL 0.5->18, support 30->1, tail 1.0->0.06, H->0 -- while 10% grounding holds all 30 modes (KL~0.6, full tail, H~0.9). Signs, not magnitudes (blueprint 3.5); the exact synthetic oracle stays the quantitative anchor. The VAE needs ~10% grounding vs the synthetic histogram's ~5%, consistent with the grounding finding that trained nets need more than the exact operator. Plugs into the existing data-agnostic contract (metrics/grounding/output reused verbatim): mnist_data (thickness bins, class x thickness bijection, MnistSampler), mnist_oracle (ClassifierOracle + confusion matrix), mnist_vae (ConvVAEGenerator), mnist_loop (run_mnist_lineage), kind= mnist_lineage dispatch, MnistCfg/OracleCfg. Figures: plot_mnist (parquet- only) + mnist_montage (eyeball diagnostic showing digits degenerate to one blurry mode). make mnist / make env-mnist, kept out of the make neural loop. 99 tests green (+5 torchvision-gated). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
39 lines
2.9 KiB
Markdown
39 lines
2.9 KiB
Markdown
# mnist_collapse — collapse and grounding-rescue on REAL MNIST images (external validity)
|
||
|
||
**Claim tested:** everything so far used a *synthetic* sandbox with a zero-error decoder oracle. Do
|
||
model collapse and its rescue by grounding also appear on **real images** with a **classifier**
|
||
oracle — i.e. is the effect real, not a synthetic artefact?
|
||
|
||
**Setup (Layer 1.5, real-data tier).** The generative model is a **convolutional VAE** (the model
|
||
in which generative collapse was first observed). Each generation a **fresh** VAE is trained from
|
||
scratch on the previous VAE's own generated digits, plus a fraction `g` of fresh **real** MNIST
|
||
images (grounding). `K = 30` **modes** = digit class × stroke-thickness bin (S=3), Zipf-resampled so
|
||
the rarest ~18 modes form a real tail. The **oracle** is a frozen CNN (digit class) + deterministic
|
||
thickness bin; its **mode accuracy ≈ 98.5%** (recorded in `manifest.json` with the full 30×30
|
||
confusion matrix) is the measurement-noise floor. Two arms — dry (`g = 0`) vs grounded (`g = 0.1`) —
|
||
`n = 6000` images/generation, 15 generations, 4 replicates.
|
||
|
||
### Symbols
|
||
- **mode** = (digit class, stroke-thickness bin); **`p*`** = Zipf truth over the 30 modes; **`p̂`** = the VAE's oracle-measured mode distribution.
|
||
- **`g`** = grounding fraction (share of real MNIST images each generation). **forward-KL** = distance from truth; **support** = distinct modes alive; **`H`** = diversity; **tail truth-mass alive** = fraction of the rare tail retained.
|
||
|
||
### The four panels (dry = red, grounded = green; band = 95% CI over 4 reps)
|
||
1. **Forward-KL.** Dry climbs from ~0.5 to **~18** (the VAE drifts far from truth); grounded stays
|
||
near the floor. Collapse is real on images.
|
||
2. **Support.** Dry collapses from all 30 modes to **~1** (the VAE ends up emitting a single blurry
|
||
mode); grounded holds all 30.
|
||
3. **Tail truth-mass alive.** Dry's rare tail is wiped out (→ 0.06); grounded keeps the whole tail.
|
||
4. **Heterozygosity.** Dry diversity → 0; grounded holds `H ≈ 0.9`.
|
||
|
||
See **`mnist_montage.png`** for the eyeball version: gen-0 digits are varied and recognisable; by
|
||
gen 12–15 the dry lineage has degenerated into one blurry blob.
|
||
|
||
### Takeaway
|
||
Model collapse and its arrest by a small dose of real data **reproduce on real MNIST images with a
|
||
learned classifier oracle** — external validity for the whole Layer-1.5 story. Note the VAE needs
|
||
~10% grounding here (vs ~5% for the synthetic histogram), consistent with the `grounding` finding
|
||
that trained neural models need somewhat more grounding than the exact operator. This is
|
||
**confirmation-only** (signs, not magnitudes; blueprint §3.5) — the exact synthetic oracle remains
|
||
the anchor for every quantitative claim, and the oracle confusion matrix is the recorded noise floor.
|
||
**Falsifier (not triggered):** if the dry VAE had shown no diversity loss, or grounding had failed to
|
||
arrest it, the external-validity claim would fail.
|