MachineSex/paper/response-to-review-3.md
Giorgio Gilestro ab3dc10587 Restructure: descriptive tier and experiment names, paper/manuscript
- paper/pnas -> paper/manuscript (venue-neutral)
- configs/layer1 -> configs/inheritance, src/knowledge -> src/inheritance
  (imported as `inheritance`), make layer1 -> make inheritance; layer2 alias dropped
- inheritance and trained-network bundles named after the manuscript figure
  they feed (fig2_grounding_sweep, figS3_rebaselining, ...), or descriptively
  where they feed none; configs keep their `experiment:` value so parquet
  hashes are unchanged, only output.dir moves
- figure scripts, SI figure sources, notebooks, REPRODUCING.md, README and the
  SI Methods/tables updated; make clean no longer deletes tracked manifests;
  reproduce.sh hashes the s{seed}/ layouts too

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y64o8FKP7rCuXzC48pxpMm
2026-09-13 17:00:40 +01:00

96 lines
7.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Response to the third review (of the PNAS-format draft)
*All five priority fixes are made, plus the presentation items. The revised draft is
`paper/manuscript/main.md` (rebuilt PDF alongside); the long-form document and the results documentation
were corrected wherever they carried the same overstatements. Point-by-point:*
## 1. The averaging proposition (your §2) — you are right, and the text now proves what it claims
Your convexity argument is correct: conservation of expected mass does not establish that averaging
cannot help, because extinction is convex in mixed mass and averaging reduces its variance. Our result
is, exactly as you diagnosed, a **first-order cancellation in the rare-item regime**, and the main
text now states the actual proposition with its quantities and assumptions: K parents with independent
retention; child draws `n` samples from one random parent vs the parents' output-mean; expected mass
identical; and in the regime `n·p/K ≪ 1`, where per-item survival is first-order in sampled mass,
expected survival is identical too. Two boundaries follow in the same paragraph: outside that regime
averaging's variance reduction can *reduce* extinction relative to a random single parent (your
argument, credited to the review process); and the union operator's renormalisation (which itself
redistributes mass) and oracle requirement are stated. "Adding parents cannot help" is deleted here
and in every other document that carried it. We agree the interesting content is the consequence for
retention, not the elementary conservation of a mean — which is how the proposition is now framed.
## 2. Grounding (your §3) — threshold made operational, floor made probabilistic, rule de-categoricalised
- `g* ≈ 0.05` is now explicitly an **operational threshold**, with the text stating what our own
analysis always showed: the immigrationdrift equilibrium is *smooth* in the grounding fraction (no
phase transition in aggregate diversity). New wording: under the tested population size and Zipf
source, `g ≈ 0.05` retained ≥95% of equilibrium diversity, with dependence on sample size, source,
and retention target (SI).
- `m·p ≳ 1` is restated as what it is: `1 e^{m·p}` observation probability per batch (~63% at
`m·p = 1`), confidence-dependent, with retention vs stationary occupancy vs reintroduction
distinguished (immigration can restore an absent item).
- The design rule now reads in your form: under unstratified grounding rare capabilities are expensive
(targeted sampling changes the cost); recombination recovers rare capabilities *still retained
across complementary parents*.
## 3. Grounded inheritance vs grounded evaluation (your §4) — separated and named
The society section now opens with the definitional distinction: **grounded inheritance** (external
samples in the reproduction process — the data channel) vs **grounded evaluation** (true fitness vs
conformity in selection — the fitness channel), related but different operators, connected only in
that both couple the lineage to a non-drifting external signal. The section is retitled to your
formulation ("…make complementary contributions"), the ablation is described as separating failure
modes *under the tested conditions*, and general joint necessity is explicitly disclaimed (alternative
mutation/restart/archive/selection schemes noted). Table 1's corresponding row now says
"complementary… in the tested society"; the same fix is propagated to the long-form document.
## 4. The alignment contradiction (your §5) — deleted, both statements reconciled
"This cannot be an alignment failure, because the same aligner succeeded on the control" is removed
everywhere (manuscript, long-form document, results documentation), replaced by your formulation: the
tested alignment removes the same-task barrier but leaves the conflict-associated barrier largely
unchanged — supporting a functional-conflict interpretation without proving optimal alignment. The
abstract now says "remaining after permutation-and-rescaling alignment" (not "surviving the full
symmetry group"), and the Methods note that the group is the alignment's *search space*, with control
recovery not establishing global optimality. The discussion's "expect specialisation alone to be
merge-safe" is replaced by the supported lesson: **do not treat divergence or specialisation alone as
evidence of incompatibility.**
## 5. Headline vs detail (your §6) — matched, and the seed-dependence analysed
The significance statement now ends with your suggested sentence (a controlled small-model test…
motivating further comparison). On the clustering point: you are right that condition-clustering does
not capture cross-condition dependence through shared task-data seeds. We added the sensitivity you
asked for (committed to the statistics script): **per-seed correlations** — each seed alone, n = 13
conditions — are stable for the functional measures (+0.37 to +0.53 in every individual seed) and ≈0
for geometry in every seed; leave-one-seed-out ranges are [+0.38, +0.56] (functional) vs
[0.04, +0.28] (geometry). One informative surprise: gradient alignment is *seed-unstable*
(0.11 to 0.55), which the manuscript now reports as its own caveat. The text also states plainly
that with three seeds, uncertainty about seed generalisation remains substantial.
## 6. Presentation (your §7) — done
Meta-language removed ("the honest statement", "sharpest honesty", "earn their place by tempering",
"honest deviations" — all gone; results are stated, not described as disclosures). "Exact" is now
reserved for closed-form mathematics — NK/simulation results are labelled "analytic model" in Table 1
and the text. The headroom relationship is stated qualitatively with "a quantitative form is
untested". "Directed sex with no biological analogue" is replaced by your phrasing (the shorthand kept,
defined as engineered recombination with flexible parent choice and pre-deployment screening).
Muller's ratchet is now a *consequence-level* correspondence, with the text stating that irreversible
loss alone does not identify the ratchet's mechanism. A compact results table (Table 2: setting/n,
outcome definition, headline with uncertainty, for the eight headline results) is added before the
Discussion. Reference numbering and the figure files accompany the rebuilt PDF; the bespoke unified
figures and journal-format reflow remain flagged as submission-time work.
## One point of information, not disagreement
On §2's closing remark — that conservation of an arithmetic mean's expectation is elementary and the
contribution must lie in its consequences — we agree, and would only note that the consequence now
stated (first-order cancellation of the multi-parent retention gain under output-mean inheritance,
against union-operator retention growth, in the regime where the deep tail actually lives) is the
claim we intended all along; the earlier wording claimed more than this and is gone.
Your bottom-line formulation — minimal models establish conditional results; neural experiments reveal
where the correspondences hold and break; a controlled predictive test motivates measuring functional
conflict before merging — is now, near-verbatim, how the paper describes itself. Thank you for three
rounds of genuinely improving review.