The decisive experiment from the external review. 39 LoRA parent pairs (0.5B, 3 seeds) on three axes decorrelated by construction: conflict (contradictory conventions on shared prompts, private budgets fixed), compat (same prompts, SAME convention — overlap without conflict), and duration (weight divergence, zero conflict). Six pre-merge predictors; primary outcome = merge penalty (parent potential − merged achieved). League table (Spearman vs penalty, n=39): functional measures predict (dis_raw +0.460, epi_conf +0.446, p<0.005); geometry collapses (delta_cos +0.03, delta_l2 +0.17 n.s.); gradient alignment weak (−0.35); performance ~0. The first grid's apparent geometry win (+0.60) was an overlap/volume artifact — the compat control axis (added for exactly this) exposed and killed it: same overlap and data volume, zero penalty. Honest riders in the README: confidence weighting does not beat raw disagreement as a rank predictor (pre-registered internal prediction not confirmed; it does double the conflict/compat level contrast), and |rho|~0.45 is bounded by 0.5B merge-outcome noise (7B is the firm-up). Also: micro-batched gradient accumulation (OOM fix on the shared 16GB GPU), exact r-space LoRA-delta geometry (brute-force-verified test, 151 green), systemd-run runbook lesson (tmux dies with the SSH session scope on this box). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BkRLcc18rwT2Lysu6PbG7v
21 lines
1.7 KiB
Markdown
21 lines
1.7 KiB
Markdown
|
|
**2026-08-11 — Stop padding time estimates.** GG: "You generally massively overestimate the time it
|
|
takes to do some work... You're fast." Phase-1 items I scoped as "Week 1" took ~an hour. Rule: state
|
|
what will be done and in what order; give a duration only when compute-bound (training walltime), and
|
|
base it on measured runtimes, not human-project heuristics.
|
|
|
|
**2026-08-11 — Correspondence is not identity; interpretation is not prediction (manuscript reviews).**
|
|
The external review's core corrections, to internalise for all paper claims: (1) distinguish
|
|
interpretation / explanation / prediction and claim only the level the evidence supports; (2) a
|
|
minimal model being exactly Wright-Fisher does not make real training "literally" WF — our own
|
|
learning-kernel result says otherwise (cite it against ourselves); (3) name the operator every claim
|
|
is about (output-mean vs weight-average vs max-with-oracle vs routing are different objects with
|
|
different budgets); (4) don't write "nobody has / none imports" — invite no priority disputes; say
|
|
"to our knowledge" and state the positive contribution; (5) negative results (E13b) are strengths —
|
|
lead with them; (6) "control theory" needs states/controls/dynamics/rule or it's a "framework".
|
|
|
|
**2026-08-11 — tmux does NOT survive SSH disconnects on this machine.** The tmux server starts inside
|
|
the SSH session's systemd scope and gets reaped on logout (lost ~20 min of the epistasis grid; GG:
|
|
"SSH disconnected"). Reliable pattern here: `systemd-run --user --collect --unit=<name>
|
|
--working-directory="$PWD" bash -c '<cmd>'` — lands in user@.service (kept alive by the desktop
|
|
session), survives disconnects; check with `systemctl --user is-active <name>`, logs via redirect.
|