hpc: PBS job scripts for Imperial CX3 (probe + scaled LLM merge run)

Enable running the LLM tier on Imperial's HPC (scheduler: PBS Pro / qsub).
- hpc/probe.pbs: 10-min 1-GPU reconnaissance job resolving the two unknowns
  the RCS docs omit -- compute-node internet access and the L40S driver's
  CUDA version -- plus TMPDIR/disk and available python/cuda modules.
- hpc/llm_merge.pbs: scaled run on an L40S (48 GB), offline HF-cache wired,
  runs configs/llm/merge_hpc.yaml.
- configs/llm/merge_hpc.yaml: Qwen2.5-7B-Instruct (fits the L40S) to reduce
  the noise that left the 0.5B prototype's overall-exceeds sign marginal.
- hpc/README.md: the git-based workflow (login-node uv env + model
  pre-download -> qsub -> rsync results back), confirmed PBS/GPU directives,
  and the code TODOs for the definitive run (multi-seed, more families,
  directed/dilution-resistant merge).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Giorgio Gilestro 2026-07-05 16:02:09 +01:00
parent 809e45a5e0
commit 38bb252c20
4 changed files with 135 additions and 0 deletions

View file

@ -0,0 +1,20 @@
experiment: llm_merge_hpc
kind: llm_merge
seed: 1
n_replicates: 1
# Scaled version of configs/llm/merge.yaml for an L40S (48 GB): a bigger, more capable base so the
# specialists decorrelate cleanly and the "recombined model exceeds any single specialist" sign is
# less noisy than the 0.5B local prototype (where it was only marginal on overall accuracy).
# NOTE (code TODOs for the *definitive* run, not yet built): loop several seeds and report mean±CI;
# add more task families; and add a dilution-resistant / offspring-selected ("directed sex") merge.
base_model: Qwen/Qwen2.5-7B-Instruct # ~15 GB bf16 + LoRA fits the L40S; fallback: 3B-Instruct
families: [lists, strings, arith]
n_train: 800
n_test: 200
epochs: 3
lora: {r: 16, alpha: 32}
merges: [soup, ties]
output: {dir: results/llm_merge_hpc}