hpc: PBS job scripts for Imperial CX3 (probe + scaled LLM merge run)
Enable running the LLM tier on Imperial's HPC (scheduler: PBS Pro / qsub). - hpc/probe.pbs: 10-min 1-GPU reconnaissance job resolving the two unknowns the RCS docs omit -- compute-node internet access and the L40S driver's CUDA version -- plus TMPDIR/disk and available python/cuda modules. - hpc/llm_merge.pbs: scaled run on an L40S (48 GB), offline HF-cache wired, runs configs/llm/merge_hpc.yaml. - configs/llm/merge_hpc.yaml: Qwen2.5-7B-Instruct (fits the L40S) to reduce the noise that left the 0.5B prototype's overall-exceeds sign marginal. - hpc/README.md: the git-based workflow (login-node uv env + model pre-download -> qsub -> rsync results back), confirmed PBS/GPU directives, and the code TODOs for the definitive run (multi-seed, more families, directed/dilution-resistant merge). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
809e45a5e0
commit
38bb252c20
4 changed files with 135 additions and 0 deletions
20
configs/llm/merge_hpc.yaml
Normal file
20
configs/llm/merge_hpc.yaml
Normal file
|
|
@ -0,0 +1,20 @@
|
|||
experiment: llm_merge_hpc
|
||||
kind: llm_merge
|
||||
seed: 1
|
||||
n_replicates: 1
|
||||
|
||||
# Scaled version of configs/llm/merge.yaml for an L40S (48 GB): a bigger, more capable base so the
|
||||
# specialists decorrelate cleanly and the "recombined model exceeds any single specialist" sign is
|
||||
# less noisy than the 0.5B local prototype (where it was only marginal on overall accuracy).
|
||||
# NOTE (code TODOs for the *definitive* run, not yet built): loop several seeds and report mean±CI;
|
||||
# add more task families; and add a dilution-resistant / offspring-selected ("directed sex") merge.
|
||||
|
||||
base_model: Qwen/Qwen2.5-7B-Instruct # ~15 GB bf16 + LoRA fits the L40S; fallback: 3B-Instruct
|
||||
families: [lists, strings, arith]
|
||||
n_train: 800
|
||||
n_test: 200
|
||||
epochs: 3
|
||||
lora: {r: 16, alpha: 32}
|
||||
merges: [soup, ties]
|
||||
|
||||
output: {dir: results/llm_merge_hpc}
|
||||
Loading…
Add table
Add a link
Reference in a new issue