# HPC Guide Full documentation: ## Cluster Selection Choose the appropriate cluster before submitting a job with `module swap cluster/`. The default login cluster is **doduo**. > Check the current queue load at . ## Storage Overview - Run outputs are written to `$VSC_SCRATCH` during the job (fast I/O) and copied to `$VSC_DATA` at the end for persistence. - `$VSC_SCRATCH` may be purged periodically — do not use it as long-term storage. Check your quota: (Usage section). ## Initial Environment Setup Run **once** after cloning the repository, but ensure you are logged into a **compute node** on the target cluster (e.g., `donphan` or `joltik`). The login node (`doduo`) will block the installation script to prevent architecture mismatches. ```bash # 1. Swap to the target cluster module swap cluster/joltik # 2. Start an interactive session on a compute node qsub -I -l nodes=1:ppn=8:gpus=1 # 3. Navigate to the project directory and run the install script cd SEL3-2026-Groep-4 bash scripts/hpc/install.sh # 4. Exit the interactive session once finished exit ``` > [!IMPORTANT] > Virtual environments are cluster-specific. If you want to switch from `joltik` to `accelgor`, you must re-run the `install.sh` script while logged into an interactive session on `accelgor`. ## Interactive Debugging on donphan ### Option A — Interactive shell session ```bash # Swap to the debug cluster (from any login node) module swap cluster/donphan # Request an interactive job (1 node, 4 cores) qsub -I -l nodes=1:ppn=4 -l walltime=1:00:00 # Once inside the job — redirect caches to scratch first export PIP_CACHE_DIR="$VSC_SCRATCH/.cache/pip" export UV_CACHE_DIR="$VSC_SCRATCH/.cache/uv" # Change to the project directory and activate environment cd "$PBS_O_WORKDIR" module load vsc-venv set +u source vsc-venv --activate \ --modules env/hpc/modules.txt \ --requirements env/hpc/requirements.txt set -u # Set headless rendering backend export MUJOCO_GL=egl # Run the smoke test python src/train.py --config configs/hpc/smoke_test.yaml ``` ### Option B — JupyterLab session (HPC web portal) 1. In the web portal go to **Interactive Apps → JupyterLab RHEL9** 2. Set the following options: | Option | Value | |---------------------|-------------------------------------------------| | Cluster | `donphan (interactive/debug)` | | Number of nodes | 1 | | Number of cores | 4 | | JupyterLab version | `4.2.5 GCCcore-13.3.0` | | Custom code | *(leave blank — vsc-venv handles modules)* | 3. Click **Launch**, wait for the session to start, then **Connect**. 4. In JupyterLab, select the kernel **`SEL3 ()`**. 5. Verify GPU access: ```python import jax print(jax.default_backend()) # expected: 'gpu' print(jax.devices()) # expected: [CudaDevice(id=0)] ``` > **Warning:** JAX can only be loaded by one kernel at a time. Shut down other kernels before switching notebooks. ## Submitting Batch Training Jobs ```bash # Default cluster (joltik — one A100 GPU slice) qsub scripts/hpc/train.pbs # Switch to a different GPU cluster first module swap cluster/accelgor qsub scripts/hpc/train.pbs ``` The job script automatically: - Writes run outputs to `$VSC_SCRATCH/runs/` during the run using the `--run_dir` argument. This ensures that frequent I/O (like tensorboard logs and checkpoints) happens on the fastest available filesystem. - Copies the final results to `$VSC_DATA/runs/` on completion for long-term persistence. - Writes PBS stdout/stderr to `runs/brittlestar-ppo.o` / `.e` (standard PBS convention, relative to the project root). Monitor your jobs: ```bash qstat # list your jobs qstat -f # detailed info for a specific job qdel # cancel a job ``` ## Managing Dependencies `env/hpc/requirements.txt` is auto-generated by CI whenever `pyproject.toml` changes. To regenerate locally: ```bash uv run scripts/export_hpc_requirements.py ``` Do **not** edit `env/hpc/requirements.txt` by hand — edit `pyproject.toml` instead.