# HPC Guide
Full documentation:
## Cluster Selection
Choose the appropriate cluster before submitting a job with `module swap cluster/`. The default login cluster is **doduo**.
> Check the current queue load at .
## Storage Overview
- Run outputs are written to `$VSC_SCRATCH` during the job (fast I/O) and copied to `$VSC_DATA` at the end for persistence.
- `$VSC_SCRATCH` may be purged periodically — do not use it as long-term storage.
Check your quota: (Usage section).
## Initial Environment Setup
Run **once** after cloning the repository, but ensure you are logged into a **compute node** on the target cluster (e.g., `donphan` or `joltik`). The login node (`doduo`) will block the installation script to prevent architecture mismatches.
```bash
# 1. Swap to the target cluster
module swap cluster/joltik
# 2. Start an interactive session on a compute node
qsub -I -l nodes=1:ppn=8:gpus=1
# 3. Navigate to the project directory and run the install script
cd SEL3-2026-Groep-4
bash scripts/hpc/install.sh
# 4. Exit the interactive session once finished
exit
```
> [!IMPORTANT]
> Virtual environments are cluster-specific. If you want to switch from `joltik` to `accelgor`, you must re-run the `install.sh` script while logged into an interactive session on `accelgor`.
## Interactive Debugging on donphan
### Option A — Interactive shell session
```bash
# Swap to the debug cluster (from any login node)
module swap cluster/donphan
# Request an interactive job (1 node, 4 cores)
qsub -I -l nodes=1:ppn=4 -l walltime=1:00:00
# Once inside the job — redirect caches to scratch first
export PIP_CACHE_DIR="$VSC_SCRATCH/.cache/pip"
export UV_CACHE_DIR="$VSC_SCRATCH/.cache/uv"
# Change to the project directory and activate environment
cd "$PBS_O_WORKDIR"
module load vsc-venv
set +u
source vsc-venv --activate \
--modules env/hpc/modules.txt \
--requirements env/hpc/requirements.txt
set -u
# Set headless rendering backend
export MUJOCO_GL=egl
# Run the smoke test
python src/train.py --config configs/hpc/smoke_test.yaml
```
### Option B — JupyterLab session (HPC web portal)
1. In the web portal go to **Interactive Apps → JupyterLab RHEL9**
2. Set the following options:
| Option | Value |
|---------------------|-------------------------------------------------|
| Cluster | `donphan (interactive/debug)` |
| Number of nodes | 1 |
| Number of cores | 4 |
| JupyterLab version | `4.2.5 GCCcore-13.3.0` |
| Custom code | *(leave blank — vsc-venv handles modules)* |
3. Click **Launch**, wait for the session to start, then **Connect**.
4. In JupyterLab, select the kernel **`SEL3 ()`**.
5. Verify GPU access:
```python
import jax
print(jax.default_backend()) # expected: 'gpu'
print(jax.devices()) # expected: [CudaDevice(id=0)]
```
> **Warning:** JAX can only be loaded by one kernel at a time. Shut down other kernels before switching notebooks.
## Submitting Batch Training Jobs
```bash
# Default cluster (joltik — one A100 GPU slice)
qsub scripts/hpc/train.pbs
# Switch to a different GPU cluster first
module swap cluster/accelgor
qsub scripts/hpc/train.pbs
```
The job script automatically:
- Writes run outputs to `$VSC_SCRATCH/runs/` during the run using the `--run_dir` argument. This ensures that frequent I/O (like tensorboard logs and checkpoints) happens on the fastest available filesystem.
- Copies the final results to `$VSC_DATA/runs/` on completion for long-term persistence.
- Writes PBS stdout/stderr to `runs/brittlestar-ppo.o` / `.e` (standard PBS convention, relative to the project root).
Monitor your jobs:
```bash
qstat # list your jobs
qstat -f # detailed info for a specific job
qdel # cancel a job
```
## Managing Dependencies
`env/hpc/requirements.txt` is auto-generated by CI whenever `pyproject.toml` changes. To regenerate locally:
```bash
uv run scripts/export_hpc_requirements.py
```
Do **not** edit `env/hpc/requirements.txt` by hand — edit `pyproject.toml` instead.