docs: add HPC integration instructions
This commit is contained in:
parent
a774649ff0
commit
0e821cf251
1 changed files with 30 additions and 0 deletions
30
README.md
30
README.md
|
|
@ -15,3 +15,33 @@ example command:
|
||||||
```bash
|
```bash
|
||||||
uv run src/train.py --model_name my_model --epochs 50 --batch_size 32
|
uv run src/train.py --model_name my_model --epochs 50 --batch_size 32
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## HPC Integration
|
||||||
|
|
||||||
|
To run training on the VSC HPC (Ghent Tier 2):
|
||||||
|
|
||||||
|
### Setup
|
||||||
|
|
||||||
|
Run the installation script once on a login node. This sets up a virtual environment in `$VSC_DATA` with the required system modules (JAX, Flax, WandB) and lightweight dependencies.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bash scripts/hpc_install.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
### Smoke Test
|
||||||
|
Verify everything runs on a compute node:
|
||||||
|
```bash
|
||||||
|
python src/train.py --config configs/hpc_smoke_test.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
### Submitting Jobs
|
||||||
|
Submit long-running experiments via the provided Slurm script:
|
||||||
|
```bash
|
||||||
|
sbatch scripts/hpc_train.slurm
|
||||||
|
```
|
||||||
|
|
||||||
|
### Syncing WandB
|
||||||
|
Since compute nodes are offline, sync your WandB logs after the job finishes:
|
||||||
|
```bash
|
||||||
|
wandb sync runs/<run_name>/wandb
|
||||||
|
```
|
||||||
|
|
|
||||||
Reference in a new issue