PACE
We use PACE's Phoenix cluster. This page covers what is specific to us. For everything else, see the PACE user documentation.
Connect
Add this to ~/.ssh/config on your laptop, using your GT username:
Host pace
HostName login-phoenix.pace.gatech.edu
User gburdell3
Host atl1-*
ProxyJump pace
User gburdell3
The second block matches every Phoenix compute node and routes through the login node, so you can reach whatever node a job gives you without editing this file again. Then ssh-copy-id pace to stop typing your password, and ssh pace to log in.
pace-quota lists the project accounts you can charge jobs to, along with your storage usage. Examples below write the account as paceship-dsgt_<project>.
Storage
Home holds 20 GB. One environment with PyTorch is 6 to 8 GB, since the CUDA libraries ship inside the wheels, so a couple of projects fill it. ~/scratch holds 15 TB. Work there.
Add to ~/.bashrc on PACE, then source ~/.bashrc:
export PATH=$HOME/.local/bin:$PATH
export XDG_CACHE_HOME=$HOME/scratch/.cache
export XDG_DATA_HOME=$HOME/scratch/.local/share
Those two variables redirect the caches that actually grow: uv's packages and interpreters, Hugging Face models and datasets, and PyTorch checkpoints. Now install uv, which picks them up:
curl -LsSf https://astral.sh/uv/install.sh | sh
Clone into scratch as well, cd ~/scratch && git clone .... Keeping the project next to the cache is not just tidiness: uv installs by hardlinking out of its cache, hardlinks cannot cross filesystems, and a project in home silently falls back to copying the full 6 to 8 GB into home anyway.
Scratch is not backed up and files unmodified for 60 days are deleted. So push code to git, rebuild environments with uv sync, and copy anything you cannot regenerate to your group's project storage, which is backed up and also listed by pace-quota.
When something on scratch starts throwing missing-file errors, a purge took part of it. Rebuild rather than debug: rm -rf .venv && uv sync.
Running jobs
Login nodes are for editing, moving data, and submitting jobs. Anything that computes belongs on a compute node.
Use embers, which is free and preemptible after a guaranteed first hour, so checkpoint your work. inferno is not preempted but spends your group's compute credits, so save it for runs that must finish.
Interactive session:
salloc --account=paceship-dsgt_<project> --qos=embers \
--nodes=1 --ntasks=1 --cpus-per-task=8 --mem-per-cpu=4G --time=2:00:00
hostname # expect atl1-..., if it says login-phoenix run: srun --pty bash
Add --gres=gpu:1 for a GPU and check it with uv run python -c "import torch; print(torch.cuda.is_available())". exit releases the allocation.
For longer work use a batch script, job.sbatch:
#!/bin/bash
#SBATCH --account=paceship-dsgt_<project>
#SBATCH --qos=embers
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=4G # per core, so 32G total here
#SBATCH --time=4:00:00 # killed at this limit
#SBATCH --gres=gpu:1 # omit for CPU-only work
#SBATCH --output=logs/%j.out
export XDG_CACHE_HOME=$HOME/scratch/.cache
export XDG_DATA_HOME=$HOME/scratch/.local/share
uv run python train.py
Batch jobs do not read ~/.bashrc, hence the exports. Submit from the project directory, since --output is relative to where you run sbatch:
cd ~/scratch/your-project && mkdir -p logs
sbatch job.sbatch
squeue -u $USER # your jobs
scancel <jobid> # stop one
tail -f logs/<jobid>.out # follow output
See the PACE user documentation for the full set of Slurm options.
Editing and notebooks
Allocate a node, note its hostname, then Remote-SSH: Connect to Host in VS Code and enter that node name. The atl1-* rule routes it. Open your project under ~/scratch and run Python: Select Interpreter to pick its .venv. Notebooks then run on the compute node with your project's packages, no separate Jupyter server needed.
Warning
Connect to a compute node, never to pace. The VS Code server is exactly the long-lived memory-hungry process login nodes forbid.
GitHub on PACE
Your laptop's authentication does not carry over, and you should never copy a private key onto a shared machine. Make PACE its own key with ssh-keygen -t ed25519, paste ~/.ssh/id_ed25519.pub into GitHub's SSH keys, and verify with ssh -T [email protected].