Skip to content

Slurm clusters

A Slurm tenancy is a dedicated cluster: a login node you ssh into, a controller triplet the platform runs, and GPU compute nodes the platform moves in and out of your cluster. You submit jobs with the standard Slurm commands (srun, sbatch, squeue, sacct); nobody runs them for you. Slurm 26.05 with the enroot and pyxis container plugins is installed on every node.

Connecting

Every Slurm cluster has one login account, tenant, with root privileges through sudo. It is a member of the GPU device groups (video, render) on every node, so GPU tools and runtimes work without sudo. The account accepts the SSH public keys the platform team registered for your organization; there is no password and no sign-in flow.

ssh tenant@<login-address>

The login address is on your cluster's page in the console under Connect. To add or remove a key, send the platform team the public key line (ssh-ed25519 AAAA…); the change reaches the login node within a minute of being registered. Keys you append to ~/.ssh/authorized_keys yourself work on the compute nodes as well, because the home directory is shared across the cluster (see Storage) — the login's managed set is the one the platform keeps.

The setup script writes an ssh <cluster> alias for the account when the cluster uses key login:

curl -fsSL https://access.nationalcompute.com/setup.sh | sh -s -- you@example.com <cluster>

Jobs

Jobs run in the main partition. Request GPUs with --gres; each requested GPU brings a default CPU allocation sized for collective libraries, so a whole-node job needs no CPU flags:

srun -N1 --gres=gpu:8 hostname
sbatch --nodes=4 --ntasks-per-node=8 --gres=gpu:8 train.sbatch
limit value
maximum wall time (--time) none — a job runs until it finishes or reaches the --time it asked for
default wall time when --time is absent unlimited
default CPUs per requested GPU 28 (currently; sized so a gpu:8 job holds the node)
GPUs per compute node 8

A job that requests fewer GPUs than a node has shares the node with other jobs; ROCR_VISIBLE_DEVICES inside the job lists the GPUs it holds, and the others are not visible to it. The whole GPU-node memory is allocatable to a full-node job.

Accounting is per cluster: every job is charged to the cluster's single account, and sacct shows it with gres/gpu in AllocTRES. There are no per-user limits, quotas or fair-share weights.

Job hooks

Every compute node runs the executables in two directories around each job, as root, with Slurm's job environment (SLURM_JOB_ID, SLURM_JOB_USER, SLURM_JOB_GPUS):

directory when time limit per hook
/etc/slurm/prolog.d/ before the job's first step, after the platform's health gate has passed 20 s
/etc/slurm/epilog.d/ after the job's last step, including cancelled and timed-out jobs 60 s

Hooks run in name order. A hook that exits non-zero or reaches its time limit is logged to the node's syslog (tags slurm-prolog and slurm-epilog) and otherwise ignored: it cannot fail the job and cannot drain the node. Files without the executable bit are skipped.

The directories are yours; write to them with sudo on each node. They are per-node state: a node that joins your cluster arrives with both directories empty, and a node that leaves is re-imaged. Keep the hook sources on the shared volume and install them onto new nodes.

Per-job GPU accounting is the intended use. Every GPU node runs AMD's ROCm Data Center daemon, rdcd, unauthenticated on localhost:28051, so every rdci call takes -u --host localhost:28051 after its subcommand and needs no certificates. The daemon is read-only: telemetry and job statistics, no power, clock or reset controls. A start hook records against a GPU group you created with rdci group -c <name> -u --host localhost:28051, and a stop hook reports:

# /etc/slurm/prolog.d/50-rdc-stats
rdci stats -s "$SLURM_JOB_ID" -g <group id> -u --host localhost:28051

# /etc/slurm/epilog.d/50-rdc-stats
rdci stats -j "$SLURM_JOB_ID" -u --host localhost:28051 >> "/mnt/shared/jobstats/$SLURM_JOB_ID.txt"
rdci stats -x "$SLURM_JOB_ID" -u --host localhost:28051

Containers

Every compute node runs jobs inside a container image when you pass --container-image; nothing is installed on the cluster for that. The GPUs and the fabric devices of the allocation are visible inside the container.

srun -N1 --gres=gpu:1 --container-image=rocm/pytorch:latest python -c "import torch; print(torch.cuda.is_available())"
flag effect
--container-image=<registry>#<repo>:<tag> or a .sqsh path image to run (registry # separates host and repository)
--container-name=<name> keep the unpacked image on the node between jobs; the first start of a 20 GB image takes minutes, later starts seconds
--container-mounts=/mnt/shared:/mnt/shared,/scratch:/scratch bind host paths into the container; only your home directory is mounted by default, the shared volume and the node-local scratch are not
--container-writable --container-save=<path>.sqsh build an image inside a job and save it to the shared volume for later jobs
--mpi=pmix the launcher for MPI and collective programs across nodes

The first pull of an image is per node. Store images you reuse on the shared volume and reference the .sqsh file.

Storage

path scope what it is
/home login and every compute node your home directory, on the cluster's shared volume; the same files everywhere
/mnt/shared login and every compute node the shared volume itself: datasets, images, checkpoints; sized at cluster creation
/scratch one compute node node-local NVMe scratch, writable by any user; not shared, wiped when the node leaves your cluster; inside a container only with --container-mounts=/scratch:/scratch

Deleting the cluster preserves the shared volume and every byte on it; a new cluster cannot take the same name while the preserved volume exists, so ask the platform team to reattach or destroy it. Node-local scratch is destroyed with the node.

Controllers and failover

Three controllers run the scheduler; one is primary and two are backups, with the scheduler state on the shared volume. If the primary stops, a backup takes over within about 2.5 minutes. Running jobs keep running through the takeover; squeue and new submissions wait until the backup is in control. When the original controller returns it resumes control within a minute, again without touching running jobs.

Nodes

The platform adds and removes compute nodes; you see them in sinfo under their assigned names. A node under maintenance is drained first, so running jobs on it finish, or reach the --time they asked for, before it leaves. There is no API for adding nodes to a Slurm cluster — ask the platform team.

Monitoring

Your cluster's page in the console has a Monitoring tab with one dashboard per view, each scoped to the selected cluster and refreshed every minute. The data comes from the scheduler itself and from an exporter on every compute node.

dashboard what it shows
GPU Nodes per node: GPU utilization, memory, power and temperature; host CPU, memory, disk and network
CPU-Only Nodes the same host view for nodes without GPUs; present only when the cluster has them
Node Network public ingress and egress volume and network errors per node

The dashboards cover the nodes, not the scheduler: queue depth, job states and drain reasons come from squeue, sinfo -R and sacct on the login node, and sacct is the record of finished jobs, including the GPUs each one held. An agent reads the same node states, the queue, a job's record and its log without a login through the MCP server's slurm_read, and a node's telemetry through metrics_read. There are no alerts and no notifications; a pending job with idle GPUs is something you notice, not something the platform tells you.