# Slurm clusters

A Slurm tenancy is a **dedicated cluster**: a login node you ssh into,
a controller triplet the platform runs, and GPU compute nodes the
platform moves in and out of your cluster. You submit jobs with the
standard Slurm commands (`srun`, `sbatch`, `squeue`, `sacct`); nobody
runs them for you. Slurm 26.05 with the enroot and pyxis container
plugins is installed on every node.

## Connecting

Every Slurm cluster has one login account, `tenant`, with root
privileges through `sudo`. It is a member of the GPU device groups
(`video`, `render`) on every node, so GPU tools and runtimes work
without `sudo`. The account accepts the SSH public keys the platform
team registered for your organization; there is no password and no
sign-in flow.

```sh
ssh tenant@<login-address>
```

The login address is on your cluster's page in the console under
Connect. To add or remove a key, send the platform team the public key
line (`ssh-ed25519 AAAA…`); the change reaches the login node within a
minute of being registered. Keys you append to `~/.ssh/authorized_keys`
yourself work on the compute nodes as well, because the home directory
is shared across the cluster (see [Storage](#storage)) — the login's
managed set is the one the platform keeps.

The setup script writes an `ssh <cluster>` alias for the account when
the cluster uses key login:

```sh
curl -fsSL https://access.nationalcompute.com/setup.sh | sh -s -- you@example.com <cluster>
```

## Jobs

Jobs run in the `main` partition. Request GPUs with `--gres`; each
requested GPU brings a default CPU allocation sized for collective
libraries, so a whole-node job needs no CPU flags:

```sh
srun -N1 --gres=gpu:8 hostname
sbatch --nodes=4 --ntasks-per-node=8 --gres=gpu:8 train.sbatch
```

| limit | value |
|---|---|
| maximum wall time (`--time`) | none — a job runs until it finishes or reaches the `--time` it asked for |
| default wall time when `--time` is absent | unlimited |
| default CPUs per requested GPU | 28 (currently; sized so a `gpu:8` job holds the node) |
| GPUs per compute node | 8 |

A job that requests fewer GPUs than a node has shares the node with
other jobs; `ROCR_VISIBLE_DEVICES` inside the job lists the GPUs it
holds, and the others are not visible to it. The whole GPU-node memory
is allocatable to a full-node job.

Accounting is per cluster: every job is charged to the cluster's single
account, and `sacct` shows it with `gres/gpu` in `AllocTRES`. There are
no per-user limits, quotas or fair-share weights.

## Job hooks

Every compute node runs the executables in two directories around each
job, as root, with Slurm's job environment (`SLURM_JOB_ID`,
`SLURM_JOB_USER`, `SLURM_JOB_GPUS`):

| directory | when | time limit per hook |
|---|---|---|
| `/etc/slurm/prolog.d/` | before the job's first step, after the platform's health gate has passed | 20 s |
| `/etc/slurm/epilog.d/` | after the job's last step, including cancelled and timed-out jobs | 60 s |

Hooks run in name order. A hook that exits non-zero or reaches its time
limit is logged to the node's syslog (tags `slurm-prolog` and
`slurm-epilog`) and otherwise ignored: it cannot fail the job and cannot
drain the node. Files without the executable bit are skipped.

The directories are yours; write to them with `sudo` on each node. They
are per-node state: a node that joins your cluster arrives with both
directories empty, and a node that leaves is re-imaged. Keep the hook
sources on the shared volume and install them onto new nodes.

Per-job GPU accounting is the intended use. Every GPU node runs AMD's
ROCm Data Center daemon, `rdcd`, unauthenticated on `localhost:28051`,
so every `rdci` call takes `-u --host localhost:28051` after its subcommand and needs no certificates. The daemon is read-only: telemetry
and job statistics, no power, clock or reset controls. A start hook
records against a GPU group you created with `rdci group -c <name> -u --host localhost:28051`, and a
stop hook reports:

```sh
# /etc/slurm/prolog.d/50-rdc-stats
rdci stats -s "$SLURM_JOB_ID" -g <group id> -u --host localhost:28051

# /etc/slurm/epilog.d/50-rdc-stats
rdci stats -j "$SLURM_JOB_ID" -u --host localhost:28051 >> "/mnt/shared/jobstats/$SLURM_JOB_ID.txt"
rdci stats -x "$SLURM_JOB_ID" -u --host localhost:28051
```

## Containers

Every compute node runs jobs inside a container image when you pass
`--container-image`; nothing is installed on the cluster for that. The
GPUs and the fabric devices of the allocation are visible inside the
container.

```sh
srun -N1 --gres=gpu:1 --container-image=rocm/pytorch:latest python -c "import torch; print(torch.cuda.is_available())"
```

| flag | effect |
|---|---|
| `--container-image=<registry>#<repo>:<tag>` or a `.sqsh` path | image to run (registry `#` separates host and repository) |
| `--container-name=<name>` | keep the unpacked image on the node between jobs; the first start of a 20 GB image takes minutes, later starts seconds |
| `--container-mounts=/mnt/shared:/mnt/shared,/scratch:/scratch` | bind host paths into the container; only your home directory is mounted by default, the shared volume and the node-local scratch are not |
| `--container-writable --container-save=<path>.sqsh` | build an image inside a job and save it to the shared volume for later jobs |
| `--mpi=pmix` | the launcher for MPI and collective programs across nodes |

The first pull of an image is per node. Store images you reuse on the
shared volume and reference the `.sqsh` file.

## Storage

| path | scope | what it is |
|---|---|---|
| `/home` | login and every compute node | your home directory, on the cluster's shared volume; the same files everywhere |
| `/mnt/shared` | login and every compute node | the shared volume itself: datasets, images, checkpoints; sized at cluster creation |
| `/scratch` | one compute node | node-local NVMe scratch, writable by any user; not shared, wiped when the node leaves your cluster; inside a container only with `--container-mounts=/scratch:/scratch` |

Deleting the cluster preserves the shared volume and every byte on it;
a new cluster cannot take the same name while the preserved volume
exists, so ask the platform team to reattach or destroy it. Node-local
scratch is destroyed with the node.

## Controllers and failover

Three controllers run the scheduler; one is primary and two are
backups, with the scheduler state on the shared volume. If the primary
stops, a backup takes over within about 2.5 minutes. Running jobs keep
running through the takeover; `squeue` and new submissions wait until
the backup is in control. When the original controller returns it
resumes control within a minute, again without touching running jobs.

## Nodes

The platform adds and removes compute nodes; you see them in `sinfo`
under their assigned names. A node under maintenance is drained first,
so running jobs on it finish, or reach the `--time` they asked for, before
it leaves. There is no API for adding nodes to a Slurm cluster — ask the
platform team.

## Monitoring

Your cluster's page in the console has a Monitoring tab with one
dashboard per view, each scoped to the selected cluster and refreshed
every minute. The data comes from the scheduler itself and from an
exporter on every compute node.

| dashboard | what it shows |
|---|---|
| GPU Nodes | per node: GPU utilization, memory, power and temperature; host CPU, memory, disk and network |
| CPU-Only Nodes | the same host view for nodes without GPUs; present only when the cluster has them |
| Node Network | public ingress and egress volume and network errors per node |

The dashboards cover the nodes, not the scheduler: queue depth, job
states and drain reasons come from `squeue`, `sinfo -R` and `sacct` on
the login node, and `sacct` is the record of finished jobs, including the
GPUs each one held. An agent reads the same node states, the queue, a
job's record and its log without a login through the MCP server's
[`slurm_read`](https://docs.nationalcompute.com/api/slurm.md), and a node's telemetry through
[`metrics_read`](https://docs.nationalcompute.com/api/metrics.md). There are no alerts and no notifications; a pending
job with idle GPUs is something you notice, not something the platform
tells you.
