Skip to content

Slurm reads

slurm_read answers what the console shows for a Slurm cluster: the status page (node states, GPU allocation, the live queue, GPU hours, shared /home usage), one job's page, that job's place in the queue, and its log. The tool lives on the MCP server. It is read only and single phase. No confirm_token is involved. Nothing on this surface submits, cancels or changes a job: sbatch, scancel and every other scheduler command stay on the login node, where your agent runs them over ssh with the connection facts grid_whoami carries (see Connection facts).

grid_whoami ──▶ the cluster's name, kind slurm, ssh facts
     │
     ▼
slurm_read status ──▶ nodes, queue, GPU hours, /home
slurm_read job    ──▶ one job's record, queue context, GPU metrics
slurm_read queue  ──▶ position and the jobs ahead
slurm_read log    ──▶ stdout tail, or a grep over the whole file
     │
     ▼
metrics_read node ──▶ a Slurm node's GPU and host telemetry

Every read is scoped to your organization. A cluster name that is not one of your organization's Slurm clusters answers 404 not-found, whether it belongs to someone else, is a Kubernetes or VM cluster, or never existed.

Arguments

Argument Values Applies to
cluster the cluster name (grid_whoami lists them) required on every call
action status, job, queue, log required on every call
jobid a Slurm job id: digits, optionally _<task> required for job, queue, log
source live (default) or mirror status; mirror answers from the platform's last snapshot without dialing the cluster's island
tail_lines 1 to 1000; defaults to 200 log; the lines kept from the end
query a fixed string, matched case insensitively log; greps the whole file instead of tailing
file an index into the answer's files log; a per task log next to the batch stdout; the default is the batch stdout
incarnation the coverage.job.incarnation of a job answer log on a centrally read island

Each argument belongs to the actions its row names. An argument present on another action refuses 422 bad-request; nothing is silently ignored. file and tail_lines must be whole numbers. A blank query tails.

action=status

{
  "age": 3,
  "error": null,
  "nodes": {
    "node-1": {"state": "mixed", "flags": [], "gpus_alloc": 4, "gpus_total": 8,
               "jobs": [{"id": "4242", "user": "ada", "gpus": 4}]},
    "node-2": {"state": "drained", "flags": ["drain"], "gpus_alloc": 0,
               "gpus_total": 8, "jobs": []}
  },
  "queue": [{"id": "4243", "name": "train", "user": "ada", "state": "pending",
             "gpus": 8, "reason": "Resources", "time": "0:00"}],
  "usage": {"month": "2026-09", "users": {"ada": 12.5}, "error": null},
  "storage": {"fs": {"used": 10995116277760, "total": 109951162777600,
                     "free": 98956046500000},
              "dirs": {"ada": 7516192768}, "error": null}
}
Field What it carries
age the age of the scheduler snapshot, seconds
error the scheduler feed's error, or null
nodes one entry per compute node, keyed by the name sinfo shows: state, flags, gpus_alloc, gpus_total, and jobs (the jobs holding GPUs on it)
queue squeue, live: id, name, user, state, gpus, and reason while pending or time while running
usage GPU hours this month per user from the scheduler's accounting, with month; error when accounting did not answer
storage the shared /home export: fs (used, total, free, bytes), dirs (bytes per member of this cluster), error when the scan did not run
mirror true when source=mirror answered

usage is the scheduler's accounting. The billing records are the authority on what a job cost. storage.dirs lists your organization's members only.

Node states

A node takes no new jobs when its state is one of down, drained, draining, error, fail, failing, unknown, invalid, inval, or when flags carries drain, not_responding, invalid_reg or fail. The console paints those red. idle means free, mixed and allocated mean busy, completing means a job is finishing on it. Every other word is Slurm's own state, verbatim.

action=job

{
  "job": {
    "id": "4242", "name": "train", "user": "ada", "state": "pending",
    "reason": "Resources", "nodes": 2, "nodelist": "", "gpus": 16, "cpus": 448,
    "mem": "1500G", "time_limit": "1-00:00:00", "elapsed": "00:00:00",
    "submitted": "2026-09-29T08:10:00", "started": "", "est_start": "2026-09-29T09:00:00",
    "ended": "", "workdir": "/home/ada/train", "stdout": "/home/ada/train/slurm-4242.out",
    "exit_code": "0:0", "partition": "main", "account": "ada", "qos": "normal",
    "gpu_alloc": {}, "source": "scontrol"
  },
  "queue": {"position": 2, "total": 5, "ahead": [{"id": "4240", "reason": "Priority"}],
            "ahead_total": 1},
  "metrics": {"gpus": [], "series": null, "error": null}
}

job is the scheduler's record, scontrol while the job is live and sacct once accounting holds it (source says which). started is empty until the job runs; est_start carries the scheduler's estimate while it waits; ended is empty until it ends. queue is present only while state is pending and is null otherwise. metrics is the node scope GPU telemetry over the job's window; its error reads unavailable when the metrics store did not answer, never a reason with internals. A job accounting no longer knows answers 404 not-found.

action=queue

{"jobid": "4242", "state": "pending",
 "queue": {"position": 2, "total": 5, "ahead": [{"id": "4240", "reason": "Priority"}],
           "ahead_total": 1}}

position counts from 1 among pending jobs in scheduler order; total is the pending count; ahead lists the nearest jobs ahead (at most 5) with their wait reasons; ahead_total is the real count ahead. queue is null when the job is not pending. This action reads no metrics.

action=log

{
  "path": "/home/ada/train/slurm-4242.out",
  "lines": ["step 1200 loss 2.31", "step 1210 loss 2.29"],
  "truncated": true,
  "clipped": true,
  "files": ["log-train_4242_0.out", "log-train_4242_1.out"],
  "selected": null,
  "note": null
}

The tail reads the batch stdout off the shared /home export. lines carries the last tail_lines lines; truncated is true when the file held more than the platform's own tail window (1000 lines or 512 KiB); clipped is true when tail_lines dropped lines from that window. files lists the per task logs written next to the stdout by wrapper stacks (srun redirections named *_<jobid>_<rank>.out), and file selects one by index; selected echoes it. note explains an empty or missing log ("log file is empty so far", a log on node local storage).

With query the answer is a grep over the whole file: lines are the matching lines (the last 500 matches at most), lnos their line numbers, matches the count of lines returned, filtered reads true. When tail_lines dropped matches, clipped reads true and total_matches keeps the file wide count. A query longer than 200 characters is cut to 200.

Centrally read islands

On an island whose reads the platform serves from its own record instead of dialing the island, every answer carries coverage: per source, availability (complete, partial, unavailable), freshness, observed_at, and a reason when something is missing. The status answer is then the platform's inventory (inventory.nodes with state, gpus_total, gpus_alloc, cpus_alloc, plus running_count and queue_count); jobs, queue, usage and storage read null with their reasons. The job answer carries coverage.job.incarnation, which log needs as incarnation. A log answer there carries first_record and last_record, the record numbers the returned lines span. A record that is not complete refuses with 503 station-unavailable carrying code (the read model's refusal, central_read_unavailable), reason (the coverage's own reason, for example job_incarnation_required or central_store_disabled) and coverage. The message says retry only when the reason is one a later call can clear.

Refusals

Code Status Meaning
bad-request 422 an unknown action or source, a missing jobid, a malformed job id, an argument outside its action, a fractional file or tail_lines, file outside 0 to 127, incarnation on an island that is not centrally read
not-found 404 no Slurm cluster of that name in your organization; a job accounting does not know; per job monitoring not enabled for the cluster (job, queue, log)
cluster-not-ready 409 the cluster is still being built or is not ready; retry
station-unavailable 503 the cluster's island did not answer, or a centrally kept record is not complete (code, reason and coverage say why)

Connection facts

grid_whoami lists every cluster your organization holds. A slurm entry adds the login endpoint the console's Connect step shows:

Field What it carries
ssh_host the login address
ssh_port the ssh port
ssh_user the shared login account, or null when the login is your own OIDC identity
ssh_auth key (keys the platform registered for your organization) or opk (a browser sign in through the setup script)

All four read null while the endpoint is unresolved. Kubernetes and VM entries carry none of these keys. GET /api/whoami by org API token answers the same body.

Telemetry

metrics_read accepts a Slurm cluster as cluster for action=node and action=flags with node; the node name is the key status.nodes uses. Workload targets (kind, name) are Kubernetes only and refuse 422 bad-request on a Slurm cluster. When your organization holds a Kubernetes cluster and a Slurm cluster of the same name, metrics_read reads the Kubernetes one.