Cluster pages by token¶
Thirteen read only routes serve what the console's Workloads, job, Cluster Overview, machine, and Storage pages show for a Kubernetes cluster, so an agent reads the platform's own view instead of rebuilding it from kubectl.
| Route | What it serves |
|---|---|
GET /api/k8s/workloads |
the Workloads page: running, pending and recent jobs with their state and waiting verdict, services with their load balancer endpoints and verdict flags, GPU totals, the market block with the price to win |
GET /api/k8s/workloads/runs?before= |
Run History, newest first |
GET /api/k8s/workloads/history?kind=&ns=&wname= |
one workload's scheduling history |
GET /api/k8s/workloads/cost?kind=&ns=&wname= |
one workload's market cost |
GET /api/k8s/workloads/metrics?kind=&ns=&wname=&r= |
one workload's utilization and traffic |
GET /api/k8s/jobs/{ns}~{pod} |
one pod's page: detail, GPU metrics, history, load test results |
GET /api/k8s/jobs/{ns}~{pod}/cost |
one pod's cost panel |
GET /api/k8s/jobs/{ns}~{pod}/network?r= |
one pod's internet traffic |
GET /api/k8s/nodes |
the Cluster Overview board: every node, its GPUs, health, pods, and the platform's verdict word per node (below) |
GET /api/k8s/nodes/{node} |
one node's machine page, with its verdict |
GET /api/k8s/nodes/{node}/network?r= |
one node's network volume |
GET /api/k8s/capacity/history?hours= |
the Cluster Overview charts: GPUs by health state over time, demand against capacity |
GET /api/k8s/storage |
the Storage page: NVMe and HBM per node, shared home and shared volume fullness, for every cluster in your organization |
Contract: GET /api/k8s/openapi.json.
The workload and job routes are also grid_api paths
on the MCP server, same bodies and refusals; the node and storage
routes are not.
All thirteen take an org API token and resolve the cluster the way the
limit price read does: a token bound to a Kubernetes cluster reads its
own; an org wide token resolves when the organization holds one
Kubernetes cluster and otherwise passes ?cluster=<name> or gets a
409 ambiguous-cluster. A token bound to a VM cluster is refused with
422 bad-request. A cluster outside your organization is a 404, the
same as one that does not exist.
What the platform knows that kubectl does not¶
These reads carry the platform's own judgement about your cluster: the market verdict on each waiting request, the price and protection state of each running gang, node health as the platform's checkers see it, cost attributed to a workload or a pod, and history that outlives the cluster's own events. The raw objects stay in kubectl.
Refusals¶
| Code | Status | Meaning |
|---|---|---|
cluster-not-ready |
409 | the cluster is still being built or is not ready; retry |
not-found |
404 | no such cluster, node, or pod in your organization |
bad-request |
422 | a malformed range (r is 1h, 6h, 24h, or 7d), workload identity, pod id, before cursor, or a non numeric hours |
station-unavailable |
503 | the cluster's island did not answer, or its reads are served in a way this token cannot reach yet |
Reading the Workloads page¶
curl -H "Authorization: Bearer $NC_TOKEN" \
https://nationalcompute.com/api/k8s/workloads
running, pending and recent list the jobs with their GPUs, nodes,
state, reason and events; services carries each Service with its
load balancer endpoint; market is the page's market block (the limit price,
each node's price, protection windows, the price to win, the reserved
block); mirror gives the age of each data kind, so a dark island's
last snapshot ages honestly instead of reading fresh.
The waiting verdicts an agent branches on ride this read. When the
market has a verdict for a waiting pod, pending[].reason is one of
these eight strings:
reason |
Meaning | What to do |
|---|---|---|
outbid |
lost on price; win_price_per_gpu_hour carries the price that wins when the market quoted one |
raise the limit price, or wait |
waiting for available supply |
protected nodes hold the supply the gang needs | wait; windows lapse on their own |
cluster too small for this request |
the site holds fewer nodes than the gang needs | a smaller gang |
waiting for capacity; none available at this site right now |
the site is out of stock for the class; a higher limit price cannot help | wait; the platform retries |
waiting for reserved GPUs |
the job is smaller than one node and waits for GPUs of your base load block to free; it never bids on the market; a higher limit price cannot help | wait |
reserved GPUs too few for this request |
the job's pods are each smaller than one node and together ask for more GPUs than one node holds; a job with pods that small runs on one reserved node only; it never bids | whole node pods to run across nodes; fewer pods to fit one node |
bid too low to ever clear |
no bid this low can ever win a node; bid_too_low reads true on the row and its Service |
check the price to win and raise the limit price |
node provisioning |
a grant is inbound | nothing |
The two reserved verdicts appear only on clusters with reservation packing enabled. There a job smaller than one node runs on your base load block. It never bids.
Any other value is scheduler state, never a market verdict: the
scheduler's own reason for a pod the market is not judging
(Unschedulable, ContainerCreating, ImagePullBackOff, a static
cluster's words), the request's condition message verbatim, or the
empty string before the first verdict. Branch on the two Service
booleans and the eight verdicts. Treat everything else as scheduler
state and watch the pod with kubectl.
services[].capacity_unavailable and services[].bid_too_low carry
the two verdicts as booleans per Service. services[].requests[].reason
is the market's full sentence for each request, naming your own bid's
numbers. Each request entry, under a job row or a Service, carries
gpus (the request's GPU total) and sub_node. sub_node reads
true when each pod of the request asks for fewer GPUs than one node.
On a cluster with reservation packing enabled such a request runs only
on your reservation.
market.price_to_win maps a job size in nodes to the price per GPU
hour that wins it right now, and market.protection lists the open
protection windows per node. market.reserved is your base load
block with its nodes and the seat each request holds on them. Its
free_gpus sums the GPUs still open for sub node work across the
block's nodes this tick. Nodes a whole node request holds are
excluded. free_gpus reads null until the platform reports it. A
request smaller than one node starts when one reserved node has its
gpus free and no whole node job is ahead of it.
Verdicts in the cluster¶
The market posts the same verdicts inside the cluster, as Events on the
Job (one per change of reason) and as the reason of the request's
Provisioned condition. The MCP job_watch tool reads them and maps
each to the sentence above. One table covers every spelling:
pending[].reason on the wire |
Event reason on the Job | job_watch field |
|---|---|---|
outbid |
Outbid |
market_events[].verdict reads outbid |
waiting for available supply |
PendingSupply |
market_events[].verdict reads waiting for available supply |
cluster too small for this request |
NoSupply |
market_events[].verdict reads cluster too small for this request |
waiting for capacity; none available at this site right now |
CapacityUnavailable, a Warning Event on the pod; the request itself reads PendingSupply, Outbid or Pending |
market_events[].verdict reads waiting for capacity; none available at this site right now |
bid too low to ever clear |
BidTooLow |
market_events[].verdict reads bid too low to ever clear; standing_bid.bid_too_low |
node provisioning |
NodeGranted, one per granted node |
node_granted_events; market_events[].verdict reads node provisioning |
waiting for reserved GPUs |
ReservedBusy |
market_events[].verdict reads waiting for reserved GPUs |
reserved GPUs too few for this request |
ReservedTooSmall |
market_events[].verdict reads reserved GPUs too few for this request |
| the empty string before the first verdict | Pending |
market_events[].verdict reads null |
billing_hold (an object; no reason sentence) |
BalanceTooLow |
market_events[].verdict reads null; capacity_read carries billing_hold |
preempt_until on the running row |
NodePreempting on the node and its pods; PreemptionRescinded on the node when the notice is withdrawn |
preempting.pods[].preempt_at, preempting.earliest |
the requeued state and requeued_at |
the Kueue Workload's Evicted and Requeued conditions and status.requeueState.count |
workloads[].requeue_count, workloads[].evicted, workloads[].requeued |
the requeued state; requests[].phase reads revoked and its message names the departed nodes |
MarketRevoked, a Warning Event on the Job when every granted node left the cluster |
market_events[].verdict reads requeued |
PreemptionNotice=True (reason MarketPreemption) is a condition on
the ProvisioningRequest. It is not an Event, so job_watch does not
read it; cluster_read does. The two reserved verdicts appear only on
clusters with reservation packing enabled.
job_watch answers market_events from one page of the namespace's
Events, scoped to the watched pods and their owning jobs (name the job
in the selector while it has no pods), at most 30, newest last;
node_granted_events comes off the same page. market_events_truncated
reads true when the page overflowed and precise per reason reads
filled it; a refused Events read answers null with
market_events_note. preempting lists every watched pod carrying the
marketplace.nationalcompute.com/preempt-at annotation with the
earliest deadline. workloads lists the Kueue Workloads behind the
watched pods, narrowed by the pods' Job uid or queue label before the
fetch, at most 20; when none names the watched job as its owner every
fetched Workload answers with workloads_note; it reads null with
workloads_note when the cluster cannot answer.
billing_hold is your organization's billing hold state, or null,
the same object the limit price read
carries. While state reads hold the market places nothing new for
the cluster. At stage notice new placements stop. At enforce
capacity is reclaimed. Pending rows then carry no market verdict until
the hold lifts. Read the hold before you treat an empty reason as
scheduler state.
Run History pages by ?before=<the previous page's last ended stamp>.
A workload's history, cost and metrics take its identity as kind,
ns and wname, the values the Workloads page shows.
Reading a pod, a node, the storage¶
A pod's page takes {namespace}~{pod} as its id. A deleted pod still
answers from the platform's record; a pod unknown to both the cluster
and the record is a 404.
The board and the machine page name nodes the way the console does.
?source=mirror on either answers from the platform's mirror without
dialing the island, the console's first paint; island_stale marks a
dark island (below).
The Storage page is organization wide: one entry per cluster with the
NVMe and HBM readings the console plots (nvme and hbm, each with
its rows and averages), pending where the island's storage telemetry
has not landed, scanned (whether a reading exists) with age (its
age in seconds), the cluster's quota, used, members, node_count,
whether the cluster trades on the market, and shared: the shared
volumes bound to the cluster with provisioned, used, mountpoint
and each node's mounted state (true, false, or null when the
node has not been swept). Sizes are bytes from the platform's scan.
null means not scanned. It is never 0.
Fullness bands¶
The Storage page paints every fullness reading in one of three bands. The body carries the same verdict, so an agent reacts to the band the console shows instead of picking its own thresholds.
| Field | Reading | Where |
|---|---|---|
thresholds |
{warn_pct: 90, crit_pct: 99}, the constants |
top level, once |
nvme.rows[].level |
node scratch: used of total |
each NVMe row |
hbm.rows[].level |
GPU memory: used of total |
each HBM row |
shared[].level |
the shared volume: used of provisioned |
each bound volume |
quota_level |
the shared home: used of quota |
each cluster |
level is crit at or above crit_pct, warn above warn_pct, ok
below, and null when the reading is unknown (an unreachable node, a
volume not yet scanned, a cluster with no quota set). A null band is
never ok. The bands are instantaneous readings of the latest scan.
metrics_read action=flags judges GPU memory over a
window with its own rule (memory_pressure). The two can disagree by
design.
Capacity history¶
GET /api/k8s/capacity/history?hours=24 serves the two charts on the
Cluster Overview: GPUs by health state over time, and demand against
capacity. hours defaults to 24 and clamps to 0.5 .. 8760 (currently
one year).
{"t0": 1759100000, "step": 216, "source": "pg",
"states": [{"label": "healthy", "data": [16, 16, 8]},
{"label": "unschedulable", "data": [0, 0, 8]}],
"total": [16, 16, 16],
"scheduled": [8, 8, 8], "demand": [0, 8, 8], "available": [8, 8, 0]}
t0 is the first bucket in epoch seconds and step the bucket width
in seconds; every array holds one value per bucket. states is one
series per health state the cluster's nodes were in during the window,
GPU weighted (a node counts its GPUs). A state absent from the window
is absent from the list. There is no zero series for it. total
stacks every state per bucket.
states[].label is the platform's health vocabulary, seven words:
label |
Meaning | Schedulable |
|---|---|---|
healthy |
nothing known bad | yes |
degraded |
hardware checks failing while the scheduler still places work | yes |
unschedulable |
the scheduler places nothing: NotReady, a platform cordon, a node in lifecycle error | no |
cordoned |
parked by your own cluster admin; Ready and fault free; still billed | no |
moving |
an operation owns the node; recorded without a cluster, so it rarely appears on a cluster's own series | no |
unreachable |
the control plane lost the node; outranks unschedulable |
no |
unknown |
the platform cannot say: the cluster snapshot was missing or the node was absent from it | no |
Schedulable capacity is healthy plus degraded. The console's GPU
Health chart merges those two into its healthy band and draws
unschedulable as unhealthy; cordoned, unreachable, unknown and
moving get their own bands only when they occurred in the window.
scheduled (GPUs held by running work), demand (GPUs asked for by
waiting work) and available (schedulable capacity minus scheduled,
floored at 0) are the demand overlay. A null bucket inside them means
no sampler beat was in reach: unknown, never 0. The three keys are
absent when the demand record is unreadable; the health series still
answers. source is pg for the record store and live for a single
current point when the store is off; error then names the
degradation.
A cluster whose reads the platform serves centrally answers 503
station-unavailable here, the same as the board and the machine page.
Node verdicts¶
Every row of the board and the node object of the machine page carry
verdict, the platform's own health word for the node. It is the word
the console tile shows, derived on the server from the same inputs the
payload carries (k8snodes, the lifecycle status, the unreachable
mark). The object is
{sev, word, detail, healthy, schedulable, lost, observed}. sev is
ok, warn or err. detail is the hover text, "" when there is
none. healthy and schedulable are the board's section split:
healthy nodes land in Healthy, unhealthy nodes that still schedule in
"Unhealthy, schedulable", the rest in Unschedulable, and lost nodes
in Unreachable. schedulable is Kubernetes truth (Ready and not
cordoned) whatever the word says. observed is false when the
scheduler feed behind the verdict did not answer on this read
(below). verdict is null only when the
platform could not derive one; the board still answers.
word |
sev |
Meaning | What to do |
|---|---|---|---|
"" |
ok |
Ready, schedulable, no platform marker | nothing |
unobserved |
warn |
the scheduler feed did not answer on this read (the island is stale, the feed is absent or erroring, or it carries no rows for the cluster); nothing judged the node, and detail names the cause |
read island_stale and feed_error; retry |
unreachable |
err |
the control plane lost the node: the kubelet has been silent for more than 300 seconds, lost is true, and detail carries the last known state |
wait; the platform already sees it |
NotReady |
err |
the node is not Ready, within the 300 second threshold | wait; a kubelet restart clears it |
Node Fix In Progress |
warn |
our health checks took the node out of service: cordoned with a platform marker; billing for it stopped; a replacement is being provisioned (mechanism) | nothing; the request row reads Replacement Pending until the replacement joins |
unhealthy |
err |
the health check daemon reports failing checks, named in detail; the node still schedules unless it is also cordoned |
move work off the node, or cordon it |
tainted |
warn |
schedulable, with a taint the platform's health checker placed | as unhealthy |
cordoned |
warn |
cordoned without a platform marker: your own cordon; detail lists the taints |
uncordon when you are done |
missing |
err |
the cluster answered and this node is absent from its node list | check kubectl get nodes |
creating, joining, error |
warn; err for error |
the platform's lifecycle status while the node is not ready. Lifecycle is never health | wait |
A node mid move leaves the list: the board drops it and its machine
read answers 404, so no moving word rides the wire. A node with a
scheduler row is judged from that row. A node with no row on a board
that lists other nodes is missing. A node with no row on a feed that
did not answer is unobserved, unless the platform's own record marks
it unreachable or its lifecycle status is not ready; those words
stand, with observed: false.
Island staleness¶
Two fields tell a dark island from a broken node. island_stale on
the board, the machine page and GET /api/k8s/cluster
is null while the platform hears the island's heartbeat and {since}
once it has lost the island. Every verdict then reads
observed: false: a node the dark feed would have read as healthy
reads unobserved, a node with a word of its own keeps it. feed_error
on GET /api/k8s/cluster is the last error of the platform's pull of
the node health feed, null when the last pull landed. feed_age_s is
the data age of what that feed serves, in seconds: the time since the
newest landed pull plus the platform's cache age at that pull, null
when no pull has ever landed. A failing pull leaves the age growing;
it never resets. k8snodes.error on the board names a node feed the
island could not answer on this read, and its rows then read
observed: false. Read these before acting on a red verdict.
On the MCP server the same two bodies are grid_api get /api/k8s/nodes
(the board) and grid_api get /api/k8s/nodes/<node> (one machine, the
REST spelling; the literal /api/k8s/nodes/{node} with node in
params is the same call), with params cluster and source
(mirror for the platform's mirror, no island dial). A malformed node
name is 422 bad-request, an unknown one 404 not-found.