# Cluster pages by token

Thirteen read only routes serve what the console's Workloads, job,
Cluster Overview, machine, and Storage pages show for a Kubernetes
cluster, so an agent reads the platform's own view instead of rebuilding
it from kubectl.

| Route | What it serves |
|---|---|
| `GET /api/k8s/workloads` | the Workloads page: running, pending and recent jobs with their state and waiting verdict, services with their load balancer endpoints and verdict flags, GPU totals, the market block with the price to win |
| `GET /api/k8s/workloads/runs?before=` | Run History, newest first |
| `GET /api/k8s/workloads/history?kind=&ns=&wname=` | one workload's scheduling history |
| `GET /api/k8s/workloads/cost?kind=&ns=&wname=` | one workload's market cost |
| `GET /api/k8s/workloads/metrics?kind=&ns=&wname=&r=` | one workload's utilization and traffic |
| `GET /api/k8s/jobs/{ns}~{pod}` | one pod's page: detail, GPU metrics, history, load test results |
| `GET /api/k8s/jobs/{ns}~{pod}/cost` | one pod's cost panel |
| `GET /api/k8s/jobs/{ns}~{pod}/network?r=` | one pod's internet traffic |
| `GET /api/k8s/nodes` | the Cluster Overview board: every node, its GPUs, health, pods, and the platform's `verdict` word per node ([below](#node-verdicts)) |
| `GET /api/k8s/nodes/{node}` | one node's machine page, with its `verdict` |
| `GET /api/k8s/nodes/{node}/network?r=` | one node's network volume |
| `GET /api/k8s/capacity/history?hours=` | the Cluster Overview charts: GPUs by health state over time, demand against capacity |
| `GET /api/k8s/storage` | the Storage page: NVMe and HBM per node, shared home and shared volume fullness, for every cluster in your organization |

Contract: [`GET /api/k8s/openapi.json`](https://nationalcompute.com/api/k8s/openapi.json).
The workload and job routes are also [`grid_api` paths](https://docs.nationalcompute.com/api/mcp.md#grid_api-paths)
on the MCP server, same bodies and refusals; the node and storage
routes are not.

All thirteen take an org API token and resolve the cluster the way the
[limit price read](https://docs.nationalcompute.com/api/k8s-bid.md) does: a token bound to a Kubernetes cluster reads its
own; an org wide token resolves when the organization holds one
Kubernetes cluster and otherwise passes `?cluster=<name>` or gets a
`409 ambiguous-cluster`. A token bound to a VM cluster is refused with
`422 bad-request`. A cluster outside your organization is a `404`, the
same as one that does not exist.

## What the platform knows that kubectl does not

These reads carry the platform's own judgement about your cluster:
the market verdict on each waiting request, the price and protection
state of each running gang, node health as the platform's checkers
see it, cost attributed to a workload or a pod, and history that
outlives the cluster's own events. The raw objects stay in kubectl.

## Refusals

| Code | Status | Meaning |
|---|---|---|
| `cluster-not-ready` | 409 | the cluster is still being built or is not ready; retry |
| `not-found` | 404 | no such cluster, node, or pod in your organization |
| `bad-request` | 422 | a malformed range (`r` is `1h`, `6h`, `24h`, or `7d`), workload identity, pod id, `before` cursor, or a non numeric `hours` |
| `station-unavailable` | 503 | the cluster's island did not answer, or its reads are served in a way this token cannot reach yet |

## Reading the Workloads page

```sh
curl -H "Authorization: Bearer $NC_TOKEN" \
  https://nationalcompute.com/api/k8s/workloads
```

`running`, `pending` and `recent` list the jobs with their GPUs, nodes,
state, reason and events; `services` carries each Service with its
load balancer endpoint; `market` is the page's market block (the limit price,
each node's price, protection windows, the price to win, the reserved
block); `mirror` gives the age of each data kind, so a dark island's
last snapshot ages honestly instead of reading fresh.

The waiting verdicts an agent branches on ride this read. When the
market has a verdict for a waiting pod, `pending[].reason` is one of
these eight strings:

| `reason` | Meaning | What to do |
|---|---|---|
| `outbid` | lost on price; `win_price_per_gpu_hour` carries the price that wins when the market quoted one | raise the limit price, or wait |
| `waiting for available supply` | protected nodes hold the supply the gang needs | wait; windows lapse on their own |
| `cluster too small for this request` | the site holds fewer nodes than the gang needs | a smaller gang |
| `waiting for capacity; none available at this site right now` | the site is out of stock for the class; a higher limit price cannot help | wait; the platform retries |
| `waiting for reserved GPUs` | the job is smaller than one node and waits for GPUs of your base load block to free; it never bids on the market; a higher limit price cannot help | wait |
| `reserved GPUs too few for this request` | the job's pods are each smaller than one node and together ask for more GPUs than one node holds; a job with pods that small runs on one reserved node only; it never bids | whole node pods to run across nodes; fewer pods to fit one node |
| `bid too low to ever clear` | no bid this low can ever win a node; `bid_too_low` reads `true` on the row and its Service | check the price to win and raise the limit price |
| `node provisioning` | a grant is inbound | nothing |

The two reserved verdicts appear only on clusters with reservation
packing enabled. There a job smaller than one node runs on your
[base load block](https://docs.nationalcompute.com/market.md#base-load-capacity). It never bids.

Any other value is scheduler state, never a market verdict: the
scheduler's own reason for a pod the market is not judging
(`Unschedulable`, `ContainerCreating`, `ImagePullBackOff`, a static
cluster's words), the request's condition message verbatim, or the
empty string before the first verdict. Branch on the two Service
booleans and the eight verdicts. Treat everything else as scheduler
state and watch the pod with kubectl.

`services[].capacity_unavailable` and `services[].bid_too_low` carry
the two verdicts as booleans per Service. `services[].requests[].reason`
is the market's full sentence for each request, naming your own bid's
numbers. Each request entry, under a job row or a Service, carries
`gpus` (the request's GPU total) and `sub_node`. `sub_node` reads
`true` when each pod of the request asks for fewer GPUs than one node.
On a cluster with reservation packing enabled such a request runs only
on your reservation.
`market.price_to_win` maps a job size in nodes to the price per GPU
hour that wins it right now, and `market.protection` lists the open
protection windows per node. `market.reserved` is your base load
block with its nodes and the seat each request holds on them. Its
`free_gpus` sums the GPUs still open for sub node work across the
block's nodes this tick. Nodes a whole node request holds are
excluded. `free_gpus` reads `null` until the platform reports it. A
request smaller than one node starts when one reserved node has its
`gpus` free and no whole node job is ahead of it.

## Verdicts in the cluster

The market posts the same verdicts inside the cluster, as Events on the
Job (one per change of reason) and as the reason of the request's
`Provisioned` condition. The MCP `job_watch` tool reads them and maps
each to the sentence above. One table covers every spelling:

| `pending[].reason` on the wire | Event reason on the Job | `job_watch` field |
|---|---|---|
| `outbid` | `Outbid` | `market_events[].verdict` reads `outbid` |
| `waiting for available supply` | `PendingSupply` | `market_events[].verdict` reads `waiting for available supply` |
| `cluster too small for this request` | `NoSupply` | `market_events[].verdict` reads `cluster too small for this request` |
| `waiting for capacity; none available at this site right now` | `CapacityUnavailable`, a Warning Event on the pod; the request itself reads `PendingSupply`, `Outbid` or `Pending` | `market_events[].verdict` reads `waiting for capacity; none available at this site right now` |
| `bid too low to ever clear` | `BidTooLow` | `market_events[].verdict` reads `bid too low to ever clear`; `standing_bid.bid_too_low` |
| `node provisioning` | `NodeGranted`, one per granted node | `node_granted_events`; `market_events[].verdict` reads `node provisioning` |
| `waiting for reserved GPUs` | `ReservedBusy` | `market_events[].verdict` reads `waiting for reserved GPUs` |
| `reserved GPUs too few for this request` | `ReservedTooSmall` | `market_events[].verdict` reads `reserved GPUs too few for this request` |
| the empty string before the first verdict | `Pending` | `market_events[].verdict` reads `null` |
| `billing_hold` (an object; no reason sentence) | `BalanceTooLow` | `market_events[].verdict` reads `null`; `capacity_read` carries `billing_hold` |
| `preempt_until` on the running row | `NodePreempting` on the node and its pods; `PreemptionRescinded` on the node when the notice is withdrawn | `preempting.pods[].preempt_at`, `preempting.earliest` |
| the `requeued` state and `requeued_at` | the Kueue Workload's `Evicted` and `Requeued` conditions and `status.requeueState.count` | `workloads[].requeue_count`, `workloads[].evicted`, `workloads[].requeued` |
| the `requeued` state; `requests[].phase` reads `revoked` and its message names the departed nodes | `MarketRevoked`, a Warning Event on the Job when every granted node left the cluster | `market_events[].verdict` reads `requeued` |

`PreemptionNotice=True` (reason `MarketPreemption`) is a condition on
the ProvisioningRequest. It is not an Event, so `job_watch` does not
read it; `cluster_read` does. The two reserved verdicts appear only on
clusters with reservation packing enabled.

`job_watch` answers `market_events` from one page of the namespace's
Events, scoped to the watched pods and their owning jobs (name the job
in the selector while it has no pods), at most 30, newest last;
`node_granted_events` comes off the same page. `market_events_truncated`
reads `true` when the page overflowed and precise per reason reads
filled it; a refused Events read answers `null` with
`market_events_note`. `preempting` lists every watched pod carrying the
`marketplace.nationalcompute.com/preempt-at` annotation with the
earliest deadline. `workloads` lists the Kueue Workloads behind the
watched pods, narrowed by the pods' Job uid or queue label before the
fetch, at most 20; when none names the watched job as its owner every
fetched Workload answers with `workloads_note`; it reads `null` with
`workloads_note` when the cluster cannot answer.

`billing_hold` is your organization's billing hold state, or `null`,
the same object the [limit price read](https://docs.nationalcompute.com/api/k8s-bid.md#setting-the-bid)
carries. While `state` reads `hold` the market places nothing new for
the cluster. At `stage` `notice` new placements stop. At `enforce`
capacity is reclaimed. Pending rows then carry no market verdict until
the hold lifts. Read the hold before you treat an empty `reason` as
scheduler state.

Run History pages by `?before=<the previous page's last ended stamp>`.
A workload's history, cost and metrics take its identity as `kind`,
`ns` and `wname`, the values the Workloads page shows.

## Reading a pod, a node, the storage

A pod's page takes `{namespace}~{pod}` as its id. A deleted pod still
answers from the platform's record; a pod unknown to both the cluster
and the record is a `404`.

The board and the machine page name nodes the way the console does.
`?source=mirror` on either answers from the platform's mirror without
dialing the island, the console's first paint; `island_stale` marks a
dark island ([below](#island-staleness)).

The Storage page is organization wide: one entry per cluster with the
NVMe and HBM readings the console plots (`nvme` and `hbm`, each with
its rows and averages), `pending` where the island's storage telemetry
has not landed, `scanned` (whether a reading exists) with `age` (its
age in seconds), the cluster's `quota`, `used`, `members`, `node_count`,
whether the cluster trades on the market, and `shared`: the shared
volumes bound to the cluster with `provisioned`, `used`, `mountpoint`
and each node's `mounted` state (`true`, `false`, or `null` when the
node has not been swept). Sizes are bytes from the platform's scan.
`null` means not scanned. It is never 0.

### Fullness bands

The Storage page paints every fullness reading in one of three bands.
The body carries the same verdict, so an agent reacts to the band the
console shows instead of picking its own thresholds.

| Field | Reading | Where |
|---|---|---|
| `thresholds` | `{warn_pct: 90, crit_pct: 99}`, the constants | top level, once |
| `nvme.rows[].level` | node scratch: `used` of `total` | each NVMe row |
| `hbm.rows[].level` | GPU memory: `used` of `total` | each HBM row |
| `shared[].level` | the shared volume: `used` of `provisioned` | each bound volume |
| `quota_level` | the shared home: `used` of `quota` | each cluster |

`level` is `crit` at or above `crit_pct`, `warn` above `warn_pct`, `ok`
below, and `null` when the reading is unknown (an unreachable node, a
volume not yet scanned, a cluster with no quota set). A `null` band is
never `ok`. The bands are instantaneous readings of the latest scan.
[`metrics_read`](https://docs.nationalcompute.com/api/metrics.md) `action=flags` judges GPU memory over a
window with its own rule (`memory_pressure`). The two can disagree by
design.

## Capacity history

`GET /api/k8s/capacity/history?hours=24` serves the two charts on the
Cluster Overview: GPUs by health state over time, and demand against
capacity. `hours` defaults to 24 and clamps to 0.5 .. 8760 (currently
one year).

```json
{"t0": 1759100000, "step": 216, "source": "pg",
 "states": [{"label": "healthy", "data": [16, 16, 8]},
            {"label": "unschedulable", "data": [0, 0, 8]}],
 "total": [16, 16, 16],
 "scheduled": [8, 8, 8], "demand": [0, 8, 8], "available": [8, 8, 0]}
```

`t0` is the first bucket in epoch seconds and `step` the bucket width
in seconds; every array holds one value per bucket. `states` is one
series per health state the cluster's nodes were in during the window,
GPU weighted (a node counts its GPUs). A state absent from the window
is absent from the list. There is no zero series for it. `total`
stacks every state per bucket.

`states[].label` is the platform's health vocabulary, seven words:

| `label` | Meaning | Schedulable |
|---|---|---|
| `healthy` | nothing known bad | yes |
| `degraded` | hardware checks failing while the scheduler still places work | yes |
| `unschedulable` | the scheduler places nothing: NotReady, a platform cordon, a node in lifecycle error | no |
| `cordoned` | parked by your own cluster admin; Ready and fault free; still billed | no |
| `moving` | an operation owns the node; recorded without a cluster, so it rarely appears on a cluster's own series | no |
| `unreachable` | the control plane lost the node; outranks `unschedulable` | no |
| `unknown` | the platform cannot say: the cluster snapshot was missing or the node was absent from it | no |

Schedulable capacity is `healthy` plus `degraded`. The console's GPU
Health chart merges those two into its healthy band and draws
`unschedulable` as unhealthy; `cordoned`, `unreachable`, `unknown` and
`moving` get their own bands only when they occurred in the window.

`scheduled` (GPUs held by running work), `demand` (GPUs asked for by
waiting work) and `available` (schedulable capacity minus `scheduled`,
floored at 0) are the demand overlay. A `null` bucket inside them means
no sampler beat was in reach: unknown, never 0. The three keys are
absent when the demand record is unreadable; the health series still
answers. `source` is `pg` for the record store and `live` for a single
current point when the store is off; `error` then names the
degradation.

A cluster whose reads the platform serves centrally answers `503
station-unavailable` here, the same as the board and the machine page.

## Node verdicts

Every row of the board and the `node` object of the machine page carry
`verdict`, the platform's own health word for the node. It is the word
the console tile shows, derived on the server from the same inputs the
payload carries (`k8snodes`, the lifecycle `status`, the `unreachable`
mark). The object is
`{sev, word, detail, healthy, schedulable, lost, observed}`. `sev` is
`ok`, `warn` or `err`. `detail` is the hover text, `""` when there is
none. `healthy` and `schedulable` are the board's section split:
healthy nodes land in Healthy, unhealthy nodes that still schedule in
"Unhealthy, schedulable", the rest in Unschedulable, and `lost` nodes
in Unreachable. `schedulable` is Kubernetes truth (Ready and not
cordoned) whatever the word says. `observed` is `false` when the
scheduler feed behind the verdict did not answer on this read
([below](#island-staleness)). `verdict` is `null` only when the
platform could not derive one; the board still answers.

| `word` | `sev` | Meaning | What to do |
|---|---|---|---|
| `""` | `ok` | Ready, schedulable, no platform marker | nothing |
| `unobserved` | `warn` | the scheduler feed did not answer on this read (the island is stale, the feed is absent or erroring, or it carries no rows for the cluster); nothing judged the node, and `detail` names the cause | read `island_stale` and `feed_error`; retry |
| `unreachable` | `err` | the control plane lost the node: the kubelet has been silent for more than 300 seconds, `lost` is `true`, and `detail` carries the last known state | wait; the platform already sees it |
| `NotReady` | `err` | the node is not Ready, within the 300 second threshold | wait; a kubelet restart clears it |
| `Node Fix In Progress` | `warn` | our health checks took the node out of service: cordoned with a platform marker; billing for it stopped; a replacement is being provisioned ([mechanism](https://docs.nationalcompute.com/kubernetes.md#when-we-take-a-node-out-of-service)) | nothing; the request row reads Replacement Pending until the replacement joins |
| `unhealthy` | `err` | the health check daemon reports failing checks, named in `detail`; the node still schedules unless it is also cordoned | move work off the node, or cordon it |
| `tainted` | `warn` | schedulable, with a taint the platform's health checker placed | as `unhealthy` |
| `cordoned` | `warn` | cordoned without a platform marker: your own cordon; `detail` lists the taints | uncordon when you are done |
| `missing` | `err` | the cluster answered and this node is absent from its node list | check `kubectl get nodes` |
| `creating`, `joining`, `error` | `warn`; `err` for `error` | the platform's lifecycle status while the node is not ready. Lifecycle is never health | wait |

A node mid move leaves the list: the board drops it and its machine
read answers `404`, so no `moving` word rides the wire. A node with a
scheduler row is judged from that row. A node with no row on a board
that lists other nodes is `missing`. A node with no row on a feed that
did not answer is `unobserved`, unless the platform's own record marks
it `unreachable` or its lifecycle status is not `ready`; those words
stand, with `observed: false`.

## Island staleness

Two fields tell a dark island from a broken node. `island_stale` on
the board, the machine page and [`GET /api/k8s/cluster`](https://docs.nationalcompute.com/api/k8s-cluster.md#the-clusters-facts)
is `null` while the platform hears the island's heartbeat and `{since}`
once it has lost the island. Every verdict then reads
`observed: false`: a node the dark feed would have read as healthy
reads `unobserved`, a node with a word of its own keeps it. `feed_error`
on `GET /api/k8s/cluster` is the last error of the platform's pull of
the node health feed, `null` when the last pull landed. `feed_age_s` is
the data age of what that feed serves, in seconds: the time since the
newest landed pull plus the platform's cache age at that pull, `null`
when no pull has ever landed. A failing pull leaves the age growing;
it never resets. `k8snodes.error` on the board names a node feed the
island could not answer on this read, and its rows then read
`observed: false`. Read these before acting on a red verdict.

On the MCP server the same two bodies are `grid_api get /api/k8s/nodes`
(the board) and `grid_api get /api/k8s/nodes/<node>` (one machine, the
REST spelling; the literal `/api/k8s/nodes/{node}` with `node` in
`params` is the same call), with `params` `cluster` and `source`
(`mirror` for the platform's mirror, no island dial). A malformed node
name is `422 bad-request`, an unknown one `404 not-found`.
