# Kubernetes limit price API

For a Kubernetes cluster, demand is derived from the cluster's own jobs:
a GPU job is one request priced in whole nodes on the market, capacity
follows the request, and idle nodes return to the pool
([Kubernetes clusters](https://docs.nationalcompute.com/kubernetes.md) covers the full behavior).
There is nothing to declare except the
price — one number per cluster, the **most you'll pay per GPU-hour** —
or per job, through the `nationalcompute.com/limit-price` label.

Contract: [`GET /api/k8s/openapi.json`](https://nationalcompute.com/api/k8s/openapi.json).

| Route | Auth | What it does |
|---|---|---|
| `GET /api/k8s/bid` | `capacity:read` | current limit price, its standing verdict, version |
| `PUT /api/k8s/bid` | `capacity:write` | set (or withdraw) the limit price; `expected_version` makes it compare and set |
| `GET /api/k8s/storage/volumes` | `capacity:read` | your shared storage volumes ([below](#shared-storage-volumes)) |
| `DELETE /api/orgs/{org_id}/k8s/storage/volumes/{volume_id}` | console session only | destroy a preserved volume ([below](#shared-storage-volumes)) |

The same contract carries the five [market analytics](https://docs.nationalcompute.com/api/k8s-market.md)
routes the console's Burst Capacity page reads, and the
[cluster facts and Base Load book](https://docs.nationalcompute.com/api/k8s-cluster.md).

There are no slots, no node counts, no node IPs, and no release verbs on
this surface — submit jobs and the market does the rest.

## Setting the limit price { #setting-the-bid }

```sh
curl -X PUT -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"max_price_per_gpu_hour": <price>}' \
  https://nationalcompute.com/api/k8s/bid
```

- The value is USD per GPU-hour, whole cents. The effective billed rate
  is often lower — this is a ceiling, not a price.
- `0` withdraws the limit price.
- A declare too low to ever clear the market is refused —
  `bid-too-low`. Check the [price-to-win ladder](https://docs.nationalcompute.com/api/k8s-cluster-pages.md)
  and the [market feed](https://docs.nationalcompute.com/api/market-feed.md) for what wins now.
- Add `"cluster": "<name>"` (PUT body) or `?cluster=<name>` (GET) only
  with an org-wide token when your org holds several Kubernetes
  clusters.
- Add `"expected_version": <n>` with the `version` the GET returned.
  The write applies only while the version still matches. Otherwise
  the PUT reads `409 version-mismatch` with the live `version` in the
  body and nothing changes. Without the field the write applies
  unconditionally. Your human on the console and your agent share one
  limit price, so send it. The version counts every accepted write on the
  cluster, a withdrawal included, and never restarts: a `0` that
  withdraws the limit price bumps it, and a version read before the withdrawal
  is stale for good.

The PUT returns the applied state — the same body the GET serves:

```json
{
  "cluster": "acme-inference",
  "max_price_per_gpu_hour": …,
  "bid_too_low": false,
  "version": 4,
  "updated_at": "2026-08-26T15:20:07Z",
  "updated_by": "token:8dcbc6",
  "billing_hold": null,
  "spend_limit_enforced": true,
  "billing_hold_enforced": true
}
```

`updated_by` names who wrote the standing limit price: `token:<token_id>`
for an agent, a member's console identity for a human, `null` without a
limit price. Read it before changing a ceiling. A value that is not a token is
a human's decision. The organization's whole trail, human and token
writes side by side, is [`GET /api/org/activity`](https://docs.nationalcompute.com/api/organization.md#the-organizations-trail).

`billing_hold` is your org's billing hold state, or `null`. When set:
`state` (`hold` or `grace`), `reason`, and for `hold` a `stage` —
`notice` (new capacity paused) or `enforce` (capacity being reclaimed) —
plus `enforce_after_ms`. A hold never alters the declared limit price itself.
`spend_limit_enforced` and `billing_hold_enforced` say whether your
organization's spend limit and a balance at or below zero hold its
capacity at all. `false` means the platform records the figure and does
not act on it for your organization's clusters, so the cap is the
instruction a person gave the agent and nothing else stops it.

The body also carries `gpu_model` (the GPU class this limit price buys; `null`
while the platform has not set one), `gpu_vendor` (its maker),
`gpus_per_node` (the auction's supply unit) and `bid_too_low` (`true`
when the standing limit price is too low to ever clear the market,
`null` while nothing is declared);
[how the bid prices the market](#how-the-bid-prices-the-market)
explains the supply unit.

A declare too low to ever clear is refused with `bid-too-low` — the
[error reference](https://docs.nationalcompute.com/api/errors.md) covers the full catalog.

## How the bid prices the market

Each request bids its gang's GPU total × your limit price — the
cluster's declared price, or the job's own
`nationalcompute.com/limit-price` label (USD per GPU-hour) when set;
`"0"` is a real zero limit price, never a fallback to the cluster's price — in
whole nodes, so a request's exposure is its node count ×
`gpus_per_node` × your ceiling. The limit price is read live every auction tick
for the request's whole life: you may change a job's limit price at any
time; an increase is always fine, and a decrease below the price the
job held when its gang started (`Provisioned=True`) voids its
[protection window](https://docs.nationalcompute.com/market.md#minimum-duration-protection).

Capacity sells **in whole nodes**.
[Launch admission](https://docs.nationalcompute.com/kubernetes.md#launch-admission) enforces the
shape that clears. Every pod template's GPU request must be a multiple
of `gpus_per_node`. A smaller request is refused with `JobSizeTooSmall`.
So a request prices its nodes whole. On a cluster with packing onto
the [base load block](https://docs.nationalcompute.com/market.md#base-load-capacity) enabled a
smaller pod is admitted. It runs on the block and bids nothing.

A ceiling can be [too low to ever clear](https://docs.nationalcompute.com/market.md#a-bid-that-can-never-clear):
the write refuses it outright (`422 bid-too-low`), because pending
demand costs nothing but a bid the market can never meet means nothing
will ever arrive. What to bid is read off the market's own numbers —
the [price-to-win ladder](https://docs.nationalcompute.com/api/k8s-cluster-pages.md) and the
[market feed](https://docs.nationalcompute.com/api/market-feed.md) — never off a posted price.

Market conditions can also move past a **standing** bid (the
write-time check never re-runs). You do not have to detect this
yourself:

- The limit price read echoes `bid_too_low`: `true` means the standing
  limit price can never clear at any packing — check the market's
  recent clearing prices and raise it. `null` means there is no
  standing limit price, or no verdict to give.
- The market also says so inside your cluster: the request's reason
  reads `BidTooLow` and its message, posted on the Job as an Event,
  names your own bid's numbers ([queue events](#queue-events)). The
  console's Workloads page shows the verdict as the reason **bid too
  low to ever clear**. The signals clear the moment the bid becomes
  viable.

**No limit price means $0.** Your jobs still create demand, but nothing
is granted until you set a real ceiling.

**CPU-only work is outside the market.** A job requesting no GPUs runs
on the cluster's CPU worker and neither bids nor holds a GPU node:
CPU-only queued work doesn't pull capacity, and a GPU node running only
CPU pods reads as idle and returns to the pool after the five-minute
grace, billed at the departed job's rate (once its one-hour
[minimum hold](https://docs.nationalcompute.com/market.md#minimum-duration-protection) has run).

## Queue events

The market posts its verdicts inside your cluster on the job's
ProvisioningRequest — as the reason of its `Provisioned=False`
condition, with the sentence as the message — and as a Normal Event on
the owning `Job`, `JobSet` or `MPIJob`, one per change of reason
(`kubectl describe job`). Kueue copies the same message into the
Workload's admission check. A new request carries no reason until the
market's first read, one tick after it appears. Reasons, highest
precedence first:

| Reason | Meaning | Clears |
|---|---|---|
| `BalanceTooLow` | the org's balance cannot fund two hours of the request at its limit price; the message names the balance available and what the request costs | after a top-up |
| `BidTooLow` | the limit price is too low to ever clear the market; the message names your own bid's numbers | when the bid becomes viable |
| `NoSupply` | the site holds fewer nodes than the gang needs | when the site grows, or with a smaller gang |
| `PendingSupply` | protected nodes hold the supply the gang needs | on its own, as windows lapse |
| `Outbid` | lost on price; the message names the limit price that wins right now (`outbid · needs over $X per GPU-hour`) | raise the limit price, or wait |
| `ReservedBusy` | the request is smaller than one node and waits for GPUs of your base load block to free; it never bids; the message names the request's GPU count and why it waits | when block GPUs free |
| `ReservedTooSmall` | the request's pods are each smaller than one node and together exceed one node; a request with pods that small runs on one reserved node only and never bids | with whole node pods or with fewer pods that fit one node |
| `Pending` | no verdict this tick (`waiting for the market`): the request was seated this very tick, or the tick left it unfilled | next tick |

The two reserved reasons appear only on clusters with reservation
packing enabled. On such a cluster a request smaller than one node
takes no other market reason. A higher limit price never clears it.

A grant reads `NodeGranted` on the Job once per node; the request's
condition reads `Provisioning` (`k of N granted nodes ready`) until it
turns `Provisioned=True` and the gang starts. On a cluster with packing
enabled a request smaller than one node reads `waiting for N GPUs to
free on reserved node <node>` while it waits for GPUs. Once seated it
reads `reserved node <node> is ready with N GPUs for this job`. A
reclaim rides the request's `PreemptionNotice` condition and
`NodePreempting` events on the node and its pods; at the deadline the
whole job is suspended and requeued as a new request
([when the market reclaims a node](https://docs.nationalcompute.com/kubernetes.md#when-the-market-reclaims-a-node)).
A granted node that leaves the cluster after `Provisioned=True` fails the
request the same way (`Failed=True`, reason `MarketRevoked`) and the job
requeues as a new request.
A site out of capacity shows on the console's Workloads page as
**capacity unavailable**; the request's own reason carries no shortage
token and reads `PendingSupply`, `Outbid` or `Pending` as usual
([when the site has no capacity](https://docs.nationalcompute.com/kubernetes.md#when-the-site-has-no-capacity)).

## Shared storage volumes

Clusters provisioned with shared storage carry a shared volume
([Kubernetes clusters](https://docs.nationalcompute.com/kubernetes.md#the-shared-volume)). The same
contract lists your organization's volumes and deletes a preserved one.
Both routes are scoped to your organization: a volume outside your org
is a `404`, the same as one that does not exist.

| Route | Auth | What it does |
|---|---|---|
| `GET /api/k8s/storage/volumes` | `capacity:read` | your volumes: attached and preserved |
| `DELETE /api/orgs/{org_id}/k8s/storage/volumes/{volume_id}` | console session only | destroy one preserved volume (irreversible) |

The list takes an API token. The organization is the token's own and
nothing identifies it on the wire; the list is organization wide
whatever the token's binding, so scripts can inventory your volumes.

```sh
curl -H "Authorization: Bearer $NC_TOKEN" \
  https://nationalcompute.com/api/k8s/storage/volumes
```

Each volume row carries `volume_id`, `name`, `label`, `cluster` (`null`
for a preserved volume), `status` (`attached` or `preserved`),
`provisioned` and `used` bytes (`null` until the platform's storage scan
lands, never `0`), `mounted` (how many nodes currently mount it, `null`
when unknown), `billing` (the current charge where storage billing is
enabled for your site, else `null`) and `preserved`
(`{former_cluster, deleted_at, used_at_delete}` for a preserved volume,
else `null`).

`label` is the name the console shows for the volume: `<cluster>-shared`,
with the former cluster for a preserved volume. `name` is the storage
name. The two can differ; when they do, the console shows the storage
name under the label. `name` is the value the delete's `confirm` takes.
`label` is a display name and is not unique: two volumes can share it
after a cluster is deleted and created again under the same name.
Scripts key on `volume_id` or `name`, never on `label`.

### Deleting a volume is a console action

The DELETE takes **no API token**. Sign in to the console and use
**Delete volume** on the Storage page
([seeing and deleting your volume](https://docs.nationalcompute.com/kubernetes.md#seeing-and-deleting-your-volume)).
A request carrying a `Bearer` token is refused with
`403 session-required`, whatever the token's scope.

A limit price is reversible. Data destruction is not. An irreversible delete
keeps a human in the loop: a member of your organization, signed in,
typing the volume's storage name back.

The console's delete rides the same route, with the same body
(`{"confirm": "<volume name>"}`) and the same refusals:

- `confirm` must be the volume's `name`, the storage name shown under
  its label. The display label is refused. Anything else is
  `confirmation-mismatch` and nothing is destroyed.
- Only a **preserved** volume can be deleted here. An attached one is
  `volume-attached`: delete the cluster first (the volume detaches with
  it and is preserved by default), or opt to delete the storage with
  the cluster.
- A preserved volume a node still holds is `volume-held`; retry in a
  few minutes.
- The delete is accepted asynchronously: the response is
  `{destroyed, operation}` and the volume leaves the list once it is
  gone. There is no undo.

## Tokens

Token minting is the normal org token flow ([authentication](https://docs.nationalcompute.com/authentication.md));
a token can be bound to a Kubernetes cluster. A k8s-bound token is
refused on the VM surfaces and a VM-bound one is refused here.
