Skip to content

Kubernetes limit price API

For a Kubernetes cluster, demand is derived from the cluster's own jobs: a GPU job is one request priced in whole nodes on the market, capacity follows the request, and idle nodes return to the pool (Kubernetes clusters covers the full behavior). There is nothing to declare except the price — one number per cluster, the most you'll pay per GPU-hour — or per job, through the nationalcompute.com/limit-price label.

Contract: GET /api/k8s/openapi.json.

Route Auth What it does
GET /api/k8s/bid capacity:read current limit price, its standing verdict, version
PUT /api/k8s/bid capacity:write set (or withdraw) the limit price; expected_version makes it compare and set
GET /api/k8s/storage/volumes capacity:read your shared storage volumes (below)
DELETE /api/orgs/{org_id}/k8s/storage/volumes/{volume_id} console session only destroy a preserved volume (below)

The same contract carries the five market analytics routes the console's Burst Capacity page reads, and the cluster facts and Base Load book.

There are no slots, no node counts, no node IPs, and no release verbs on this surface — submit jobs and the market does the rest.

Setting the limit price

curl -X PUT -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"max_price_per_gpu_hour": <price>}' \
  https://nationalcompute.com/api/k8s/bid
  • The value is USD per GPU-hour, whole cents. The effective billed rate is often lower — this is a ceiling, not a price.
  • 0 withdraws the limit price.
  • A declare too low to ever clear the market is refused — bid-too-low. Check the price-to-win ladder and the market feed for what wins now.
  • Add "cluster": "<name>" (PUT body) or ?cluster=<name> (GET) only with an org-wide token when your org holds several Kubernetes clusters.
  • Add "expected_version": <n> with the version the GET returned. The write applies only while the version still matches. Otherwise the PUT reads 409 version-mismatch with the live version in the body and nothing changes. Without the field the write applies unconditionally. Your human on the console and your agent share one limit price, so send it. The version counts every accepted write on the cluster, a withdrawal included, and never restarts: a 0 that withdraws the limit price bumps it, and a version read before the withdrawal is stale for good.

The PUT returns the applied state — the same body the GET serves:

{
  "cluster": "acme-inference",
  "max_price_per_gpu_hour": …,
  "bid_too_low": false,
  "version": 4,
  "updated_at": "2026-08-26T15:20:07Z",
  "updated_by": "token:8dcbc6",
  "billing_hold": null,
  "spend_limit_enforced": true,
  "billing_hold_enforced": true
}

updated_by names who wrote the standing limit price: token:<token_id> for an agent, a member's console identity for a human, null without a limit price. Read it before changing a ceiling. A value that is not a token is a human's decision. The organization's whole trail, human and token writes side by side, is GET /api/org/activity.

billing_hold is your org's billing hold state, or null. When set: state (hold or grace), reason, and for hold a stage — notice (new capacity paused) or enforce (capacity being reclaimed) — plus enforce_after_ms. A hold never alters the declared limit price itself. spend_limit_enforced and billing_hold_enforced say whether your organization's spend limit and a balance at or below zero hold its capacity at all. false means the platform records the figure and does not act on it for your organization's clusters, so the cap is the instruction a person gave the agent and nothing else stops it.

The body also carries gpu_model (the GPU class this limit price buys; null while the platform has not set one), gpu_vendor (its maker), gpus_per_node (the auction's supply unit) and bid_too_low (true when the standing limit price is too low to ever clear the market, null while nothing is declared); how the bid prices the market explains the supply unit.

A declare too low to ever clear is refused with bid-too-low — the error reference covers the full catalog.

How the bid prices the market

Each request bids its gang's GPU total × your limit price — the cluster's declared price, or the job's own nationalcompute.com/limit-price label (USD per GPU-hour) when set; "0" is a real zero limit price, never a fallback to the cluster's price — in whole nodes, so a request's exposure is its node count × gpus_per_node × your ceiling. The limit price is read live every auction tick for the request's whole life: you may change a job's limit price at any time; an increase is always fine, and a decrease below the price the job held when its gang started (Provisioned=True) voids its protection window.

Capacity sells in whole nodes. Launch admission enforces the shape that clears. Every pod template's GPU request must be a multiple of gpus_per_node. A smaller request is refused with JobSizeTooSmall. So a request prices its nodes whole. On a cluster with packing onto the base load block enabled a smaller pod is admitted. It runs on the block and bids nothing.

A ceiling can be too low to ever clear: the write refuses it outright (422 bid-too-low), because pending demand costs nothing but a bid the market can never meet means nothing will ever arrive. What to bid is read off the market's own numbers — the price-to-win ladder and the market feed — never off a posted price.

Market conditions can also move past a standing bid (the write-time check never re-runs). You do not have to detect this yourself:

  • The limit price read echoes bid_too_low: true means the standing limit price can never clear at any packing — check the market's recent clearing prices and raise it. null means there is no standing limit price, or no verdict to give.
  • The market also says so inside your cluster: the request's reason reads BidTooLow and its message, posted on the Job as an Event, names your own bid's numbers (queue events). The console's Workloads page shows the verdict as the reason bid too low to ever clear. The signals clear the moment the bid becomes viable.

No limit price means $0. Your jobs still create demand, but nothing is granted until you set a real ceiling.

CPU-only work is outside the market. A job requesting no GPUs runs on the cluster's CPU worker and neither bids nor holds a GPU node: CPU-only queued work doesn't pull capacity, and a GPU node running only CPU pods reads as idle and returns to the pool after the five-minute grace, billed at the departed job's rate (once its one-hour minimum hold has run).

Queue events

The market posts its verdicts inside your cluster on the job's ProvisioningRequest — as the reason of its Provisioned=False condition, with the sentence as the message — and as a Normal Event on the owning Job, JobSet or MPIJob, one per change of reason (kubectl describe job). Kueue copies the same message into the Workload's admission check. A new request carries no reason until the market's first read, one tick after it appears. Reasons, highest precedence first:

Reason Meaning Clears
BalanceTooLow the org's balance cannot fund two hours of the request at its limit price; the message names the balance available and what the request costs after a top-up
BidTooLow the limit price is too low to ever clear the market; the message names your own bid's numbers when the bid becomes viable
NoSupply the site holds fewer nodes than the gang needs when the site grows, or with a smaller gang
PendingSupply protected nodes hold the supply the gang needs on its own, as windows lapse
Outbid lost on price; the message names the limit price that wins right now (outbid · needs over $X per GPU-hour) raise the limit price, or wait
ReservedBusy the request is smaller than one node and waits for GPUs of your base load block to free; it never bids; the message names the request's GPU count and why it waits when block GPUs free
ReservedTooSmall the request's pods are each smaller than one node and together exceed one node; a request with pods that small runs on one reserved node only and never bids with whole node pods or with fewer pods that fit one node
Pending no verdict this tick (waiting for the market): the request was seated this very tick, or the tick left it unfilled next tick

The two reserved reasons appear only on clusters with reservation packing enabled. On such a cluster a request smaller than one node takes no other market reason. A higher limit price never clears it.

A grant reads NodeGranted on the Job once per node; the request's condition reads Provisioning (k of N granted nodes ready) until it turns Provisioned=True and the gang starts. On a cluster with packing enabled a request smaller than one node reads waiting for N GPUs to free on reserved node <node> while it waits for GPUs. Once seated it reads reserved node <node> is ready with N GPUs for this job. A reclaim rides the request's PreemptionNotice condition and NodePreempting events on the node and its pods; at the deadline the whole job is suspended and requeued as a new request (when the market reclaims a node). A granted node that leaves the cluster after Provisioned=True fails the request the same way (Failed=True, reason MarketRevoked) and the job requeues as a new request. A site out of capacity shows on the console's Workloads page as capacity unavailable; the request's own reason carries no shortage token and reads PendingSupply, Outbid or Pending as usual (when the site has no capacity).

Shared storage volumes

Clusters provisioned with shared storage carry a shared volume (Kubernetes clusters). The same contract lists your organization's volumes and deletes a preserved one. Both routes are scoped to your organization: a volume outside your org is a 404, the same as one that does not exist.

Route Auth What it does
GET /api/k8s/storage/volumes capacity:read your volumes: attached and preserved
DELETE /api/orgs/{org_id}/k8s/storage/volumes/{volume_id} console session only destroy one preserved volume (irreversible)

The list takes an API token. The organization is the token's own and nothing identifies it on the wire; the list is organization wide whatever the token's binding, so scripts can inventory your volumes.

curl -H "Authorization: Bearer $NC_TOKEN" \
  https://nationalcompute.com/api/k8s/storage/volumes

Each volume row carries volume_id, name, label, cluster (null for a preserved volume), status (attached or preserved), provisioned and used bytes (null until the platform's storage scan lands, never 0), mounted (how many nodes currently mount it, null when unknown), billing (the current charge where storage billing is enabled for your site, else null) and preserved ({former_cluster, deleted_at, used_at_delete} for a preserved volume, else null).

label is the name the console shows for the volume: <cluster>-shared, with the former cluster for a preserved volume. name is the storage name. The two can differ; when they do, the console shows the storage name under the label. name is the value the delete's confirm takes. label is a display name and is not unique: two volumes can share it after a cluster is deleted and created again under the same name. Scripts key on volume_id or name, never on label.

Deleting a volume is a console action

The DELETE takes no API token. Sign in to the console and use Delete volume on the Storage page (seeing and deleting your volume). A request carrying a Bearer token is refused with 403 session-required, whatever the token's scope.

A limit price is reversible. Data destruction is not. An irreversible delete keeps a human in the loop: a member of your organization, signed in, typing the volume's storage name back.

The console's delete rides the same route, with the same body ({"confirm": "<volume name>"}) and the same refusals:

  • confirm must be the volume's name, the storage name shown under its label. The display label is refused. Anything else is confirmation-mismatch and nothing is destroyed.
  • Only a preserved volume can be deleted here. An attached one is volume-attached: delete the cluster first (the volume detaches with it and is preserved by default), or opt to delete the storage with the cluster.
  • A preserved volume a node still holds is volume-held; retry in a few minutes.
  • The delete is accepted asynchronously: the response is {destroyed, operation} and the volume leaves the list once it is gone. There is no undo.

Tokens

Token minting is the normal org token flow (authentication); a token can be bound to a Kubernetes cluster. A k8s-bound token is refused on the VM surfaces and a VM-bound one is refused here.