Kubernetes limit price API¶
For a Kubernetes cluster, demand is derived from the cluster's own jobs:
a GPU job is one request priced in whole nodes on the market, capacity
follows the request, and idle nodes return to the pool
(Kubernetes clusters covers the full behavior).
There is nothing to declare except the
price — one number per cluster, the most you'll pay per GPU-hour —
or per job, through the nationalcompute.com/limit-price label.
Contract: GET /api/k8s/openapi.json.
| Route | Auth | What it does |
|---|---|---|
GET /api/k8s/bid |
capacity:read |
current limit price, its standing verdict, version |
PUT /api/k8s/bid |
capacity:write |
set (or withdraw) the limit price; expected_version makes it compare and set |
GET /api/k8s/storage/volumes |
capacity:read |
your shared storage volumes (below) |
DELETE /api/orgs/{org_id}/k8s/storage/volumes/{volume_id} |
console session only | destroy a preserved volume (below) |
The same contract carries the five market analytics routes the console's Burst Capacity page reads, and the cluster facts and Base Load book.
There are no slots, no node counts, no node IPs, and no release verbs on this surface — submit jobs and the market does the rest.
Setting the limit price¶
curl -X PUT -H "Authorization: Bearer $NC_TOKEN" \
-H "Content-Type: application/json" \
-d '{"max_price_per_gpu_hour": <price>}' \
https://nationalcompute.com/api/k8s/bid
- The value is USD per GPU-hour, whole cents. The effective billed rate is often lower — this is a ceiling, not a price.
0withdraws the limit price.- A declare too low to ever clear the market is refused —
bid-too-low. Check the price-to-win ladder and the market feed for what wins now. - Add
"cluster": "<name>"(PUT body) or?cluster=<name>(GET) only with an org-wide token when your org holds several Kubernetes clusters. - Add
"expected_version": <n>with theversionthe GET returned. The write applies only while the version still matches. Otherwise the PUT reads409 version-mismatchwith the liveversionin the body and nothing changes. Without the field the write applies unconditionally. Your human on the console and your agent share one limit price, so send it. The version counts every accepted write on the cluster, a withdrawal included, and never restarts: a0that withdraws the limit price bumps it, and a version read before the withdrawal is stale for good.
The PUT returns the applied state — the same body the GET serves:
{
"cluster": "acme-inference",
"max_price_per_gpu_hour": …,
"bid_too_low": false,
"version": 4,
"updated_at": "2026-08-26T15:20:07Z",
"updated_by": "token:8dcbc6",
"billing_hold": null,
"spend_limit_enforced": true,
"billing_hold_enforced": true
}
updated_by names who wrote the standing limit price: token:<token_id>
for an agent, a member's console identity for a human, null without a
limit price. Read it before changing a ceiling. A value that is not a token is
a human's decision. The organization's whole trail, human and token
writes side by side, is GET /api/org/activity.
billing_hold is your org's billing hold state, or null. When set:
state (hold or grace), reason, and for hold a stage —
notice (new capacity paused) or enforce (capacity being reclaimed) —
plus enforce_after_ms. A hold never alters the declared limit price itself.
spend_limit_enforced and billing_hold_enforced say whether your
organization's spend limit and a balance at or below zero hold its
capacity at all. false means the platform records the figure and does
not act on it for your organization's clusters, so the cap is the
instruction a person gave the agent and nothing else stops it.
The body also carries gpu_model (the GPU class this limit price buys; null
while the platform has not set one), gpu_vendor (its maker),
gpus_per_node (the auction's supply unit) and bid_too_low (true
when the standing limit price is too low to ever clear the market,
null while nothing is declared);
how the bid prices the market
explains the supply unit.
A declare too low to ever clear is refused with bid-too-low — the
error reference covers the full catalog.
How the bid prices the market¶
Each request bids its gang's GPU total × your limit price — the
cluster's declared price, or the job's own
nationalcompute.com/limit-price label (USD per GPU-hour) when set;
"0" is a real zero limit price, never a fallback to the cluster's price — in
whole nodes, so a request's exposure is its node count ×
gpus_per_node × your ceiling. The limit price is read live every auction tick
for the request's whole life: you may change a job's limit price at any
time; an increase is always fine, and a decrease below the price the
job held when its gang started (Provisioned=True) voids its
protection window.
Capacity sells in whole nodes.
Launch admission enforces the
shape that clears. Every pod template's GPU request must be a multiple
of gpus_per_node. A smaller request is refused with JobSizeTooSmall.
So a request prices its nodes whole. On a cluster with packing onto
the base load block enabled a
smaller pod is admitted. It runs on the block and bids nothing.
A ceiling can be too low to ever clear:
the write refuses it outright (422 bid-too-low), because pending
demand costs nothing but a bid the market can never meet means nothing
will ever arrive. What to bid is read off the market's own numbers —
the price-to-win ladder and the
market feed — never off a posted price.
Market conditions can also move past a standing bid (the write-time check never re-runs). You do not have to detect this yourself:
- The limit price read echoes
bid_too_low:truemeans the standing limit price can never clear at any packing — check the market's recent clearing prices and raise it.nullmeans there is no standing limit price, or no verdict to give. - The market also says so inside your cluster: the request's reason
reads
BidTooLowand its message, posted on the Job as an Event, names your own bid's numbers (queue events). The console's Workloads page shows the verdict as the reason bid too low to ever clear. The signals clear the moment the bid becomes viable.
No limit price means $0. Your jobs still create demand, but nothing is granted until you set a real ceiling.
CPU-only work is outside the market. A job requesting no GPUs runs on the cluster's CPU worker and neither bids nor holds a GPU node: CPU-only queued work doesn't pull capacity, and a GPU node running only CPU pods reads as idle and returns to the pool after the five-minute grace, billed at the departed job's rate (once its one-hour minimum hold has run).
Queue events¶
The market posts its verdicts inside your cluster on the job's
ProvisioningRequest — as the reason of its Provisioned=False
condition, with the sentence as the message — and as a Normal Event on
the owning Job, JobSet or MPIJob, one per change of reason
(kubectl describe job). Kueue copies the same message into the
Workload's admission check. A new request carries no reason until the
market's first read, one tick after it appears. Reasons, highest
precedence first:
| Reason | Meaning | Clears |
|---|---|---|
BalanceTooLow |
the org's balance cannot fund two hours of the request at its limit price; the message names the balance available and what the request costs | after a top-up |
BidTooLow |
the limit price is too low to ever clear the market; the message names your own bid's numbers | when the bid becomes viable |
NoSupply |
the site holds fewer nodes than the gang needs | when the site grows, or with a smaller gang |
PendingSupply |
protected nodes hold the supply the gang needs | on its own, as windows lapse |
Outbid |
lost on price; the message names the limit price that wins right now (outbid · needs over $X per GPU-hour) |
raise the limit price, or wait |
ReservedBusy |
the request is smaller than one node and waits for GPUs of your base load block to free; it never bids; the message names the request's GPU count and why it waits | when block GPUs free |
ReservedTooSmall |
the request's pods are each smaller than one node and together exceed one node; a request with pods that small runs on one reserved node only and never bids | with whole node pods or with fewer pods that fit one node |
Pending |
no verdict this tick (waiting for the market): the request was seated this very tick, or the tick left it unfilled |
next tick |
The two reserved reasons appear only on clusters with reservation packing enabled. On such a cluster a request smaller than one node takes no other market reason. A higher limit price never clears it.
A grant reads NodeGranted on the Job once per node; the request's
condition reads Provisioning (k of N granted nodes ready) until it
turns Provisioned=True and the gang starts. On a cluster with packing
enabled a request smaller than one node reads waiting for N GPUs to
free on reserved node <node> while it waits for GPUs. Once seated it
reads reserved node <node> is ready with N GPUs for this job. A
reclaim rides the request's PreemptionNotice condition and
NodePreempting events on the node and its pods; at the deadline the
whole job is suspended and requeued as a new request
(when the market reclaims a node).
A granted node that leaves the cluster after Provisioned=True fails the
request the same way (Failed=True, reason MarketRevoked) and the job
requeues as a new request.
A site out of capacity shows on the console's Workloads page as
capacity unavailable; the request's own reason carries no shortage
token and reads PendingSupply, Outbid or Pending as usual
(when the site has no capacity).
Shared storage volumes¶
Clusters provisioned with shared storage carry a shared volume
(Kubernetes clusters). The same
contract lists your organization's volumes and deletes a preserved one.
Both routes are scoped to your organization: a volume outside your org
is a 404, the same as one that does not exist.
| Route | Auth | What it does |
|---|---|---|
GET /api/k8s/storage/volumes |
capacity:read |
your volumes: attached and preserved |
DELETE /api/orgs/{org_id}/k8s/storage/volumes/{volume_id} |
console session only | destroy one preserved volume (irreversible) |
The list takes an API token. The organization is the token's own and nothing identifies it on the wire; the list is organization wide whatever the token's binding, so scripts can inventory your volumes.
curl -H "Authorization: Bearer $NC_TOKEN" \
https://nationalcompute.com/api/k8s/storage/volumes
Each volume row carries volume_id, name, label, cluster (null
for a preserved volume), status (attached or preserved),
provisioned and used bytes (null until the platform's storage scan
lands, never 0), mounted (how many nodes currently mount it, null
when unknown), billing (the current charge where storage billing is
enabled for your site, else null) and preserved
({former_cluster, deleted_at, used_at_delete} for a preserved volume,
else null).
label is the name the console shows for the volume: <cluster>-shared,
with the former cluster for a preserved volume. name is the storage
name. The two can differ; when they do, the console shows the storage
name under the label. name is the value the delete's confirm takes.
label is a display name and is not unique: two volumes can share it
after a cluster is deleted and created again under the same name.
Scripts key on volume_id or name, never on label.
Deleting a volume is a console action¶
The DELETE takes no API token. Sign in to the console and use
Delete volume on the Storage page
(seeing and deleting your volume).
A request carrying a Bearer token is refused with
403 session-required, whatever the token's scope.
A limit price is reversible. Data destruction is not. An irreversible delete keeps a human in the loop: a member of your organization, signed in, typing the volume's storage name back.
The console's delete rides the same route, with the same body
({"confirm": "<volume name>"}) and the same refusals:
confirmmust be the volume'sname, the storage name shown under its label. The display label is refused. Anything else isconfirmation-mismatchand nothing is destroyed.- Only a preserved volume can be deleted here. An attached one is
volume-attached: delete the cluster first (the volume detaches with it and is preserved by default), or opt to delete the storage with the cluster. - A preserved volume a node still holds is
volume-held; retry in a few minutes. - The delete is accepted asynchronously: the response is
{destroyed, operation}and the volume leaves the list once it is gone. There is no undo.
Tokens¶
Token minting is the normal org token flow (authentication); a token can be bound to a Kubernetes cluster. A k8s-bound token is refused on the VM surfaces and a VM-bound one is refused here.