Kubernetes clusters¶
A Kubernetes tenancy is a dedicated cluster: you hold the kubeconfig, you deploy your own workloads, and the platform moves GPU nodes in and out of your cluster as your jobs demand them. Nobody runs your containers for you — the cluster is yours; the platform only decides how many nodes it has.
Cluster nodes run the Kubernetes agent directly on the metal — there is no hypervisor layer in the Kubernetes offering.
Connecting¶
The kubeconfig is a public client and the cluster's certificate. It carries no secret. Fetch it with your org API token:
curl -fsSL -H "Authorization: Bearer $NC_TOKEN" \
"https://nationalcompute.com/api/k8s/cluster/kubeconfig?cluster=<cluster>" \
-o ~/.kube/<cluster>.yaml
The token names your organization, so the route resolves <cluster>
inside it. Another organization's cluster can never answer, whatever its
name. The first kubectl call prints a sign in link. A person opens
the link once and approves. After that kubectl refreshes its own sign
in while it is in use. A day without a kubectl call needs one more
approval: the next call prints a fresh link and waits on it. A
kubectl call that seems hung is showing a sign in link — agents run
kubectl where its output stays readable while the command runs, and
relay the link to a person.
The setup script installs kubectl and the sign in plugin when they are
missing and merges a <cluster> context into ~/.kube/config:
curl -fsSL https://access.nationalcompute.com/setup.sh | sh -s -- you@example.com <cluster>
The script reads the token from NC_TOKEN or from
~/.config/nationalcompute/token and fetches the kubeconfig by token
when it finds one. Without a token it asks the public onboarding path by
name. That path refuses a name that more than one cluster carries with
409 ambiguous-cluster and points at the token route.
The script never waits on the sign in: run without a terminal, it
prints the link and exits while sign-in polls in the background.
Approval lands the token on its own, and
kubectl --context <cluster> get nodes then answers at once. A fresh
link in its place means the code expired; the new link replaces the
old.
Workloads and automation authenticate to the cluster with ServiceAccount tokens rather than a sign in, and every cluster is a public OIDC issuer for them — see Authentication.
Demand comes from your jobs¶
There is nothing to declare except your price. The market derives your demand from the jobs you submit, through Kueue — the queue installed in your cluster that holds a job until the market grants its nodes:
- A GPU job is one request for whole nodes.
kubectl applyaJob,JobSetorMPIJobwhose pods request full-node GPU multiples; Kueue creates it suspended and files one request for the whole gang — every pod the job needs, priced together. When your limit price clears, the nodes join and every pod starts in the same moment, on exactly those nodes. Nothing else is needed: no queue name, nosuspend: true, and the toleration for the GPU worker taint is injected on admission. - A job smaller than one node runs on your base load block. On a cluster where the platform has enabled packing onto the base load block, a pod may request fewer GPUs than one node. Kueue files one request for it. The market seats it on a block node beside your other small jobs. It never bids. It never runs on a market node. Without a live block it waits.
- A waiting job has no pods.
kubectl --context <cluster> get podsis empty until the whole gang is granted; a gang is fully waiting or fully running, never "1 of 2". The exceptions are divisible: aDeploymentreplica or a bare GPU pod is its own request and exists as a pod heldSchedulingGateduntil granted. - Idle nodes return to the pool. Once a node's one-hour minimum hold has run, a node whose job has ended sheds after an idle grace of five minutes, billed at that job's rate, unless another job takes it first — the grace covers between-job gaps, so batchy workloads don't churn nodes.
- CPU-only work is outside the market. A job requesting no GPUs
runs at once on the cluster's CPU worker (below) and neither bids nor
holds a GPU node. It passes through the queue without a market check,
so it can show
Suspendedfor a moment.
The result is that capacity — and billing — tracks what you actually run: when the queue drains, nodes shed and the meter stops, with no action from you.
Reading the queue¶
The queue is read on the Job, never on pods. kubectl --context <cluster> get job shows
the Job Suspended; kubectl --context <cluster> get workloads lists its Workload (Kueue's
queue entry for the job, <kind>-<jobname>-<hash>) and whether it is
admitted; kubectl describe job carries the market's verdicts as
Events — one per change of reason, then NodeGranted per granted
node. The same verdict is the Provisioned=False condition of the
ProvisioningRequest — the object Kueue files to ask the market for the
gang's nodes, <workload>-market-<attempt> — read with
kubectl describe provisioningrequest, and is copied into the
Workload's admission check; once granted, the condition reads
Provisioning (k of N granted nodes ready) until every node is ready
and it turns Provisioned=True. The reason is empty until the market's
first read, one tick after the request appears; the
limit price API tables every reason. The
console's Workloads page lists the queued job before it has pods, with
the same reason, and marks a running gang gang scheduled.
The cluster CPU worker¶
Newly provisioned clusters come with one small CPU-only worker node in
addition to the control plane. It is a plain schedulable node with the
cpu role in kubectl --context <cluster> get nodes. It carries no GPUs and no taints, so
untolerating pods (queues, controllers, data prep, small services)
land there by default while GPU jobs wait in the queue.
The CPU worker sits outside the market: it never bids, never sheds on idleness, and is never reclaimed when the market clears above your ceiling. It is created with the cluster and removed with the cluster. It cannot be resized or multiplied; for CPU capacity beyond it, run CPU work on GPU nodes you already hold or contact support.
The limit price¶
PUT /api/k8s/bid sets the most you'll pay per
GPU-hour — your limit price. Each request bids its gang's GPU total ×
your limit price, in whole nodes, so a request's exposure is its node
count × gpus_per_node × your ceiling. A request smaller than one
node bids nothing. It waits for GPUs of your base load block. A job
carries its own limit price in the label
nationalcompute.com/limit-price: "<price>" (USD per
GPU-hour, as a quoted string), which replaces the cluster's for that job; "0" is a real
zero limit price — the job waits with the reason BidTooLow and never
falls back to the cluster's price. The limit price is read live at every tick for the
request's whole life, waiting or granted: you may change a job's limit
price at any time. An increase is always fine. A decrease below the
price the job held when its gang started (Provisioned=True) voids its
protection window (below).
With no limit price set, your jobs bid $0, which never clears:
the demand is visible, but nothing is granted until you declare a real
ceiling. A limit price too low to ever clear the market is refused at
write time (bid-too-low); a standing one the market moved past reads
bid_too_low: true on the limit price read, and its requests wait
with the reason BidTooLow — they never fail — and clear at the first
tick after you raise the price. See
how the bid prices the market
for the arithmetic.
The max cluster size¶
Every cluster has a max cluster size in GPUs — set by the platform,
read-only to you (an edit of the ClusterQueue, or of any other Kueue
object the platform owns, is refused with StationOwned), shown on the
cluster's console page and in-cluster as the quota of Kueue's
ClusterQueue market (kubectl --context <cluster> get clusterqueue market). While the
GPUs of your running and requested jobs stay under it, every new job
files its request at once. Beyond it, jobs wait in your own queue and
file their requests as room frees: the Workload's QuotaReserved
condition names the quota as the cause (kubectl describe workload),
and the console's Workloads page lists the job waiting, with no request
filed yet. A job larger than the limit on its own never reaches the
market; ask support to raise it.
Priority orders only your own queue¶
Priority orders a job against your other jobs, and only while the max
cluster size binds: set the label kueue.x-k8s.io/priority-class to
low, normal or high on a suspended Job — a job without the label
ranks below low — and the higher class files its request first when
room frees. Every job under the limit reaches the market at once, where
price decides — priority never buys market position over another tenant
and never evicts a running job. On a cluster with a
base load block the same order decides who takes the
GPUs of the block as they free: the higher class first, oldest first within a
class — so a serving Deployment marked high reclaims its node in the block
after a rollout ahead of waiting batch work.
When the site has no capacity¶
A shortage is not a pricing problem. When the site runs out of
capacity of the class your jobs request, no bid clears it: the request
stays in the queue, the console's Workloads page shows the verdict as
the reason capacity unavailable, and the
market feed's ticks read null while nothing
clears. Raising your limit price does not start the job sooner. The request's
own reason does not distinguish a shortage — it reads PendingSupply,
Outbid or Pending as usual; the console's Workloads page is where
the shortage verdict shows.
There is nothing to do. Pending demand costs nothing. Leave the limit price where it is; the platform checks for returning capacity on its own, and the job starts when capacity returns, with no action from you.
Launch admission¶
Submitting a GPU workload is checked at admission on market-managed
clusters — a violation is refused at kubectl apply with the exact
numbers in the error, never left silently queued:
- Full nodes on the market (
JobSizeTooSmall). Every pod template of aJob,JobSetorMPIJobmust request a multiple ofgpus_per_nodeGPUs. The same rule binds a bare pod. A multi node job runs as several full node pods. Templates requesting no GPUs (an MPI launcher, for one) are exempt. The market assigns whole nodes for strong isolation. A partial node GPU pod is refused rather than priced. The exception is a cluster with packing onto the base load block enabled (above). There a pod may request fewer GPUs than one node. Such a pod runs on the block only. A pod larger than one node must still request a whole number of nodes (JobSizeNotNodeAligned). - Minimum balance (
BalanceTooLow). The org's credit balance must cover two hours (currently) of the job at your limit price — 2 × limit price × the job's GPU total, counted acrossparallelism(or replicas). The denial names the required and current balance. The rule reduces churn — a job that takes nodes must be fundable at its own limit price for more than minutes. The denial binds at apply only — a job already in the queue is never refused later; a shortfall while it waits pends the request instead (below).
The market runs the same two-hour balance check again at every tick
while the request waits, in case the situation has changed: a request
the org's balance can no longer fund stays in the queue with the reason
BalanceTooLow, whose message names the balance available and what the
request costs (queue events); a top-up
clears it at the next tick and nothing fails.
The platform can switch packing onto the block off for a cluster. From
that moment the full node rule binds again at admission: a new pod
smaller than one node is refused with JobSizeTooSmall. A Deployment
update and a Job retry create new pods and are refused the same way.
Pods already running are never touched. Resubmit the work as whole
node pods until packing is switched on again.
When the market reclaims a node¶
If the market clears above your ceiling (or you withdraw the limit price), a
gang's nodes are reclaimed: the whole job is suspended, its pods are
deleted together, and the job requeues as a new request at your current
limit price — it resumes when granted again, on whatever nodes the
market grants then. A reclaim is never a failed job: the Workload goes
back to pending, and with the spec below in place it does not consume
the Job's backoffLimit. On the console's Workloads page the requeued
job's row shows the current episode — queued since the requeue, running
since the last grant — and its Run History the episodes and the totals.
A reclaim caused by a higher bid from another tenant against your
standing limit price comes with one minute's warning. A reclaim
caused by your own account — your limit price becomes too low to ever
clear, you withdraw or lower it, or your credit balance runs out — is
immediate, even when another tenant takes the node: no advance
warning, the drain starts at once, and pods get the same 60-second
eviction grace described below. At notice time the request
gains a PreemptionNotice=True condition (reason MarketPreemption)
naming the deadline, the node is cordoned and tainted, a
NodePreempting event is posted on the node and on each of its pods,
every pod gets a marketplace.nationalcompute.com/preempt-at
annotation carrying the deadline (readable in-pod through a
downward-API volume — annotation files update live), and the pods are
evicted through the Eviction API with the window as their grace period
— SIGTERM at notice is your checkpoint signal, and preempt-at the
deadline to checkpoint by. PodDisruptionBudgets are honored inside the
window, never past it. At the deadline the whole job is suspended and
requeued: the market revokes the request (it reads Failed=True,
reason MarketRevoked), and Kueue evicts the Workload, deletes
whatever pods remain and files the new request. Set
terminationGracePeriodSeconds: 60 on your pods so their own grace
matches the window. A reclaim withdrawn mid-notice uncordons the node,
flips the condition to PreemptionNotice=False (reason Rescinded)
and posts a PreemptionRescinded event on the node.
A granted node that leaves your cluster after the request turns
Provisioned=True ends the request the same way: the request reads
Failed=True with reason MarketRevoked and the job requeues under a
fresh request at your current limit price.
When we take a node out of service¶
A node our health checks cordon is replaced, not reclaimed. The market
stops counting it toward your grant the moment the cordon lands: billing
for it stops, your gang reads short by one node, and a replacement is
granted that inherits the remaining protection window at the locked
rate. The faulty node stays in your cluster, cordoned and tainted, while
we investigate; the pods on it are evicted with a 60-second grace so the
job reschedules them onto the replacement once it joins — the
replacement carries your request's node label, so the pinned pods place
themselves. Nothing changes on the request: it stays Provisioned=True
throughout. In the console the node's tile reads Node Fix In
Progress and the request's row shows Replacement Pending until the
replacement joins. The replacement claim belongs to your cluster, not to
one job: whichever of your pending requests the market seats first takes
it, at the rate your cluster locked for the node that failed, and that
request's row says "protected at the cluster's locked rate" when that
rate is above its own bid. When our checks clear the node it is
uncordoned and returns as idle capacity under the usual idle grace. Cordoning a node yourself
does none of this: an unmarked cordon is read as your own scheduling
choice, the node keeps billing, and no replacement is owed.
Design for reclaims the way you would for any preemptible capacity:
- Checkpoint long-running training; make jobs resumable.
- Treat node-local disk as ephemeral; keep durable state on shared or external storage.
- Watch the market feed and keep your ceiling above the going rate for work that shouldn't be interrupted.
Without the spec below, every evicted pod counts against the Job's
backoffLimit (default 6) and the Job creates replacement pods while
the evicted ones are still terminating — a long-running Job could fail
from reclaims alone. Exempt reclaims from the retry budget with
podReplacementPolicy: Failed (no replacement pod while an evicted pod
is terminating) plus podFailurePolicy Ignore rules for the
DisruptionTarget condition the eviction sets and for your SIGTERM
exit code (143 for sh and any process that dies on the signal;
podFailurePolicy is valid only with restartPolicy: Never):
spec:
podReplacementPolicy: Failed
podFailurePolicy:
rules:
- action: Ignore
onPodConditions:
- type: DisruptionTarget
status: "True"
- action: Ignore
onExitCodes:
operator: In
values: [143]
template:
spec:
restartPolicy: Never
terminationGracePeriodSeconds: 60
A granted Kubernetes node carries the same minimum-duration protection
window as a VM node, anchored at grant — counted from the moment the
node is usable; currently two hours on Kubernetes against one on VM
clusters (minimum-duration protection).
Its first hour is a minimum hold: the node stays in your cluster and
bills for at least one hour even if the job finishes sooner — which also
means a job that crashes restarts on the nodes it already had. The
window belongs to the node grant, not to the job, so it survives job
turnover on the node — and a node that stops responding to the cluster
is owed a replacement for the remainder of its window (a node an
operator moves out is not). The window does not lock the price: while a
job's limit price sits below the price it held when its gang started
(Provisioned=True), its nodes carry no window and are preemptible at
once; raising the price never resets or voids the window.
Exposing services: type: LoadBalancer¶
Serving needs a stable public endpoint that survives node churn. Create
a standard Kubernetes Service with type: LoadBalancer and the
platform provisions an external TCP load balancer for it:
apiVersion: v1
kind: Service
metadata:
name: serve
spec:
type: LoadBalancer
selector:
app: serve
ports:
- port: 443
targetPort: 8443
Within a couple of minutes the Service's status.loadBalancer.ingress
carries a static public IP. The console shows the same endpoint on
the cluster's Workloads page beside the Service. The IP is stable for
the life of the Service. It holds through node churn as the market
moves nodes in and out of your cluster. Deleting the Service, or
changing its type away from LoadBalancer, releases it.
What to know:
- TCP only. UDP ports on a Service are skipped. A Warning Event on the Service says so.
- Every Ready worker backs the endpoint at the Service's NodePort.
Keep
allocateLoadBalancerNodePortsenabled (the default).externalTrafficPolicy: Localworks as expected: health checks route around nodes with no serving pod. - Your pods see the client's real source IP. The path performs no source NAT. IP allowlists and client keyed rate limits work inside your workload.
- Refusals are Events. When the external IP stays
<pending>, runkubectl --context <cluster> describe service <name>. The Event names the reason: unsupported protocol, no allocated NodePort, no Ready workers, or a limit. - Limits exist. Each cluster has a load balancer cap. The platform also has a shared pool limit. The Event names which one you hit; ask support to raise it.
- A Service that sets
spec.loadBalancerClassbelongs to whatever controller you run for that class. The platform leaves it alone.
Load balancer provisioning is enabled per region. Where it is not yet
enabled, the Service stays <pending>; ask support.
The shared volume¶
Clusters provisioned with shared storage come with one 2 TiB shared
volume, mounted at the same path on every node (the control plane
tells you where: /mnt/shared today). Clusters created before
September 2026 carry 1 TiB. The Storage page shows each volume's size. Use it from pods through a
hostPath volume at that path. It is the place for datasets,
checkpoints and outputs that must outlive any one node. The disk on a
node is ephemeral and goes with the node when the market reclaims it.
The volume is created with the cluster and attached to every node that joins, GPU nodes and the CPU worker alike. Where storage billing is enabled for your site, the volume is billed on the bytes it holds and not on its size (billing); your Storage page shows the current charge when that is the case. When you delete the cluster the volume is preserved by default, so your data survives the teardown. A preserved volume stays yours until you delete it.
Seeing and deleting your volume¶
The console's Storage page lists every shared volume your
organization owns. Each volume shows its label (<cluster>-shared),
the storage name under it when the two differ, the cluster it is
attached to (or "preserved, no cluster" once that cluster is gone),
the bytes it holds, and its current charge where storage billing is
enabled for your site.
To delete a volume, use Delete volume on its row. The delete is irreversible: it destroys the volume and every file on it, and the platform keeps no copy. You confirm by typing the volume's storage name back exactly, the name shown under its label. Where storage billing is enabled, billing for the volume stops once the delete is accepted.
Two rules apply:
- An attached volume cannot be deleted. Delete the cluster first. The volume detaches with the cluster and is preserved by default, so it then appears on the Storage page as "preserved, no cluster" and can be deleted from there.
- Deleting a cluster preserves its volume unless you opt in. The cluster delete offers a checkbox, "also delete my storage volume", unchecked by default. Leave it unchecked and the volume and its data survive (and stay billable where storage billing is enabled). Tick it and the volume is destroyed with the cluster. The platform still refuses to destroy a volume while any node holds it; in that case the volume is preserved and can be deleted later.
Listing your volumes is available to automation through the
Kubernetes API. Deleting one
is a console action only: the API refuses a token with
session-required, whatever its scope. A limit price is reversible. Data
destruction is not, so the irreversible delete keeps a human in the
loop with the typed confirmation.
What you'll never see on this surface¶
No node IPs, no slots, no release verbs — node membership is entirely platform-managed. Your controls are the limit price, your job specs (GPU requests, limit price, priority), and your own cluster's scheduling configuration.