Skip to content

Kubernetes clusters

A Kubernetes tenancy is a dedicated cluster: you hold the kubeconfig, you deploy your own workloads, and the platform moves GPU nodes in and out of your cluster as your jobs demand them. Nobody runs your containers for you — the cluster is yours; the platform only decides how many nodes it has.

Cluster nodes run the Kubernetes agent directly on the metal — there is no hypervisor layer in the Kubernetes offering.

Connecting

The kubeconfig is a public client and the cluster's certificate. It carries no secret. Fetch it with your org API token:

curl -fsSL -H "Authorization: Bearer $NC_TOKEN" \
  "https://nationalcompute.com/api/k8s/cluster/kubeconfig?cluster=<cluster>" \
  -o ~/.kube/<cluster>.yaml

The token names your organization, so the route resolves <cluster> inside it. Another organization's cluster can never answer, whatever its name. The first kubectl call prints a sign in link. A person opens the link once and approves. After that kubectl refreshes its own sign in while it is in use. A day without a kubectl call needs one more approval: the next call prints a fresh link and waits on it. A kubectl call that seems hung is showing a sign in link — agents run kubectl where its output stays readable while the command runs, and relay the link to a person.

The setup script installs kubectl and the sign in plugin when they are missing and merges a <cluster> context into ~/.kube/config:

curl -fsSL https://access.nationalcompute.com/setup.sh | sh -s -- you@example.com <cluster>

The script reads the token from NC_TOKEN or from ~/.config/nationalcompute/token and fetches the kubeconfig by token when it finds one. Without a token it asks the public onboarding path by name. That path refuses a name that more than one cluster carries with 409 ambiguous-cluster and points at the token route.

The script never waits on the sign in: run without a terminal, it prints the link and exits while sign-in polls in the background. Approval lands the token on its own, and kubectl --context <cluster> get nodes then answers at once. A fresh link in its place means the code expired; the new link replaces the old.

Workloads and automation authenticate to the cluster with ServiceAccount tokens rather than a sign in, and every cluster is a public OIDC issuer for them — see Authentication.

Demand comes from your jobs

There is nothing to declare except your price. The market derives your demand from the jobs you submit, through Kueue — the queue installed in your cluster that holds a job until the market grants its nodes:

  • A GPU job is one request for whole nodes. kubectl apply a Job, JobSet or MPIJob whose pods request full-node GPU multiples; Kueue creates it suspended and files one request for the whole gang — every pod the job needs, priced together. When your limit price clears, the nodes join and every pod starts in the same moment, on exactly those nodes. Nothing else is needed: no queue name, no suspend: true, and the toleration for the GPU worker taint is injected on admission.
  • A job smaller than one node runs on your base load block. On a cluster where the platform has enabled packing onto the base load block, a pod may request fewer GPUs than one node. Kueue files one request for it. The market seats it on a block node beside your other small jobs. It never bids. It never runs on a market node. Without a live block it waits.
  • A waiting job has no pods. kubectl --context <cluster> get pods is empty until the whole gang is granted; a gang is fully waiting or fully running, never "1 of 2". The exceptions are divisible: a Deployment replica or a bare GPU pod is its own request and exists as a pod held SchedulingGated until granted.
  • Idle nodes return to the pool. Once a node's one-hour minimum hold has run, a node whose job has ended sheds after an idle grace of five minutes, billed at that job's rate, unless another job takes it first — the grace covers between-job gaps, so batchy workloads don't churn nodes.
  • CPU-only work is outside the market. A job requesting no GPUs runs at once on the cluster's CPU worker (below) and neither bids nor holds a GPU node. It passes through the queue without a market check, so it can show Suspended for a moment.

The result is that capacity — and billing — tracks what you actually run: when the queue drains, nodes shed and the meter stops, with no action from you.

Reading the queue

The queue is read on the Job, never on pods. kubectl --context <cluster> get job shows the Job Suspended; kubectl --context <cluster> get workloads lists its Workload (Kueue's queue entry for the job, <kind>-<jobname>-<hash>) and whether it is admitted; kubectl describe job carries the market's verdicts as Events — one per change of reason, then NodeGranted per granted node. The same verdict is the Provisioned=False condition of the ProvisioningRequest — the object Kueue files to ask the market for the gang's nodes, <workload>-market-<attempt> — read with kubectl describe provisioningrequest, and is copied into the Workload's admission check; once granted, the condition reads Provisioning (k of N granted nodes ready) until every node is ready and it turns Provisioned=True. The reason is empty until the market's first read, one tick after the request appears; the limit price API tables every reason. The console's Workloads page lists the queued job before it has pods, with the same reason, and marks a running gang gang scheduled.

The cluster CPU worker

Newly provisioned clusters come with one small CPU-only worker node in addition to the control plane. It is a plain schedulable node with the cpu role in kubectl --context <cluster> get nodes. It carries no GPUs and no taints, so untolerating pods (queues, controllers, data prep, small services) land there by default while GPU jobs wait in the queue.

The CPU worker sits outside the market: it never bids, never sheds on idleness, and is never reclaimed when the market clears above your ceiling. It is created with the cluster and removed with the cluster. It cannot be resized or multiplied; for CPU capacity beyond it, run CPU work on GPU nodes you already hold or contact support.

The limit price

PUT /api/k8s/bid sets the most you'll pay per GPU-hour — your limit price. Each request bids its gang's GPU total × your limit price, in whole nodes, so a request's exposure is its node count × gpus_per_node × your ceiling. A request smaller than one node bids nothing. It waits for GPUs of your base load block. A job carries its own limit price in the label nationalcompute.com/limit-price: "<price>" (USD per GPU-hour, as a quoted string), which replaces the cluster's for that job; "0" is a real zero limit price — the job waits with the reason BidTooLow and never falls back to the cluster's price. The limit price is read live at every tick for the request's whole life, waiting or granted: you may change a job's limit price at any time. An increase is always fine. A decrease below the price the job held when its gang started (Provisioned=True) voids its protection window (below).

With no limit price set, your jobs bid $0, which never clears: the demand is visible, but nothing is granted until you declare a real ceiling. A limit price too low to ever clear the market is refused at write time (bid-too-low); a standing one the market moved past reads bid_too_low: true on the limit price read, and its requests wait with the reason BidTooLow — they never fail — and clear at the first tick after you raise the price. See how the bid prices the market for the arithmetic.

The max cluster size

Every cluster has a max cluster size in GPUs — set by the platform, read-only to you (an edit of the ClusterQueue, or of any other Kueue object the platform owns, is refused with StationOwned), shown on the cluster's console page and in-cluster as the quota of Kueue's ClusterQueue market (kubectl --context <cluster> get clusterqueue market). While the GPUs of your running and requested jobs stay under it, every new job files its request at once. Beyond it, jobs wait in your own queue and file their requests as room frees: the Workload's QuotaReserved condition names the quota as the cause (kubectl describe workload), and the console's Workloads page lists the job waiting, with no request filed yet. A job larger than the limit on its own never reaches the market; ask support to raise it.

Priority orders only your own queue

Priority orders a job against your other jobs, and only while the max cluster size binds: set the label kueue.x-k8s.io/priority-class to low, normal or high on a suspended Job — a job without the label ranks below low — and the higher class files its request first when room frees. Every job under the limit reaches the market at once, where price decides — priority never buys market position over another tenant and never evicts a running job. On a cluster with a base load block the same order decides who takes the GPUs of the block as they free: the higher class first, oldest first within a class — so a serving Deployment marked high reclaims its node in the block after a rollout ahead of waiting batch work.

When the site has no capacity

A shortage is not a pricing problem. When the site runs out of capacity of the class your jobs request, no bid clears it: the request stays in the queue, the console's Workloads page shows the verdict as the reason capacity unavailable, and the market feed's ticks read null while nothing clears. Raising your limit price does not start the job sooner. The request's own reason does not distinguish a shortage — it reads PendingSupply, Outbid or Pending as usual; the console's Workloads page is where the shortage verdict shows.

There is nothing to do. Pending demand costs nothing. Leave the limit price where it is; the platform checks for returning capacity on its own, and the job starts when capacity returns, with no action from you.

Launch admission

Submitting a GPU workload is checked at admission on market-managed clusters — a violation is refused at kubectl apply with the exact numbers in the error, never left silently queued:

  • Full nodes on the market (JobSizeTooSmall). Every pod template of a Job, JobSet or MPIJob must request a multiple of gpus_per_node GPUs. The same rule binds a bare pod. A multi node job runs as several full node pods. Templates requesting no GPUs (an MPI launcher, for one) are exempt. The market assigns whole nodes for strong isolation. A partial node GPU pod is refused rather than priced. The exception is a cluster with packing onto the base load block enabled (above). There a pod may request fewer GPUs than one node. Such a pod runs on the block only. A pod larger than one node must still request a whole number of nodes (JobSizeNotNodeAligned).
  • Minimum balance (BalanceTooLow). The org's credit balance must cover two hours (currently) of the job at your limit price — 2 × limit price × the job's GPU total, counted across parallelism (or replicas). The denial names the required and current balance. The rule reduces churn — a job that takes nodes must be fundable at its own limit price for more than minutes. The denial binds at apply only — a job already in the queue is never refused later; a shortfall while it waits pends the request instead (below).

The market runs the same two-hour balance check again at every tick while the request waits, in case the situation has changed: a request the org's balance can no longer fund stays in the queue with the reason BalanceTooLow, whose message names the balance available and what the request costs (queue events); a top-up clears it at the next tick and nothing fails.

The platform can switch packing onto the block off for a cluster. From that moment the full node rule binds again at admission: a new pod smaller than one node is refused with JobSizeTooSmall. A Deployment update and a Job retry create new pods and are refused the same way. Pods already running are never touched. Resubmit the work as whole node pods until packing is switched on again.

When the market reclaims a node

If the market clears above your ceiling (or you withdraw the limit price), a gang's nodes are reclaimed: the whole job is suspended, its pods are deleted together, and the job requeues as a new request at your current limit price — it resumes when granted again, on whatever nodes the market grants then. A reclaim is never a failed job: the Workload goes back to pending, and with the spec below in place it does not consume the Job's backoffLimit. On the console's Workloads page the requeued job's row shows the current episode — queued since the requeue, running since the last grant — and its Run History the episodes and the totals.

A reclaim caused by a higher bid from another tenant against your standing limit price comes with one minute's warning. A reclaim caused by your own account — your limit price becomes too low to ever clear, you withdraw or lower it, or your credit balance runs out — is immediate, even when another tenant takes the node: no advance warning, the drain starts at once, and pods get the same 60-second eviction grace described below. At notice time the request gains a PreemptionNotice=True condition (reason MarketPreemption) naming the deadline, the node is cordoned and tainted, a NodePreempting event is posted on the node and on each of its pods, every pod gets a marketplace.nationalcompute.com/preempt-at annotation carrying the deadline (readable in-pod through a downward-API volume — annotation files update live), and the pods are evicted through the Eviction API with the window as their grace period — SIGTERM at notice is your checkpoint signal, and preempt-at the deadline to checkpoint by. PodDisruptionBudgets are honored inside the window, never past it. At the deadline the whole job is suspended and requeued: the market revokes the request (it reads Failed=True, reason MarketRevoked), and Kueue evicts the Workload, deletes whatever pods remain and files the new request. Set terminationGracePeriodSeconds: 60 on your pods so their own grace matches the window. A reclaim withdrawn mid-notice uncordons the node, flips the condition to PreemptionNotice=False (reason Rescinded) and posts a PreemptionRescinded event on the node.

A granted node that leaves your cluster after the request turns Provisioned=True ends the request the same way: the request reads Failed=True with reason MarketRevoked and the job requeues under a fresh request at your current limit price.

When we take a node out of service

A node our health checks cordon is replaced, not reclaimed. The market stops counting it toward your grant the moment the cordon lands: billing for it stops, your gang reads short by one node, and a replacement is granted that inherits the remaining protection window at the locked rate. The faulty node stays in your cluster, cordoned and tainted, while we investigate; the pods on it are evicted with a 60-second grace so the job reschedules them onto the replacement once it joins — the replacement carries your request's node label, so the pinned pods place themselves. Nothing changes on the request: it stays Provisioned=True throughout. In the console the node's tile reads Node Fix In Progress and the request's row shows Replacement Pending until the replacement joins. The replacement claim belongs to your cluster, not to one job: whichever of your pending requests the market seats first takes it, at the rate your cluster locked for the node that failed, and that request's row says "protected at the cluster's locked rate" when that rate is above its own bid. When our checks clear the node it is uncordoned and returns as idle capacity under the usual idle grace. Cordoning a node yourself does none of this: an unmarked cordon is read as your own scheduling choice, the node keeps billing, and no replacement is owed.

Design for reclaims the way you would for any preemptible capacity:

  • Checkpoint long-running training; make jobs resumable.
  • Treat node-local disk as ephemeral; keep durable state on shared or external storage.
  • Watch the market feed and keep your ceiling above the going rate for work that shouldn't be interrupted.

Without the spec below, every evicted pod counts against the Job's backoffLimit (default 6) and the Job creates replacement pods while the evicted ones are still terminating — a long-running Job could fail from reclaims alone. Exempt reclaims from the retry budget with podReplacementPolicy: Failed (no replacement pod while an evicted pod is terminating) plus podFailurePolicy Ignore rules for the DisruptionTarget condition the eviction sets and for your SIGTERM exit code (143 for sh and any process that dies on the signal; podFailurePolicy is valid only with restartPolicy: Never):

spec:
  podReplacementPolicy: Failed
  podFailurePolicy:
    rules:
      - action: Ignore
        onPodConditions:
          - type: DisruptionTarget
            status: "True"
      - action: Ignore
        onExitCodes:
          operator: In
          values: [143]
  template:
    spec:
      restartPolicy: Never
      terminationGracePeriodSeconds: 60

A granted Kubernetes node carries the same minimum-duration protection window as a VM node, anchored at grant — counted from the moment the node is usable; currently two hours on Kubernetes against one on VM clusters (minimum-duration protection). Its first hour is a minimum hold: the node stays in your cluster and bills for at least one hour even if the job finishes sooner — which also means a job that crashes restarts on the nodes it already had. The window belongs to the node grant, not to the job, so it survives job turnover on the node — and a node that stops responding to the cluster is owed a replacement for the remainder of its window (a node an operator moves out is not). The window does not lock the price: while a job's limit price sits below the price it held when its gang started (Provisioned=True), its nodes carry no window and are preemptible at once; raising the price never resets or voids the window.

Exposing services: type: LoadBalancer

Serving needs a stable public endpoint that survives node churn. Create a standard Kubernetes Service with type: LoadBalancer and the platform provisions an external TCP load balancer for it:

apiVersion: v1
kind: Service
metadata:
  name: serve
spec:
  type: LoadBalancer
  selector:
    app: serve
  ports:
    - port: 443
      targetPort: 8443

Within a couple of minutes the Service's status.loadBalancer.ingress carries a static public IP. The console shows the same endpoint on the cluster's Workloads page beside the Service. The IP is stable for the life of the Service. It holds through node churn as the market moves nodes in and out of your cluster. Deleting the Service, or changing its type away from LoadBalancer, releases it.

What to know:

  • TCP only. UDP ports on a Service are skipped. A Warning Event on the Service says so.
  • Every Ready worker backs the endpoint at the Service's NodePort. Keep allocateLoadBalancerNodePorts enabled (the default). externalTrafficPolicy: Local works as expected: health checks route around nodes with no serving pod.
  • Your pods see the client's real source IP. The path performs no source NAT. IP allowlists and client keyed rate limits work inside your workload.
  • Refusals are Events. When the external IP stays <pending>, run kubectl --context <cluster> describe service <name>. The Event names the reason: unsupported protocol, no allocated NodePort, no Ready workers, or a limit.
  • Limits exist. Each cluster has a load balancer cap. The platform also has a shared pool limit. The Event names which one you hit; ask support to raise it.
  • A Service that sets spec.loadBalancerClass belongs to whatever controller you run for that class. The platform leaves it alone.

Load balancer provisioning is enabled per region. Where it is not yet enabled, the Service stays <pending>; ask support.

The shared volume

Clusters provisioned with shared storage come with one 2 TiB shared volume, mounted at the same path on every node (the control plane tells you where: /mnt/shared today). Clusters created before September 2026 carry 1 TiB. The Storage page shows each volume's size. Use it from pods through a hostPath volume at that path. It is the place for datasets, checkpoints and outputs that must outlive any one node. The disk on a node is ephemeral and goes with the node when the market reclaims it.

The volume is created with the cluster and attached to every node that joins, GPU nodes and the CPU worker alike. Where storage billing is enabled for your site, the volume is billed on the bytes it holds and not on its size (billing); your Storage page shows the current charge when that is the case. When you delete the cluster the volume is preserved by default, so your data survives the teardown. A preserved volume stays yours until you delete it.

Seeing and deleting your volume

The console's Storage page lists every shared volume your organization owns. Each volume shows its label (<cluster>-shared), the storage name under it when the two differ, the cluster it is attached to (or "preserved, no cluster" once that cluster is gone), the bytes it holds, and its current charge where storage billing is enabled for your site.

To delete a volume, use Delete volume on its row. The delete is irreversible: it destroys the volume and every file on it, and the platform keeps no copy. You confirm by typing the volume's storage name back exactly, the name shown under its label. Where storage billing is enabled, billing for the volume stops once the delete is accepted.

Two rules apply:

  • An attached volume cannot be deleted. Delete the cluster first. The volume detaches with the cluster and is preserved by default, so it then appears on the Storage page as "preserved, no cluster" and can be deleted from there.
  • Deleting a cluster preserves its volume unless you opt in. The cluster delete offers a checkbox, "also delete my storage volume", unchecked by default. Leave it unchecked and the volume and its data survive (and stay billable where storage billing is enabled). Tick it and the volume is destroyed with the cluster. The platform still refuses to destroy a volume while any node holds it; in that case the volume is preserved and can be deleted later.

Listing your volumes is available to automation through the Kubernetes API. Deleting one is a console action only: the API refuses a token with session-required, whatever its scope. A limit price is reversible. Data destruction is not, so the irreversible delete keeps a human in the loop with the typed confirmation.

What you'll never see on this surface

No node IPs, no slots, no release verbs — node membership is entirely platform-managed. Your controls are the limit price, your job specs (GPU requests, limit price, priority), and your own cluster's scheduling configuration.