# Kubernetes clusters

A Kubernetes tenancy is a **dedicated cluster**: you hold the
kubeconfig, you deploy your own workloads, and the platform moves GPU
nodes in and out of your cluster as your jobs demand them. Nobody runs
your containers for you — the cluster is yours; the platform only decides
how many nodes it has.

Cluster nodes run the Kubernetes agent directly on the metal — there is
no hypervisor layer in the Kubernetes offering.

## Connecting

The kubeconfig is a public client and the cluster's certificate. It
carries no secret. Fetch it with your org API token:

```sh
curl -fsSL -H "Authorization: Bearer $NC_TOKEN" \
  "https://nationalcompute.com/api/k8s/cluster/kubeconfig?cluster=<cluster>" \
  -o ~/.kube/<cluster>.yaml
```

The token names your organization, so the route resolves `<cluster>`
inside it. Another organization's cluster can never answer, whatever its
name. The first `kubectl` call prints a sign in link. A person opens
the link once and approves. After that `kubectl` refreshes its own sign
in while it is in use. A day without a `kubectl` call needs one more
approval: the next call prints a fresh link and waits on it. A
`kubectl` call that seems hung is showing a sign in link — agents run
`kubectl` where its output stays readable while the command runs, and
relay the link to a person.

The setup script installs `kubectl` and the sign in plugin when they are
missing and merges a `<cluster>` context into `~/.kube/config`:

```sh
curl -fsSL https://access.nationalcompute.com/setup.sh | sh -s -- you@example.com <cluster>
```

The script reads the token from `NC_TOKEN` or from
`~/.config/nationalcompute/token` and fetches the kubeconfig by token
when it finds one. Without a token it asks the public onboarding path by
name. That path refuses a name that more than one cluster carries with
`409 ambiguous-cluster` and points at the token route.

The script never waits on the sign in: run without a terminal, it
prints the link and exits while sign-in polls in the background.
Approval lands the token on its own, and
`kubectl --context <cluster> get nodes` then answers at once. A fresh
link in its place means the code expired; the new link replaces the
old.

Workloads and automation authenticate to the cluster with ServiceAccount
tokens rather than a sign in, and every cluster is a public OIDC issuer
for them — see [Authentication](https://docs.nationalcompute.com/authentication.md#kubernetes-service-account-tokens).

## Demand comes from your jobs

There is nothing to declare except your price. The market derives your
demand from the jobs you submit, through Kueue — the queue installed in
your cluster that holds a job until the market grants its nodes:

- **A GPU job is one request for whole nodes.** `kubectl apply` a
  `Job`, `JobSet` or `MPIJob` whose pods request full-node GPU
  multiples; Kueue creates it suspended and files one request for the
  whole gang — every pod the job needs, priced together. When your limit price
  clears, the nodes join and every pod starts in the same moment, on
  exactly those nodes. Nothing else is needed: no queue name, no
  `suspend: true`, and the toleration for the GPU worker taint is
  injected on admission.
- **A job smaller than one node runs on your base load block.** On a
  cluster where the platform has enabled packing onto the
  [base load block](https://docs.nationalcompute.com/market.md#base-load-capacity), a pod may request
  fewer GPUs than one node. Kueue files one request for it. The market
  seats it on a block node beside your other small jobs. It never
  bids. It never runs on a market node. Without a live block it waits.
- **A waiting job has no pods.** `kubectl --context <cluster> get pods` is empty until the
  whole gang is granted; a gang is fully waiting or fully running, never
  "1 of 2". The exceptions are divisible: a `Deployment` replica or a
  bare GPU pod is its own request and exists as a pod held
  `SchedulingGated` until granted.
- **Idle nodes return to the pool.** Once a node's one-hour
  [minimum hold](https://docs.nationalcompute.com/market.md#minimum-duration-protection) has run, a node
  whose job has ended sheds after an idle grace of five minutes, billed
  at that job's rate, unless another job takes it first — the grace
  covers between-job gaps, so batchy workloads don't churn nodes.
- **CPU-only work is outside the market.** A job requesting no GPUs
  runs at once on the cluster's CPU worker (below) and neither bids nor
  holds a GPU node. It passes through the queue without a market check,
  so it can show `Suspended` for a moment.

The result is that capacity — and [billing](https://docs.nationalcompute.com/billing.md) — tracks what
you actually run: when the queue drains, nodes shed and the meter
stops, with no action from you.

### Reading the queue

The queue is read on the Job, never on pods. `kubectl --context <cluster> get job` shows
the Job `Suspended`; `kubectl --context <cluster> get workloads` lists its Workload (Kueue's
queue entry for the job, `<kind>-<jobname>-<hash>`) and whether it is
admitted; `kubectl describe job` carries the market's verdicts as
Events — one per change of reason, then `NodeGranted` per granted
node. The same verdict is the `Provisioned=False` condition of the
ProvisioningRequest — the object Kueue files to ask the market for the
gang's nodes, `<workload>-market-<attempt>` — read with
`kubectl describe provisioningrequest`, and is copied into the
Workload's admission check; once granted, the condition reads
`Provisioning` (`k of N granted nodes ready`) until every node is ready
and it turns `Provisioned=True`. The reason is empty until the market's
first read, one tick after the request appears; the
[limit price API](https://docs.nationalcompute.com/api/k8s-bid.md#queue-events) tables every reason. The
console's Workloads page lists the queued job before it has pods, with
the same reason, and marks a running gang **gang scheduled**.

## The cluster CPU worker

Newly provisioned clusters come with one small CPU-only worker node in
addition to the control plane. It is a plain schedulable node with the
`cpu` role in `kubectl --context <cluster> get nodes`. It carries no GPUs and no taints, so
untolerating pods (queues, controllers, data prep, small services)
land there by default while GPU jobs wait in the queue.

The CPU worker sits outside the market: it never bids, never sheds on
idleness, and is never reclaimed when the market clears above your
ceiling. It is created with the cluster and removed with the cluster.
It cannot be resized or multiplied; for CPU capacity beyond it, run
CPU work on GPU nodes you already hold or contact support.

## The limit price { #the-bid }

[`PUT /api/k8s/bid`](https://docs.nationalcompute.com/api/k8s-bid.md) sets the most you'll pay per
GPU-hour — your limit price. Each request bids its gang's GPU total ×
your limit price, in whole nodes, so a request's exposure is its node
count × `gpus_per_node` × your ceiling. A request smaller than one
node bids nothing. It waits for GPUs of your base load block. A job
carries its own limit price in the label
`nationalcompute.com/limit-price: "<price>"` (USD per
GPU-hour, as a quoted string), which replaces the cluster's for that job; `"0"` is a real
zero limit price — the job waits with the reason `BidTooLow` and never
falls back to the cluster's price. The limit price is read live at every tick for the
request's whole life, waiting or granted: you may change a job's limit
price at any time. An increase is always fine. A decrease below the
price the job held when its gang started (`Provisioned=True`) voids its
protection window ([below](#when-the-market-reclaims-a-node)).

With **no limit price set, your jobs bid $0**, which never clears:
the demand is visible, but nothing is granted until you declare a real
ceiling. A limit price too low to ever clear the market is refused at
write time (`bid-too-low`); a standing one the market moved past reads
`bid_too_low: true` on the limit price read, and its requests wait
with the reason `BidTooLow` — they never fail — and clear at the first
tick after you raise the price. See
[how the bid prices the market](https://docs.nationalcompute.com/api/k8s-bid.md#how-the-bid-prices-the-market)
for the arithmetic.

## The max cluster size

Every cluster has a max cluster size in GPUs — set by the platform,
read-only to you (an edit of the ClusterQueue, or of any other Kueue
object the platform owns, is refused with `StationOwned`), shown on the
cluster's console page and in-cluster as the quota of Kueue's
ClusterQueue `market` (`kubectl --context <cluster> get clusterqueue market`). While the
GPUs of your running and requested jobs stay under it, every new job
files its request at once. Beyond it, jobs wait in your own queue and
file their requests as room frees: the Workload's `QuotaReserved`
condition names the quota as the cause (`kubectl describe workload`),
and the console's Workloads page lists the job waiting, with no request
filed yet. A job larger than the limit on its own never reaches the
market; ask support to raise it.

## Priority orders only your own queue

Priority orders a job against your other jobs, and only while the max
cluster size binds: set the label `kueue.x-k8s.io/priority-class` to
`low`, `normal` or `high` on a suspended Job — a job without the label
ranks below `low` — and the higher class files its request first when
room frees. Every job under the limit reaches the market at once, where
price decides — priority never buys market position over another tenant
and never evicts a running job. On a cluster with a
[base load block](https://docs.nationalcompute.com/market.md#base-load-capacity) the same order decides who takes the
GPUs of the block as they free: the higher class first, oldest first within a
class — so a serving Deployment marked `high` reclaims its node in the block
after a rollout ahead of waiting batch work.

## When the site has no capacity

A shortage is not a pricing problem. When the site runs out of
capacity of the class your jobs request, no bid clears it: the request
stays in the queue, the console's Workloads page shows the verdict as
the reason **capacity unavailable**, and the
[market feed](https://docs.nationalcompute.com/api/market-feed.md)'s ticks read `null` while nothing
clears. Raising your limit price does not start the job sooner. The request's
own reason does not distinguish a shortage — it reads `PendingSupply`,
`Outbid` or `Pending` as usual; the console's Workloads page is where
the shortage verdict shows.

There is nothing to do. Pending demand costs nothing. Leave the limit
price where it is; the platform checks for returning capacity on its own, and
the job starts when capacity returns, with no action from you.

## Launch admission

Submitting a GPU workload is checked at admission on market-managed
clusters — a violation is refused at `kubectl apply` with the exact
numbers in the error, never left silently queued:

- **Full nodes on the market** (`JobSizeTooSmall`). Every pod template
  of a `Job`, `JobSet` or `MPIJob` must request a multiple of
  `gpus_per_node` GPUs. The same rule binds a bare pod. A multi node
  job runs as several full node pods. Templates requesting no GPUs (an
  MPI launcher, for one) are exempt. The market assigns whole nodes
  for strong isolation. A partial node GPU pod is refused rather than
  priced. The exception is a cluster with packing onto the base load
  block enabled ([above](#demand-comes-from-your-jobs)). There a pod
  may request fewer GPUs than one node. Such a pod runs on the block
  only. A pod larger than one node must still request a whole number
  of nodes (`JobSizeNotNodeAligned`).
- **Minimum balance** (`BalanceTooLow`). The org's credit balance must
  cover **two hours** (currently) of the job at your limit price — 2 × limit price ×
  the job's GPU total, counted across `parallelism` (or replicas). The
  denial names the required and current balance. The rule reduces
  churn — a job that takes nodes must be fundable at its own limit price for
  more than minutes. The denial binds at apply only — a job already in
  the queue is never refused later; a shortfall while it waits pends
  the request instead (below).

The market runs the same two-hour balance check again at every tick
while the request waits, in case the situation has changed: a request
the org's balance can no longer fund stays in the queue with the reason
`BalanceTooLow`, whose message names the balance available and what the
request costs ([queue events](https://docs.nationalcompute.com/api/k8s-bid.md#queue-events)); a top-up
clears it at the next tick and nothing fails.

The platform can switch packing onto the block off for a cluster. From
that moment the full node rule binds again at admission: a new pod
smaller than one node is refused with `JobSizeTooSmall`. A Deployment
update and a Job retry create new pods and are refused the same way.
Pods already running are never touched. Resubmit the work as whole
node pods until packing is switched on again.

## When the market reclaims a node

If the market clears above your ceiling (or you withdraw the limit price), a
gang's nodes are reclaimed: the whole job is suspended, its pods are
deleted together, and the job requeues as a new request at your current
limit price — it resumes when granted again, on whatever nodes the
market grants then. A reclaim is never a failed job: the Workload goes
back to pending, and with the spec below in place it does not consume
the Job's `backoffLimit`. On the console's Workloads page the requeued
job's row shows the current episode — queued since the requeue, running
since the last grant — and its Run History the episodes and the totals.

A reclaim caused by a **higher bid from another tenant** against your
standing limit price comes with **one minute's warning**. A reclaim
caused by your own account — your limit price becomes too low to ever
clear, you withdraw or lower it, or your credit balance runs out — is
**immediate**, even when another tenant takes the node: no advance
warning, the drain starts at once, and pods get the same 60-second
eviction grace described below. At notice time the request
gains a `PreemptionNotice=True` condition (reason `MarketPreemption`)
naming the deadline, the node is cordoned and tainted, a
`NodePreempting` event is posted on the node and on each of its pods,
every pod gets a `marketplace.nationalcompute.com/preempt-at`
annotation carrying the deadline (readable in-pod through a
downward-API volume — annotation files update live), and the pods are
evicted through the Eviction API with the window as their grace period
— SIGTERM at notice is your checkpoint signal, and `preempt-at` the
deadline to checkpoint by. PodDisruptionBudgets are honored inside the
window, never past it. At the deadline the whole job is suspended and
requeued: the market revokes the request (it reads `Failed=True`,
reason `MarketRevoked`), and Kueue evicts the Workload, deletes
whatever pods remain and files the new request. Set
`terminationGracePeriodSeconds: 60` on your pods so their own grace
matches the window. A reclaim withdrawn mid-notice uncordons the node,
flips the condition to `PreemptionNotice=False` (reason `Rescinded`)
and posts a `PreemptionRescinded` event on the node.

A granted node that leaves your cluster after the request turns
`Provisioned=True` ends the request the same way: the request reads
`Failed=True` with reason `MarketRevoked` and the job requeues under a
fresh request at your current limit price.

## When we take a node out of service

A node our health checks cordon is replaced, not reclaimed. The market
stops counting it toward your grant the moment the cordon lands: billing
for it stops, your gang reads short by one node, and a replacement is
granted that inherits the remaining protection window at the locked
rate. The faulty node stays in your cluster, cordoned and tainted, while
we investigate; the pods on it are evicted with a 60-second grace so the
job reschedules them onto the replacement once it joins — the
replacement carries your request's node label, so the pinned pods place
themselves. Nothing changes on the request: it stays `Provisioned=True`
throughout. In the console the node's tile reads **Node Fix In
Progress** and the request's row shows **Replacement Pending** until the
replacement joins. The replacement claim belongs to your cluster, not to
one job: whichever of your pending requests the market seats first takes
it, at the rate your cluster locked for the node that failed, and that
request's row says "protected at the cluster's locked rate" when that
rate is above its own bid. When our checks clear the node it is
uncordoned and returns as idle capacity under the usual idle grace. Cordoning a node yourself
does none of this: an unmarked cordon is read as your own scheduling
choice, the node keeps billing, and no replacement is owed.

Design for reclaims the way you would for any preemptible capacity:

- Checkpoint long-running training; make jobs resumable.
- Treat node-local disk as ephemeral; keep durable state on shared or
  external storage.
- Watch the [market feed](https://docs.nationalcompute.com/api/market-feed.md) and keep your ceiling
  above the going rate for work that shouldn't be interrupted.

Without the spec below, every evicted pod counts against the Job's
`backoffLimit` (default 6) and the Job creates replacement pods while
the evicted ones are still terminating — a long-running Job could fail
from reclaims alone. Exempt reclaims from the retry budget with
`podReplacementPolicy: Failed` (no replacement pod while an evicted pod
is terminating) plus `podFailurePolicy` Ignore rules for the
`DisruptionTarget` condition the eviction sets and for your SIGTERM
exit code (143 for `sh` and any process that dies on the signal;
`podFailurePolicy` is valid only with `restartPolicy: Never`):

```yaml
spec:
  podReplacementPolicy: Failed
  podFailurePolicy:
    rules:
      - action: Ignore
        onPodConditions:
          - type: DisruptionTarget
            status: "True"
      - action: Ignore
        onExitCodes:
          operator: In
          values: [143]
  template:
    spec:
      restartPolicy: Never
      terminationGracePeriodSeconds: 60
```

A granted Kubernetes node carries the same minimum-duration protection
window as a VM node, anchored at grant — counted from the moment the
node is usable; currently two hours on Kubernetes against one on VM
clusters ([minimum-duration protection](https://docs.nationalcompute.com/market.md#minimum-duration-protection)).
Its first hour is a minimum hold: the node stays in your cluster and
bills for at least one hour even if the job finishes sooner — which also
means a job that crashes restarts on the nodes it already had. The
window belongs to the node grant, not to the job, so it survives job
turnover on the node — and a node that stops responding to the cluster
is owed a replacement for the remainder of its window (a node an
operator moves out is not). The window does not lock the price: while a
job's limit price sits below the price it held when its gang started
(`Provisioned=True`), its nodes carry no window and are preemptible at
once; raising the price never resets or voids the window.

## Exposing services: `type: LoadBalancer`

Serving needs a stable public endpoint that survives node churn. Create
a standard Kubernetes Service with `type: LoadBalancer` and the
platform provisions an external TCP load balancer for it:

```yaml
apiVersion: v1
kind: Service
metadata:
  name: serve
spec:
  type: LoadBalancer
  selector:
    app: serve
  ports:
    - port: 443
      targetPort: 8443
```

Within a couple of minutes the Service's `status.loadBalancer.ingress`
carries a **static public IP**. The console shows the same endpoint on
the cluster's Workloads page beside the Service. The IP is stable for
the life of the Service. It holds through node churn as the market
moves nodes in and out of your cluster. Deleting the Service, or
changing its `type` away from `LoadBalancer`, releases it.

What to know:

- **TCP only.** UDP ports on a Service are skipped. A Warning Event on
  the Service says so.
- **Every Ready worker backs the endpoint** at the Service's NodePort.
  Keep `allocateLoadBalancerNodePorts` enabled (the default).
  `externalTrafficPolicy: Local` works as expected: health checks
  route around nodes with no serving pod.
- **Your pods see the client's real source IP.** The path performs no
  source NAT. IP allowlists and client keyed rate limits work inside
  your workload.
- **Refusals are Events.** When the external IP stays `<pending>`, run
  `kubectl --context <cluster> describe service <name>`. The Event names the reason:
  unsupported protocol, no allocated NodePort, no Ready workers, or a
  limit.
- **Limits exist.** Each cluster has a load balancer cap. The platform
  also has a shared pool limit. The Event names which one you hit; ask
  support to raise it.
- A Service that sets `spec.loadBalancerClass` belongs to whatever
  controller you run for that class. The platform leaves it alone.

Load balancer provisioning is enabled per region. Where it is not yet
enabled, the Service stays `<pending>`; ask support.

## The shared volume

Clusters provisioned with shared storage come with one **2 TiB shared
volume**, mounted at the same path on every node (the control plane
tells you where: `/mnt/shared` today). Clusters created before
September 2026 carry 1 TiB. The Storage page shows each volume's size. Use it from pods through a
`hostPath` volume at that path. It is the place for datasets,
checkpoints and outputs that must outlive any one node. The disk on a
node is ephemeral and goes with the node when the market reclaims it.

The volume is created with the cluster and attached to every node that
joins, GPU nodes and the CPU worker alike. Where storage billing is
enabled for your site, the volume is billed on the bytes it holds and
not on its size ([billing](https://docs.nationalcompute.com/billing.md)); your Storage
page shows the current charge when that is the case. When you delete
the cluster the volume is preserved by default, so your data survives
the teardown. A preserved volume stays yours until you delete it.

### Seeing and deleting your volume

The console's **Storage** page lists every shared volume your
organization owns. Each volume shows its label (`<cluster>-shared`),
the storage name under it when the two differ, the cluster it is
attached to (or "preserved, no cluster" once that cluster is gone),
the bytes it holds, and its current charge where storage billing is
enabled for your site.

To delete a volume, use **Delete volume** on its row. The delete is
irreversible: it destroys the volume and every file on it, and the
platform keeps no copy. You confirm by typing the volume's storage name
back exactly, the name shown under its label. Where storage billing is
enabled, billing for the volume stops once the delete is accepted.

Two rules apply:

- **An attached volume cannot be deleted.** Delete the cluster first.
  The volume detaches with the cluster and is preserved by default, so
  it then appears on the Storage page as "preserved, no cluster" and
  can be deleted from there.
- **Deleting a cluster preserves its volume unless you opt in.** The
  cluster delete offers a checkbox, "also delete my storage volume",
  unchecked by default. Leave it unchecked and the volume and its data
  survive (and stay billable where storage billing is enabled). Tick
  it and the volume is destroyed with the cluster. The platform still
  refuses to destroy a volume while any node holds it; in that case the
  volume is preserved and can be deleted later.

Listing your volumes is available to automation through the
[Kubernetes API](https://docs.nationalcompute.com/api/k8s-bid.md#shared-storage-volumes). Deleting one
is a console action only: the API refuses a token with
`session-required`, whatever its scope. A limit price is reversible. Data
destruction is not, so the irreversible delete keeps a human in the
loop with the typed confirmation.

## What you'll never see on this surface

No node IPs, no slots, no release verbs — node membership is entirely
platform-managed. Your controls are the limit price, your job specs (GPU
requests, limit price, priority), and your own cluster's scheduling
configuration.
