Skip to content

VM capacity API

Declare how many GPUs your organization wants and the most you'll pay per GPU-hour; the market grants and reclaims nodes to match. Granted nodes are yours over ssh until the market reclaims them or you scale down.

Contract: GET /api/vm/openapi.json.

Route Auth What it does
GET /api/vm/capacity capacity:read current declaration, allocations, nominations
PUT /api/vm/capacity capacity:write declare or update capacity; also scales down
POST /api/vm/capacity/swap capacity:write release ONE node and let the market refill its slot
GET /api/vm/activity token what you declared and what the platform did about it, newest first
GET /api/vm/history token declared versus granted, and your ceiling versus the clearing rate, over time
GET /api/vm/orgs/{org}/ssh-keys token list the org's ssh public keys
POST /api/vm/orgs/{org}/ssh-keys token add an ssh public key
DELETE /api/vm/orgs/{org}/ssh-keys/{key_id} token revoke a key

{org} is your organization id — returned as org_id by /api/whoami, and shown on the console's API → Docs page.

Before you declare: register an ssh key

A grant installs your org's registered ssh keys on the node — with none registered the node would be unreachable, so a capacity raise without a key is refused (422 no-ssh-keys):

curl -X POST -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"name": "ci", "public_key": "ssh-ed25519 AAAA... ci@example"}' \
  https://nationalcompute.com/api/vm/orgs/$ORG_ID/ssh-keys

Keys install on future grants only — adding or revoking a key never touches nodes you already hold. A key can be scoped to one VM cluster with the optional cluster field; omitted means org-wide. The console's SSH Keys page manages the same list.

Declaring capacity

curl -X PUT -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "max_gpus": 16,
        "max_price_per_gpu_hour": <price>,
        "expected_version": 7
      }' \
  https://nationalcompute.com/api/vm/capacity
Field Meaning
max_gpus how many GPUs to hold — a multiple of gpus_per_node (whole nodes; a non-multiple is a 422, never rounded). Lowering it is how you scale down; 0 releases everything.
max_price_per_gpu_hour your ceiling per GPU-hour, USD, whole cents. A whole node clears at gpus_per_node × this. The billed rate is the clearing rate — often lower, never higher.
release optional: node public IPs to release first on scale-down. Replaces the standing nomination set; omitted = keep current; [] = clear all (a nomination tied to an in-flight swap persists until the swap lands); an IP you don't hold is a 409 and nothing is written.
expected_version compare-and-set guard — see conventions.
cluster only needed with an org-wide token when the org holds several VM clusters.

Raising max_gpus requires the market-managed posture (409 static-posture otherwise — on an operator-managed cluster nothing fills new capacity, so a raise would pend forever). Lowering and release always work.

The MCP server's bid_write kind=vm is the same write in two phases: max_gpus, max_price_per_gpu_hour and the optional release list. The preview echoes the nominated nodes, the standing nominations they replace, and flags any entry that matches no node you hold.

Reading state

GET /api/vm/capacity returns the full picture:

Field Meaning
version echo back as expected_version on writes
writes_armed whether the capacity PUT is accepted for this cluster's island at all. false means every PUT answers 503 capacity-writes-disabled while reads keep serving
gpus_per_node the node ↔ GPU conversion — read it, never hardcode it
gpu_vendor, gpu_model the maker and the GPU class this declaration buys. gpu_model is null while the platform has not set one
declared your standing declaration (ceiling, max_gpus, bid_too_low, capacity_unavailable, who updated it and when); null until the first PUT
pending declared capacity (in whole nodes) not yet covered by an open allocation
allocations[] the nodes you hold: public_ip (what you ssh to), state (provisioning / active / reclaiming), slot, swap
planned_release[] what a shrink will destroy, as of this read
nominations[] your standing release nominations, echoed as public IPs
swaps[] public IPs with a swap in flight
billing_hold your org's billing hold state, or null. When set: state (hold or grace), reason, and for hold a stage — notice (new capacity paused) or enforce (capacity being reclaimed) — plus enforce_after_ms. A hold never alters declared: the response always echoes exactly what you asked for.
spend_limit_enforced, billing_hold_enforced whether your organization's monthly spend limit and a balance at or below zero hold its capacity at all. false means the platform records the figure and does not act on it for your organization's clusters. The billing summary carries the same pair
environment site-specific onboarding facts (ssh user, environment variables, storage paths)

A representative response:

{
  "cluster": "acme-research",
  "version": 8,
  "writes_armed": true,
  "gpus_per_node": 8,
  "gpu_vendor": "…",
  "gpu_model": "…",
  "declared": {
    "max_price_per_gpu_hour": …,
    "max_gpus": 24,
    "bid_too_low": false,
    "capacity_unavailable": null,
    "updated_at": "2026-08-26T14:02:11Z",
    "updated_by": "api-token:ci"
  },
  "pending": 1,
  "allocations": [
    {"public_ip": "203.0.113.7",  "granted_at": "2026-08-25T09:14:02Z",
     "state": "active",       "slot": 0, "swap": false},
    {"public_ip": "203.0.113.21", "granted_at": "2026-08-26T13:58:40Z",
     "state": "provisioning", "slot": 1, "swap": false}
  ],
  "planned_release": [],
  "nominations": [],
  "swaps": [],
  "billing_hold": null,
  "environment": {"ssh_user": "tenant"}
}

Here the org wants 24 GPUs (3 nodes): one node is active, one is still provisioning, and one slot is pending — declared but not yet granted by the market.

Allocation lifecycle

Each held node moves through three states:

  1. provisioning — the market granted the slot and the node is being prepared; public_ip may still be null. Your registered ssh keys are installed during this phase.
  2. active — the node is yours over ssh at its public_ip. A fresh grant carries the minimum-duration protection window; after it lapses the node competes in the market normally.
  3. reclaiming — the node is being taken back: the market cleared above your ceiling, you scaled down, or a swap is landing. The node will be destroyed — anything on local disk is lost, so treat node state as ephemeral and keep durable data elsewhere.

You can see a reclaim coming from three signals on the capacity read: planned_release (what a shrink will destroy as of this read), nominations (your own standing choices for who goes first), and swaps (destroys in flight). Billing runs from grant completion to the start of reclaim.

Swapping a node

Releases one node — the node is destroyed and all data on it is lost — while its capacity slot stays declared, so the market grants a replacement once the release lands:

curl -X POST -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"node": "203.0.113.7"}' \
  https://nationalcompute.com/api/vm/capacity/swap

The release happens when the swap actuates; a swap that can't actuate within its freshness window (currently an hour) lapses and the node simply stays yours. The replacement depends on supply and your limit price. Retrying the same node while it swaps is an idempotent 200; swaps on different nodes stack, spaced by the shared write cadence. To shrink instead of swap, lower max_gpus on the capacity PUT (with release to pick the victim).

History and activity

GET /api/vm/activity is the Instances page's History: applied declarations and allocation events, newest first, 30 per page. Each source is read 400 rows deep before the merge, so very old history ages out. Pages carry 30 entries (?page=, from 0; more says whether a next page exists). An entry of kind: request is a declaration the platform applied, with its actor, max_gpus, and max_price_per_gpu_hour; who is user for your organization's own declares and system for a platform override. An entry of kind: action is an allocation event: a node granted or released, named by its public IP in note, or a grant attempt unwound. ?who=user or ?who=system filters. Capacity reads never appear.

GET /api/vm/history?hours=24 is the Instances page's charts as bucketed series over a window of 0.5 to 8760 hours: t0 (epoch seconds of the first bucket), step (bucket width in seconds, max(60, span / 400)), max_nodes (the declared node count in force: max_gpus divided by gpus_per_node; null before the first declaration), granted (nodes held; 0 when none, never null), max_price (your declared ceiling) and price (the market's clearing rate), both in USD per node-hour. Divide by gpus_per_node to read them per GPU-hour. A null price is a bucket where nothing was declared or nothing cleared.

Errors

Status When
401 / 403 / 404 see authentication
409 version-mismatch, unknown-node (a release/swap entry that matches no held node), ambiguous-cluster, static-posture; swaps can also refuse swap-disabled, cluster-gated, or node-pinned
422 validation — price-precision, bid-too-low (any max_price_per_gpu_hour set too low to ever clear the market, raising or lowering max_gpus; a standing price left unchanged can still lower max_gpus or release, so a market that moved past your ceiling never traps held capacity), price-above-maximum (a max_price_per_gpu_hour over the cap the GET echoes as max_price_cap_per_gpu_hour; the response carries maximum), no-ssh-keys, a non-multiple max_gpus; the detail says exactly what
429 faster than the per cluster write cadence, or past the token's request budget for the minute; wait retry_after_s (the Retry-After header carries the same figure)
503 temporary platform condition — reads keep serving, retry writes later

The error reference covers every code with handling advice.

The too-low check runs at write time only. When market conditions move past a standing declaration, the ask keeps losing and its unfilled capacity never arrives. The GET reports this as declared.bid_too_low: true; check the market feed and raise max_price_per_gpu_hour to fill again.

declared.capacity_unavailable is the other reason a standing ask can sit unfilled, and it is not about price. true means the site has no capacity of the requested class right now, so raising the price will not fill pending slots sooner. The system retries automatically and slots fill when capacity returns. false means capacity is flowing normally. null means no verdict (the detector is not armed at this site). The same verdict reaches Kubernetes clusters on the console's Workloads page — see when the site has no capacity.