# VM capacity API

Declare how many GPUs your organization wants and the most you'll pay
per GPU-hour; the market grants and reclaims nodes to match. Granted
nodes are yours over ssh until the market reclaims them or you scale
down.

Contract: [`GET /api/vm/openapi.json`](https://nationalcompute.com/api/vm/openapi.json).

| Route | Auth | What it does |
|---|---|---|
| `GET /api/vm/capacity` | `capacity:read` | current declaration, allocations, nominations |
| `PUT /api/vm/capacity` | `capacity:write` | declare or update capacity; also scales down |
| `POST /api/vm/capacity/swap` | `capacity:write` | release ONE node and let the market refill its slot |
| `GET /api/vm/activity` | token | what you declared and what the platform did about it, newest first |
| `GET /api/vm/history` | token | declared versus granted, and your ceiling versus the clearing rate, over time |
| `GET /api/vm/orgs/{org}/ssh-keys` | token | list the org's ssh public keys |
| `POST /api/vm/orgs/{org}/ssh-keys` | token | add an ssh public key |
| `DELETE /api/vm/orgs/{org}/ssh-keys/{key_id}` | token | revoke a key |

`{org}` is your organization id — returned as `org_id` by
[`/api/whoami`](https://docs.nationalcompute.com/authentication.md#discovering-what-a-token-can-address),
and shown on the console's **API → Docs** page.

## Before you declare: register an ssh key

A grant installs your org's registered ssh keys on the node — with none
registered the node would be unreachable, so a capacity raise without a
key is refused (`422 no-ssh-keys`):

```sh
curl -X POST -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"name": "ci", "public_key": "ssh-ed25519 AAAA... ci@example"}' \
  https://nationalcompute.com/api/vm/orgs/$ORG_ID/ssh-keys
```

Keys install on **future** grants only — adding or revoking a key never
touches nodes you already hold. A key can be scoped to one VM cluster
with the optional `cluster` field; omitted means org-wide. The console's
**SSH Keys** page manages the same list.

## Declaring capacity

```sh
curl -X PUT -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "max_gpus": 16,
        "max_price_per_gpu_hour": <price>,
        "expected_version": 7
      }' \
  https://nationalcompute.com/api/vm/capacity
```

| Field | Meaning |
|---|---|
| `max_gpus` | how many GPUs to hold — a multiple of `gpus_per_node` (whole nodes; a non-multiple is a `422`, never rounded). Lowering it is how you scale down; `0` releases everything. |
| `max_price_per_gpu_hour` | your ceiling per GPU-hour, USD, whole cents. A whole node clears at `gpus_per_node` × this. The billed rate is the clearing rate — often lower, never higher. |
| `release` | optional: node public IPs to release first on scale-down. **Replaces** the standing nomination set; omitted = keep current; `[]` = clear all (a nomination tied to an in-flight swap persists until the swap lands); an IP you don't hold is a `409` and nothing is written. |
| `expected_version` | compare-and-set guard — see [conventions](https://docs.nationalcompute.com/api/index.md). |
| `cluster` | only needed with an org-wide token when the org holds several VM clusters. |

Raising `max_gpus` requires the market-managed posture (`409
static-posture` otherwise — on an operator-managed cluster nothing fills
new capacity, so a raise would pend forever). Lowering and `release`
always work.

The [MCP server](https://docs.nationalcompute.com/api/mcp.md)'s `bid_write kind=vm` is the same write in two
phases: `max_gpus`, `max_price_per_gpu_hour` and the optional `release`
list. The preview echoes the nominated nodes, the standing nominations
they replace, and flags any entry that matches no node you hold.

## Reading state

`GET /api/vm/capacity` returns the full picture:

| Field | Meaning |
|---|---|
| `version` | echo back as `expected_version` on writes |
| `writes_armed` | whether the capacity PUT is accepted for this cluster's island at all. `false` means every PUT answers `503 capacity-writes-disabled` while reads keep serving |
| `gpus_per_node` | the node ↔ GPU conversion — read it, never hardcode it |
| `gpu_vendor`, `gpu_model` | the maker and the GPU class this declaration buys. `gpu_model` is `null` while the platform has not set one |
| `declared` | your standing declaration (ceiling, `max_gpus`, `bid_too_low`, `capacity_unavailable`, who updated it and when); `null` until the first PUT |
| `pending` | declared capacity (in whole nodes) not yet covered by an open allocation |
| `allocations[]` | the nodes you hold: `public_ip` (what you ssh to), `state` (`provisioning` / `active` / `reclaiming`), `slot`, `swap` |
| `planned_release[]` | what a shrink will destroy, as of this read |
| `nominations[]` | your standing release nominations, echoed as public IPs |
| `swaps[]` | public IPs with a swap in flight |
| `billing_hold` | your org's billing hold state, or `null`. When set: `state` (`hold` or `grace`), `reason`, and for `hold` a `stage` — `notice` (new capacity paused) or `enforce` (capacity being reclaimed) — plus `enforce_after_ms`. A hold never alters `declared`: the response always echoes exactly what you asked for. |
| `spend_limit_enforced`, `billing_hold_enforced` | whether your organization's monthly spend limit and a balance at or below zero hold its capacity at all. `false` means the platform records the figure and does not act on it for your organization's clusters. The [billing summary](https://docs.nationalcompute.com/api/billing.md#balance-and-totals) carries the same pair |
| `environment` | site-specific onboarding facts (ssh user, environment variables, storage paths) |

A representative response:

```json
{
  "cluster": "acme-research",
  "version": 8,
  "writes_armed": true,
  "gpus_per_node": 8,
  "gpu_vendor": "…",
  "gpu_model": "…",
  "declared": {
    "max_price_per_gpu_hour": …,
    "max_gpus": 24,
    "bid_too_low": false,
    "capacity_unavailable": null,
    "updated_at": "2026-08-26T14:02:11Z",
    "updated_by": "api-token:ci"
  },
  "pending": 1,
  "allocations": [
    {"public_ip": "203.0.113.7",  "granted_at": "2026-08-25T09:14:02Z",
     "state": "active",       "slot": 0, "swap": false},
    {"public_ip": "203.0.113.21", "granted_at": "2026-08-26T13:58:40Z",
     "state": "provisioning", "slot": 1, "swap": false}
  ],
  "planned_release": [],
  "nominations": [],
  "swaps": [],
  "billing_hold": null,
  "environment": {"ssh_user": "tenant"}
}
```

Here the org wants 24 GPUs (3 nodes): one node is active, one is still
provisioning, and one slot is `pending` — declared but not yet granted
by the market.

## Allocation lifecycle

Each held node moves through three states:

1. **`provisioning`** — the market granted the slot and the node is
   being prepared; `public_ip` may still be `null`. Your registered ssh
   keys are installed during this phase.
2. **`active`** — the node is yours over ssh at its `public_ip`. A
   fresh grant carries the [minimum-duration protection
   window](https://docs.nationalcompute.com/market.md#minimum-duration-protection); after it lapses
   the node competes in the market normally.
3. **`reclaiming`** — the node is being taken back: the market cleared
   above your ceiling, you scaled down, or a swap is landing. The node
   will be **destroyed** — anything on local disk is lost, so treat
   node state as ephemeral and keep durable data elsewhere.

You can see a reclaim coming from three signals on the capacity read:
`planned_release` (what a shrink will destroy as of this read),
`nominations` (your own standing choices for who goes first), and
`swaps` (destroys in flight). [Billing](https://docs.nationalcompute.com/billing.md) runs from grant
completion to the start of reclaim.

## Swapping a node

Releases one node — **the node is destroyed and all data on it is
lost** — while its capacity slot stays declared, so the market grants a
replacement once the release lands:

```sh
curl -X POST -H "Authorization: Bearer $NC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"node": "203.0.113.7"}' \
  https://nationalcompute.com/api/vm/capacity/swap
```

The release happens when the swap actuates; a swap that can't actuate
within its freshness window (currently an hour) lapses and the node
simply stays yours. The replacement depends on supply and your limit price.
Retrying the same node while it swaps is an idempotent `200`; swaps on
different nodes stack, spaced by the shared write cadence. To shrink
instead of swap, lower `max_gpus` on the capacity PUT (with `release`
to pick the victim).

## History and activity

`GET /api/vm/activity` is the Instances page's History: applied
declarations and allocation events, newest first, 30 per page. Each
source is read 400 rows deep before the merge, so very old history ages
out. Pages carry 30 entries
(`?page=`, from 0; `more` says whether a next page exists). An entry of
`kind: request` is a declaration the platform applied, with its
`actor`, `max_gpus`, and `max_price_per_gpu_hour`; `who` is `user` for
your organization's own declares and `system` for a platform override.
An entry of `kind: action` is an allocation event: a node granted or
released, named by its public IP in `note`, or a grant attempt unwound.
`?who=user` or `?who=system` filters. Capacity reads never appear.

`GET /api/vm/history?hours=24` is the Instances page's charts as
bucketed series over a window of 0.5 to 8760 hours: `t0` (epoch seconds
of the first bucket), `step` (bucket width in seconds, `max(60, span /
400)`), `max_nodes` (the declared node count in force: `max_gpus`
divided by `gpus_per_node`; `null` before the first declaration),
`granted` (nodes held; `0` when none, never `null`), `max_price` (your
declared ceiling) and `price` (the market's clearing rate), both in USD
per node-hour. Divide by `gpus_per_node` to read them per GPU-hour. A
`null` price is a bucket where nothing was declared or nothing cleared.

## Errors

| Status | When |
|---|---|
| `401` / `403` / `404` | see [authentication](https://docs.nationalcompute.com/authentication.md) |
| `409` | `version-mismatch`, `unknown-node` (a `release`/swap entry that matches no held node), `ambiguous-cluster`, `static-posture`; swaps can also refuse `swap-disabled`, `cluster-gated`, or `node-pinned` |
| `422` | validation — `price-precision`, `bid-too-low` (any `max_price_per_gpu_hour` set too low to ever clear the market, raising or lowering `max_gpus`; a standing price left unchanged can still lower `max_gpus` or release, so a market that moved past your ceiling never traps held capacity), `price-above-maximum` (a `max_price_per_gpu_hour` over the cap the GET echoes as `max_price_cap_per_gpu_hour`; the response carries `maximum`), `no-ssh-keys`, a non-multiple `max_gpus`; the `detail` says exactly what |
| `429` | faster than the per cluster write cadence, or past the token's request budget for the minute; wait `retry_after_s` (the `Retry-After` header carries the same figure) |
| `503` | temporary platform condition — reads keep serving, retry writes later |

The [error reference](https://docs.nationalcompute.com/api/errors.md) covers every code with handling
advice.

The too-low check runs at write time only. When market conditions
move past a standing declaration, the ask keeps losing and its
unfilled capacity never arrives. The GET reports this as
`declared.bid_too_low: true`; check the
[market feed](https://docs.nationalcompute.com/api/market-feed.md) and raise `max_price_per_gpu_hour` to
fill again.

`declared.capacity_unavailable` is the other reason a standing ask can
sit unfilled, and it is not about price. `true` means the site has no
capacity of the requested class right now, so raising the price will
not fill pending slots sooner. The
system retries automatically and slots fill when capacity returns.
`false` means capacity is flowing normally. `null` means no verdict
(the detector is not armed at this site). The same verdict reaches
Kubernetes clusters on the console's Workloads page — see
[when the site has no capacity](https://docs.nationalcompute.com/kubernetes.md#when-the-site-has-no-capacity).
