MCP server¶
Point your own agent at National Compute. The platform hosts an MCP (Model Context Protocol) server that bundles the whole tenant surface — market data, capacity and limit prices, billing, hardware metrics, and server-side Kubernetes access — behind one OAuth sign-in. No API token to mint, no kubeconfig to install, no kubectl on the agent's machine.
https://nationalcompute.com/mcp
Any MCP capable agent connects: marshall, Claude Code, Codex, Cursor,
claude.ai, or your own. The orientation the server hands every client on
connect states the framing in one line. You are talking to a market
rather than a reservation system, so read market_read or
price_estimate before pricing anything.
Connecting¶
The endpoint is Streamable HTTP with standard MCP OAuth: add it to your client and complete the browser sign-in when prompted. You sign in as yourself — a member of your organization — and every action the agent takes is attributed to you.
Claude Code
claude mcp add --transport http national-compute https://nationalcompute.com/mcp
claude.ai — Settings → Connectors → Add custom connector with the URL above.
Cursor — add to mcp.json:
{"mcpServers": {"national-compute": {"url": "https://nationalcompute.com/mcp"}}}
If you belong to several organizations, configure the endpoint as
https://nationalcompute.com/mcp/o/<org> (your org's name or id); the
server tells you so on the first tool call.
Headless machines (a VM with no browser): the OAuth callback lands
on localhost, so forward it over your SSH session while you sign in
from your laptop's browser:
ssh -L <port>:localhost:<port> you@your-vm
where <port> is the callback port your MCP client prints. Sign-ins are
long-lived — an idle agent re-authenticates after ~90 days, not daily.
The wire¶
An unauthenticated POST answers 401 with the RFC 9728 challenge that
points a client at sign-in:
WWW-Authenticate: Bearer resource_metadata="https://nationalcompute.com/.well-known/oauth-protected-resource/mcp"
That metadata document names the authorization server,
https://auth.nationalcompute.com/application/o/nc-mcp/. It accepts the
authorization code flow with PKCE, which is what MCP clients use. It
also accepts the device code flow, for a terminal client that cannot
open a browser.
Protocol version 2025-06-18.
| Fact | Value |
|---|---|
| Method | POST; GET and DELETE answer 405 |
| Response | one JSON-RPC response per POST, application/json |
| Sessions | none: no session id, no server-sent events, no server push |
| Batches | refused |
| Methods | initialize, ping, tools/list, tools/call; notifications/* answer 202 |
| Result size | each tool result is clipped at 40,000 characters, with a suffix naming the limit |
$NC_MCP_TOKEN below is the OAuth access token your client obtained,
never an org API token. Start with initialize:
curl -sS https://nationalcompute.com/mcp \
-H "Authorization: Bearer $NC_MCP_TOKEN" \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{
"protocolVersion":"2025-06-18","capabilities":{},
"clientInfo":{"name":"my-agent","version":"1.0"}}}'
{"jsonrpc": "2.0", "id": 1, "result": {
"protocolVersion": "2025-06-18",
"capabilities": {"tools": {}},
"serverInfo": {"name": "national-compute", "version": "1.0.0"},
"instructions": "National Compute MCP server. Call grid_whoami first: …"
}}
tools/list answers the whole catalog in one response:
{"jsonrpc": "2.0", "id": 2, "method": "tools/list"}
A tools/call carries the tool name and its arguments. The result is
one text block holding the tool's JSON:
{"jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": {
"name": "metrics_read",
"arguments": {"action": "flags", "cluster": "aurora-prod",
"kind": "Job", "name": "train-llm", "range": "6h"}}}
{"jsonrpc": "2.0", "id": 3, "result": {
"content": [{"type": "text", "text": "{\n \"cluster\": \"aurora-prod\", …"}],
"isError": false
}}
A refused tool answers isError: true with the refusal JSON in that
same text block. The JSON-RPC error envelope is reserved for protocol
faults.
The tools¶
Call grid_whoami first: it names your organization and every cluster
you can address. Each cluster entry carries the facts the console gates
its pages on: market (true = the cluster trades on the market and a
limit price or capacity declaration applies; false = a fixed cluster
the platform sizes; null = not known yet), gpus_per_node,
shared_volume (null when the storage record could not be read) and
cpu_worker. orgs lists the organizations you belong to as
{id, display}, and org_bound with orgs_note say whether you can
switch: an OAuth session with several organizations pins one with the
/mcp/o/<org> endpoint; a member or workspace token is bound to its
organization and cannot switch. Kubernetes tools take an explicit
cluster argument and run server-side under an agent service
account minted for you on demand — the credential never reaches the
client.
Read tools:
| Tool | What it does |
|---|---|
grid_whoami |
your org (org_id, org_display, org_kind), your identity, your own orgs, every addressable cluster with its kind, GPU class, market posture, gpus_per_node, shared_volume and cpu_worker; a Slurm entry adds its login endpoint |
market_read |
clearing-price ticks (20 rows by default, limit up to 500; tail with after_id, page back with before_id, as on the price feed), bucketed history, or the price-to-win ladder |
capacity_read |
the standing limit price (k8s) or capacity declaration (VM) |
roadmap_read |
expected landing quarter for a GPU model, from planning; an expectation and never a commitment; no quantities, dates or sources |
price_estimate |
a cost band for a job shape from recent clearings — an index, not a quote |
billing_read |
transactions, daily rollups, summary (the balance card, the hold state, the wire transfer rail), balance + burn rate, per-workload usage, rates |
cluster_read |
live Kubernetes state: pods and events as one summary row per object (raw: true for the full objects), logs, any resource, raw GET paths |
job_watch |
a bounded poll over a job's pods: phases, restarts, the market's own events with their verdicts, preemption deadlines, Kueue Workload requeue state, the standing limit price |
service_call |
one HTTP request to a cluster-internal Service (the port-forward stand-in) |
metrics_read |
what the platform's monitoring saw on a node, or on every node a workload touched, plus flags with the evidence behind them; a node may sit in a Kubernetes or a Slurm cluster |
slurm_read |
a Slurm cluster as the console shows it: node states and GPU allocation, the live queue, GPU hours, shared /home, one job's record and queue position, its log tail or a grep over it |
metrics_query |
one PromQL expression of your own over your organization's monitoring, bounded and read only |
dashboard_read |
your own Grafana dashboard documents and the platform's tenant dashboards, with panels and variables |
docs_read |
this documentation: table of contents, search, pages |
recipe_read |
the starter recipes catalog and files |
skill_read |
the skill collection: instruction sets Marshall follows for one task, the global collection plus your organization's own skills; index lists them with label, source, install count and proposal status, body returns one skill's full instructions by id |
escalation_read |
your organization's own escalation records with their status (requested is the only status; there is no acknowledged state) |
context_read |
the guidance the console agent loads at session start: the escalation rules and the notes the platform team wrote for your organization |
feedback_read |
your organization's own feedback submissions and whether your next substantive submission earns the feedback reward |
checkpoint_read |
your organization's published checkpoints in every state; list, search (q over title, description and file name) and get by id. The public feed itself is the anonymous route, never a member tool's list |
grid_api |
the live OpenAPI contracts, plus a fixed set of GET paths from the REST API: the console's Workloads, Run History, job, reservation, kubeconfig, storage volume, token and activity reads beside the market, billing and organization feeds |
grid_api action=get serves the Storage page (/api/k8s/storage,
with its fullness bands) and the
Cluster Overview charts (/api/k8s/capacity/history, with the
seven health states) among its
paths; an unlisted path refuses and names the set.
Single-phase writes, each on records of your own:
| Tool | What it does |
|---|---|
dashboard_write |
create, update, keep or delete a Grafana dashboard document of your own from a small panel spec; scratch documents expire after 24 hours unless kept |
feedback_write |
record product feedback for the National Compute team in one call; nobody replies, and substantive feedback earns credits under the feedback rewards terms |
skill_write |
create, update and delete one of your organization's own skills (update and delete by the author only; a delete takes the skill's installs with it); propose is two-phase and asks the National Compute team to add the skill to the global collection, after a check against the public page policy |
Two-phase tools, the seven that spend money or mutate state:
| Tool | What it does |
|---|---|
bid_write |
k8s: set the limit price (release is a VM argument: an empty or null release on a k8s call is ignored, a non-empty one is refused). VM: declare capacity plus an optional release list naming the nodes to give back first (victim nomination) |
vm_swap |
destroy-and-replace one VM node |
ssh_keys |
list org SSH keys (free); add and revoke are two-phase |
cluster_apply |
server-side apply of a YAML manifest, with server dry-run previews |
cluster_delete |
delete one named object |
billing_write |
the billing actions: checkout, setup and portal answer a hosted page URL in one call for a person to complete; limit (tighten only) and reload (auto reload, both directions here because a person approves the preview) run in two phases under the human confirmation class |
escalate |
send a request the platform cannot self serve to the National Compute team (a Base Load reservation ask, a capacity watch, an unsupported ask, feedback); the preview is the message as the team receives it |
checkpoint_write |
delete one of your own published checkpoints (action: delete); the preview names the item, the confirm removes it for the organization and from the public feed. Publishing carries bytes the wire cannot: that is the REST family or Marshall's checkpoint_publish inside a workspace |
Node lifecycle is never available: no tool can cordon, drain, taint or
delete a node, on any path, with any confirmation. cluster_apply and
cluster_delete refuse Node kinds before they reach the cluster.
metrics_read has its own page. Telemetry and flags
covers its three actions, the answer shapes, the flag rules and the
refusals, and metrics_query, the free-form read next to it.
Dashboards covers dashboard_write and
dashboard_read: the panel spec, ownership, scratch expiry and the
caps. roadmap_read has its own page as well.
Checkpoints covers checkpoint_read and checkpoint_write, the public feed anyone can read and the REST family that publishes.
Hardware roadmap states what the roadmap publishes and
what it never publishes. grid_api get /api/k8s/nodes serves the
Cluster Overview board, every
node with the platform's verdict word; grid_api get
/api/k8s/nodes/<node> answers one machine.
grid_api paths¶
grid_api {action: get, path, params} serves a fixed set of the
REST API reads in process, org keyed like every tool. Any
other path reads 404 not-found with the whole set in detail.
Kubernetes paths take params.cluster the way the REST routes take
?cluster=, and refuse the way those routes refuse: 409
cluster-not-ready, 404 not-found for a cluster outside your
organization or a VM cluster, 422 bad-request for a malformed
parameter, 503 station-unavailable when the island cannot answer.
| Path | Page |
|---|---|
/api/whoami |
Index |
/api/market/ticks, /api/market/history |
Market feed |
/api/k8s/bid, /api/k8s/storage/volumes |
Limit price |
/api/k8s/cluster, /api/k8s/cluster/kubeconfig, /api/k8s/reservations |
Kubernetes cluster |
/api/k8s/market/ladder, /api/k8s/market/demand, /api/k8s/market/spend, /api/k8s/market/paid, /api/k8s/market/bidhistory |
Kubernetes market |
/api/k8s/workloads, /api/k8s/workloads/runs, /api/k8s/workloads/history, /api/k8s/workloads/cost, /api/k8s/workloads/metrics |
Cluster pages |
/api/k8s/jobs/{jobid}, /api/k8s/jobs/{jobid}/cost, /api/k8s/jobs/{jobid}/network |
Cluster pages |
/api/vm/capacity, /api/vm/activity, /api/vm/history |
VM capacity |
/api/billing/summary, /api/billing/balance, /api/billing/balance/history, /api/billing/transactions, /api/billing/daily, /api/billing/rates |
Billing |
/api/org/members, /api/org/tokens, /api/org/tokens/{token_id}/activity, /api/org/activity |
Organization |
A path with {jobid} or {token_id} takes the REST spelling
(/api/k8s/jobs/<namespace>~<pod>,
/api/org/tokens/<token_id>/activity) or the templated key itself with
the value in params. Both spellings are URL decoded the same way, so
ml%7Etrain-0 and ml~train-0 name one pod. A value given both ways
must agree; a disagreement reads 422 bad-request before anything is
read. A null in params is an omitted value. A pasted REST URL keeps
its query string: ?cluster=<name>&r=1h fills params, and a key set
in params wins over the same key in the query string.
Two paths answer differently from their REST routes on this leg:
/api/k8s/cluster/kubeconfiganswers JSON,{cluster, filename, media_type, kubeconfig}:kubeconfigholds the same YAML the REST route streams,filenamethe name the route's download carries. When the platform's own record served the file the route's headers ride as fields:freshness,observed_at,ingested_at, andcache_control(no-store)./api/org/tokens/self/activityreads422 bad-request: the MCP sign in carries no org API token, soselfnames nothing. Pass atoken_idfrom/api/org/tokens.
The nodes, machine, storage and capacity history reads
(/api/k8s/nodes, /api/k8s/nodes/{node}, /api/k8s/storage,
/api/k8s/capacity/history) are served on this leg as well.
Slurm reads covers slurm_read:
its four actions, the answer shapes, the node state vocabulary and the
connection facts grid_whoami carries for a Slurm cluster.
Two-phase confirmation¶
Every tool that spends money or mutates state runs in two phases. The
first call returns a preview — the current state, the proposed
change, its projected cost — and a single-use confirm_token valid for
300 seconds. The same call repeated with identical arguments plus
the token executes. Changed arguments, a reused token, or an expired one
refuse with a typed error; nothing executes without a fresh preview.
Your agent should show you the preview before confirming — but the pause is enforced server-side either way.
Human confirmation¶
A preview that spends money or destroys data carries one more key
beside confirm_required:
"human_confirmation": {"required": true,
"reason": "destroys the node and every file on its disks"}
reason is one sentence stating what happens to your organization. A
preview without the key is the ordinary gate. The server decides per
call, so a mixed tool marks only the calls that qualify:
| Call | Marked |
|---|---|
vm_swap |
yes |
bid_write with kind: vm and a max_gpus below the GPUs your granted nodes hold (0 included), or with a non empty release list |
yes; the reason names how many nodes the declaration releases; release: [] (clear the nominations) alone is not marked |
bid_write with kind: vm that grows or keeps the declaration, or with kind: k8s |
no |
cluster_delete of a PersistentVolumeClaim, Namespace, VolumeSnapshot or VolumeSnapshotContent (any case, singular or plural, pvc, ns, group qualified spellings) |
yes |
cluster_delete of a StatefulSet whose spec.persistentVolumeClaimRetentionPolicy.whenDeleted is Delete |
yes |
cluster_delete of a PersistentVolume (pv) |
refused today as an unknown kind (422 bad-request); marked if the kind is ever added |
cluster_delete of any other kind (Job, Pod, Deployment, Service, a StatefulSet that keeps its claims, ...) |
no |
cluster_apply of a StatefulSet document that sets persistentVolumeClaimRetentionPolicy.whenDeleted or whenScaled to Delete, or that lowers replicas on a live set whose whenScaled is Delete |
yes |
billing_write with action: limit |
yes; the reason states that the call lowers what the organization may spend this month |
billing_write with action: reload |
yes; with enabled: true the reason states that the call arms automatic card charges that spend the organization's money without further approval, with enabled: false that it changes the automatic top up setting |
billing_write with action: checkout, setup or portal |
no; one call answering a hosted page URL, which a person completes in a browser |
every other cluster_apply, ssh_keys |
no |
The marker, the instruction and the confirm_token precede preview in
the envelope, so the 40,000 character clip on a large preview never
removes them. The preview of a marked call also carries a warning
built from the same reason; a cluster_delete preview names the
canonical resource plural beside the kind you spelled. A cluster
scoped object (a Namespace, a VolumeSnapshotContent) has no namespace
in its preview. A decision read from the live cluster (a StatefulSet's
retention policy, the number of granted nodes) is read again at
execute; when the answer changed inside the window the call refuses
409 object-changed and needs a fresh preview.
The server marks the class and verifies nothing more: on the wire a
human's answer and an agent's are the same confirm_token. The
platform's own clients (marshall and the workspace agent) ask you in
the terminal for every marked preview. Full Access does not answer it.
A remembered answer does not answer it. A third party client is told
the same in the server's instructions and should hold the call for
your explicit answer. The execute leg is unchanged: the same token, the
same 300 second expiry, the same typed refusals.
The token a workspace agent holds cannot run a marked call at all. The
server refuses 403 workspace-human-class-refused before a preview
exists. Nothing in the workspace can prove a human answered. Use the
console or an MCP client signed in as yourself for those calls; every
other two phase call stays open to the workspace agent.
Tool annotations¶
Every tools/list entry carries the standard MCP annotations object,
so a client can gate by hint before it reads a description:
| Hint | Meaning here |
|---|---|
readOnlyHint |
true on every read tool; the tool changes nothing |
destructiveHint |
true on every write tool, since none is purely additive: vm_swap, cluster_delete, ssh_keys (revoke), dashboard_write (delete), service_call (a POST reaches your own Service), bid_write (a VM withdraw releases nodes and their disks), cluster_apply (an apply can replace a template or shrink replicas), billing_write (limit and reload change what the organization spends; the hosted pages mint a fresh Stripe session) |
idempotentHint |
true when repeating the same arguments changes nothing more: every read, bid_write, cluster_apply |
openWorldHint |
false on every tool; they talk to the platform only |
The hint and the class are different facts: a k8s bid_write, a
growing VM declaration and an ordinary cluster_apply are destructive
by hint and stay the ordinary gate.
Never available through any tool, with any confirmation: minting an org
API token or a member token (your human does it on the console's API
Keys page); inviting or removing members, join rules and domain rules;
buying Base Load; destroying the shared NFS volume (the Storage page's
volume destroy). A Kubernetes PersistentVolumeClaim delete is a
different object: it goes through cluster_delete under the human
confirmation class. The hosted billing pages (checkout, card setup,
payment portal) are URLs you complete in a browser.
Feedback rewards¶
feedback_write records how the platform worked for you — the cluster,
scheduling, pricing, the docs, the agent tools, billing — for the
National Compute team to read. Nobody replies to feedback; a question
that needs a human answer is escalate. It takes body (1 to 8000
characters) and optionally category (cluster, scheduling,
pricing, docs, agent_tools, billing, other), rating (1 to
5), cluster and contact_ok.
Public-research organizations cannot earn feedback rewards, even after
buying credits. Their feedback is still recorded and read; the reward
is skipped with reason public_research.
For eligible standard organizations, substantive feedback (500 or more characters after whitespace is collapsed) deposits 25 credits to the organization's balance when submitted. Two limits apply: one rewarded submission per member per rolling seven days, and at most one reward per $1,000 of credits your organization has loaded or received, over its lifetime, with the rewards themselves not counted. Shorter feedback is recorded and read, and earns nothing.
The response says whether the reward was granted and, if not, why:
short, member_weekly, org_cap, ineligible, disabled or
public_research, with next_eligible_at and org_rewards_remaining.
feedback_read shows
the same eligibility before you write, beside your organization's
earlier submissions. A rewarded submission appears on the
billing ledger as a platform grant with the note
"Feedback reward".
Rewards exist to hear real experience. Padded, duplicated or generated submissions make an organization ineligible for rewards, and that determination is entirely at National Compute's discretion. An ineligible organization can still submit feedback; it earns nothing.
Limits and semantics¶
- Rate limit: 300 requests per minute per member. Every
429carries aRetry-Afterheader andretry_after_sin the body — whether from your budget (rate-limited) or from a busy replica (server-busy, and per-toolwatch-busy/service-call-busyon the long-running tools). Back off and retry. - Attribution: every action is yours — the console's Activity page shows tool calls "via MCP", and Kubernetes audit logs name your agent service account.
- Access follows membership: removing a member from the org cuts their MCP access within seconds, sign-in or not.
- Errors use the API error catalog; optimistic concurrency
rides the previews. Capacity writes capture a version at preview time
and check it at execute — concurrent changes surface as
409 version-mismatch.cluster_deletecaptures the object's uid the same way: if a same-name object replaced the previewed one, execute refuses with409 object-changed— re-preview and look again. cluster_applymanifests: at most 1.5 MB of YAML and 20 objects per call.- The protocol is stateless: no sessions to manage, and previews / confirmations work across reconnects.
Prefer raw REST? Everything here wraps the same documented APIs — org API tokens keep working unchanged.