Skip to content

MCP server

Point your own agent at National Compute. The platform hosts an MCP (Model Context Protocol) server that bundles the whole tenant surface — market data, capacity and limit prices, billing, hardware metrics, and server-side Kubernetes access — behind one OAuth sign-in. No API token to mint, no kubeconfig to install, no kubectl on the agent's machine.

https://nationalcompute.com/mcp

Any MCP capable agent connects: marshall, Claude Code, Codex, Cursor, claude.ai, or your own. The orientation the server hands every client on connect states the framing in one line. You are talking to a market rather than a reservation system, so read market_read or price_estimate before pricing anything.

Connecting

The endpoint is Streamable HTTP with standard MCP OAuth: add it to your client and complete the browser sign-in when prompted. You sign in as yourself — a member of your organization — and every action the agent takes is attributed to you.

Claude Code

claude mcp add --transport http national-compute https://nationalcompute.com/mcp

claude.ai — Settings → Connectors → Add custom connector with the URL above.

Cursor — add to mcp.json:

{"mcpServers": {"national-compute": {"url": "https://nationalcompute.com/mcp"}}}

If you belong to several organizations, configure the endpoint as https://nationalcompute.com/mcp/o/<org> (your org's name or id); the server tells you so on the first tool call.

Headless machines (a VM with no browser): the OAuth callback lands on localhost, so forward it over your SSH session while you sign in from your laptop's browser:

ssh -L <port>:localhost:<port> you@your-vm

where <port> is the callback port your MCP client prints. Sign-ins are long-lived — an idle agent re-authenticates after ~90 days, not daily.

The wire

An unauthenticated POST answers 401 with the RFC 9728 challenge that points a client at sign-in:

WWW-Authenticate: Bearer resource_metadata="https://nationalcompute.com/.well-known/oauth-protected-resource/mcp"

That metadata document names the authorization server, https://auth.nationalcompute.com/application/o/nc-mcp/. It accepts the authorization code flow with PKCE, which is what MCP clients use. It also accepts the device code flow, for a terminal client that cannot open a browser.

Protocol version 2025-06-18.

Fact Value
Method POST; GET and DELETE answer 405
Response one JSON-RPC response per POST, application/json
Sessions none: no session id, no server-sent events, no server push
Batches refused
Methods initialize, ping, tools/list, tools/call; notifications/* answer 202
Result size each tool result is clipped at 40,000 characters, with a suffix naming the limit

$NC_MCP_TOKEN below is the OAuth access token your client obtained, never an org API token. Start with initialize:

curl -sS https://nationalcompute.com/mcp \
  -H "Authorization: Bearer $NC_MCP_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{
       "protocolVersion":"2025-06-18","capabilities":{},
       "clientInfo":{"name":"my-agent","version":"1.0"}}}'
{"jsonrpc": "2.0", "id": 1, "result": {
  "protocolVersion": "2025-06-18",
  "capabilities": {"tools": {}},
  "serverInfo": {"name": "national-compute", "version": "1.0.0"},
  "instructions": "National Compute MCP server. Call grid_whoami first: …"
}}

tools/list answers the whole catalog in one response:

{"jsonrpc": "2.0", "id": 2, "method": "tools/list"}

A tools/call carries the tool name and its arguments. The result is one text block holding the tool's JSON:

{"jsonrpc": "2.0", "id": 3, "method": "tools/call", "params": {
  "name": "metrics_read",
  "arguments": {"action": "flags", "cluster": "aurora-prod",
                "kind": "Job", "name": "train-llm", "range": "6h"}}}
{"jsonrpc": "2.0", "id": 3, "result": {
  "content": [{"type": "text", "text": "{\n  \"cluster\": \"aurora-prod\", …"}],
  "isError": false
}}

A refused tool answers isError: true with the refusal JSON in that same text block. The JSON-RPC error envelope is reserved for protocol faults.

The tools

Call grid_whoami first: it names your organization and every cluster you can address. Each cluster entry carries the facts the console gates its pages on: market (true = the cluster trades on the market and a limit price or capacity declaration applies; false = a fixed cluster the platform sizes; null = not known yet), gpus_per_node, shared_volume (null when the storage record could not be read) and cpu_worker. orgs lists the organizations you belong to as {id, display}, and org_bound with orgs_note say whether you can switch: an OAuth session with several organizations pins one with the /mcp/o/<org> endpoint; a member or workspace token is bound to its organization and cannot switch. Kubernetes tools take an explicit cluster argument and run server-side under an agent service account minted for you on demand — the credential never reaches the client.

Read tools:

Tool What it does
grid_whoami your org (org_id, org_display, org_kind), your identity, your own orgs, every addressable cluster with its kind, GPU class, market posture, gpus_per_node, shared_volume and cpu_worker; a Slurm entry adds its login endpoint
market_read clearing-price ticks (20 rows by default, limit up to 500; tail with after_id, page back with before_id, as on the price feed), bucketed history, or the price-to-win ladder
capacity_read the standing limit price (k8s) or capacity declaration (VM)
roadmap_read expected landing quarter for a GPU model, from planning; an expectation and never a commitment; no quantities, dates or sources
price_estimate a cost band for a job shape from recent clearings — an index, not a quote
billing_read transactions, daily rollups, summary (the balance card, the hold state, the wire transfer rail), balance + burn rate, per-workload usage, rates
cluster_read live Kubernetes state: pods and events as one summary row per object (raw: true for the full objects), logs, any resource, raw GET paths
job_watch a bounded poll over a job's pods: phases, restarts, the market's own events with their verdicts, preemption deadlines, Kueue Workload requeue state, the standing limit price
service_call one HTTP request to a cluster-internal Service (the port-forward stand-in)
metrics_read what the platform's monitoring saw on a node, or on every node a workload touched, plus flags with the evidence behind them; a node may sit in a Kubernetes or a Slurm cluster
slurm_read a Slurm cluster as the console shows it: node states and GPU allocation, the live queue, GPU hours, shared /home, one job's record and queue position, its log tail or a grep over it
metrics_query one PromQL expression of your own over your organization's monitoring, bounded and read only
dashboard_read your own Grafana dashboard documents and the platform's tenant dashboards, with panels and variables
docs_read this documentation: table of contents, search, pages
recipe_read the starter recipes catalog and files
skill_read the skill collection: instruction sets Marshall follows for one task, the global collection plus your organization's own skills; index lists them with label, source, install count and proposal status, body returns one skill's full instructions by id
escalation_read your organization's own escalation records with their status (requested is the only status; there is no acknowledged state)
context_read the guidance the console agent loads at session start: the escalation rules and the notes the platform team wrote for your organization
feedback_read your organization's own feedback submissions and whether your next substantive submission earns the feedback reward
checkpoint_read your organization's published checkpoints in every state; list, search (q over title, description and file name) and get by id. The public feed itself is the anonymous route, never a member tool's list
grid_api the live OpenAPI contracts, plus a fixed set of GET paths from the REST API: the console's Workloads, Run History, job, reservation, kubeconfig, storage volume, token and activity reads beside the market, billing and organization feeds

grid_api action=get serves the Storage page (/api/k8s/storage, with its fullness bands) and the Cluster Overview charts (/api/k8s/capacity/history, with the seven health states) among its paths; an unlisted path refuses and names the set.

Single-phase writes, each on records of your own:

Tool What it does
dashboard_write create, update, keep or delete a Grafana dashboard document of your own from a small panel spec; scratch documents expire after 24 hours unless kept
feedback_write record product feedback for the National Compute team in one call; nobody replies, and substantive feedback earns credits under the feedback rewards terms
skill_write create, update and delete one of your organization's own skills (update and delete by the author only; a delete takes the skill's installs with it); propose is two-phase and asks the National Compute team to add the skill to the global collection, after a check against the public page policy

Two-phase tools, the seven that spend money or mutate state:

Tool What it does
bid_write k8s: set the limit price (release is a VM argument: an empty or null release on a k8s call is ignored, a non-empty one is refused). VM: declare capacity plus an optional release list naming the nodes to give back first (victim nomination)
vm_swap destroy-and-replace one VM node
ssh_keys list org SSH keys (free); add and revoke are two-phase
cluster_apply server-side apply of a YAML manifest, with server dry-run previews
cluster_delete delete one named object
billing_write the billing actions: checkout, setup and portal answer a hosted page URL in one call for a person to complete; limit (tighten only) and reload (auto reload, both directions here because a person approves the preview) run in two phases under the human confirmation class
escalate send a request the platform cannot self serve to the National Compute team (a Base Load reservation ask, a capacity watch, an unsupported ask, feedback); the preview is the message as the team receives it
checkpoint_write delete one of your own published checkpoints (action: delete); the preview names the item, the confirm removes it for the organization and from the public feed. Publishing carries bytes the wire cannot: that is the REST family or Marshall's checkpoint_publish inside a workspace

Node lifecycle is never available: no tool can cordon, drain, taint or delete a node, on any path, with any confirmation. cluster_apply and cluster_delete refuse Node kinds before they reach the cluster.

metrics_read has its own page. Telemetry and flags covers its three actions, the answer shapes, the flag rules and the refusals, and metrics_query, the free-form read next to it. Dashboards covers dashboard_write and dashboard_read: the panel spec, ownership, scratch expiry and the caps. roadmap_read has its own page as well. Checkpoints covers checkpoint_read and checkpoint_write, the public feed anyone can read and the REST family that publishes. Hardware roadmap states what the roadmap publishes and what it never publishes. grid_api get /api/k8s/nodes serves the Cluster Overview board, every node with the platform's verdict word; grid_api get /api/k8s/nodes/<node> answers one machine.

grid_api paths

grid_api {action: get, path, params} serves a fixed set of the REST API reads in process, org keyed like every tool. Any other path reads 404 not-found with the whole set in detail. Kubernetes paths take params.cluster the way the REST routes take ?cluster=, and refuse the way those routes refuse: 409 cluster-not-ready, 404 not-found for a cluster outside your organization or a VM cluster, 422 bad-request for a malformed parameter, 503 station-unavailable when the island cannot answer.

Path Page
/api/whoami Index
/api/market/ticks, /api/market/history Market feed
/api/k8s/bid, /api/k8s/storage/volumes Limit price
/api/k8s/cluster, /api/k8s/cluster/kubeconfig, /api/k8s/reservations Kubernetes cluster
/api/k8s/market/ladder, /api/k8s/market/demand, /api/k8s/market/spend, /api/k8s/market/paid, /api/k8s/market/bidhistory Kubernetes market
/api/k8s/workloads, /api/k8s/workloads/runs, /api/k8s/workloads/history, /api/k8s/workloads/cost, /api/k8s/workloads/metrics Cluster pages
/api/k8s/jobs/{jobid}, /api/k8s/jobs/{jobid}/cost, /api/k8s/jobs/{jobid}/network Cluster pages
/api/vm/capacity, /api/vm/activity, /api/vm/history VM capacity
/api/billing/summary, /api/billing/balance, /api/billing/balance/history, /api/billing/transactions, /api/billing/daily, /api/billing/rates Billing
/api/org/members, /api/org/tokens, /api/org/tokens/{token_id}/activity, /api/org/activity Organization

A path with {jobid} or {token_id} takes the REST spelling (/api/k8s/jobs/<namespace>~<pod>, /api/org/tokens/<token_id>/activity) or the templated key itself with the value in params. Both spellings are URL decoded the same way, so ml%7Etrain-0 and ml~train-0 name one pod. A value given both ways must agree; a disagreement reads 422 bad-request before anything is read. A null in params is an omitted value. A pasted REST URL keeps its query string: ?cluster=<name>&r=1h fills params, and a key set in params wins over the same key in the query string.

Two paths answer differently from their REST routes on this leg:

  • /api/k8s/cluster/kubeconfig answers JSON, {cluster, filename, media_type, kubeconfig}: kubeconfig holds the same YAML the REST route streams, filename the name the route's download carries. When the platform's own record served the file the route's headers ride as fields: freshness, observed_at, ingested_at, and cache_control (no-store).
  • /api/org/tokens/self/activity reads 422 bad-request: the MCP sign in carries no org API token, so self names nothing. Pass a token_id from /api/org/tokens.

The nodes, machine, storage and capacity history reads (/api/k8s/nodes, /api/k8s/nodes/{node}, /api/k8s/storage, /api/k8s/capacity/history) are served on this leg as well.

Slurm reads covers slurm_read: its four actions, the answer shapes, the node state vocabulary and the connection facts grid_whoami carries for a Slurm cluster.

Two-phase confirmation

Every tool that spends money or mutates state runs in two phases. The first call returns a preview — the current state, the proposed change, its projected cost — and a single-use confirm_token valid for 300 seconds. The same call repeated with identical arguments plus the token executes. Changed arguments, a reused token, or an expired one refuse with a typed error; nothing executes without a fresh preview.

Your agent should show you the preview before confirming — but the pause is enforced server-side either way.

Human confirmation

A preview that spends money or destroys data carries one more key beside confirm_required:

"human_confirmation": {"required": true,
                       "reason": "destroys the node and every file on its disks"}

reason is one sentence stating what happens to your organization. A preview without the key is the ordinary gate. The server decides per call, so a mixed tool marks only the calls that qualify:

Call Marked
vm_swap yes
bid_write with kind: vm and a max_gpus below the GPUs your granted nodes hold (0 included), or with a non empty release list yes; the reason names how many nodes the declaration releases; release: [] (clear the nominations) alone is not marked
bid_write with kind: vm that grows or keeps the declaration, or with kind: k8s no
cluster_delete of a PersistentVolumeClaim, Namespace, VolumeSnapshot or VolumeSnapshotContent (any case, singular or plural, pvc, ns, group qualified spellings) yes
cluster_delete of a StatefulSet whose spec.persistentVolumeClaimRetentionPolicy.whenDeleted is Delete yes
cluster_delete of a PersistentVolume (pv) refused today as an unknown kind (422 bad-request); marked if the kind is ever added
cluster_delete of any other kind (Job, Pod, Deployment, Service, a StatefulSet that keeps its claims, ...) no
cluster_apply of a StatefulSet document that sets persistentVolumeClaimRetentionPolicy.whenDeleted or whenScaled to Delete, or that lowers replicas on a live set whose whenScaled is Delete yes
billing_write with action: limit yes; the reason states that the call lowers what the organization may spend this month
billing_write with action: reload yes; with enabled: true the reason states that the call arms automatic card charges that spend the organization's money without further approval, with enabled: false that it changes the automatic top up setting
billing_write with action: checkout, setup or portal no; one call answering a hosted page URL, which a person completes in a browser
every other cluster_apply, ssh_keys no

The marker, the instruction and the confirm_token precede preview in the envelope, so the 40,000 character clip on a large preview never removes them. The preview of a marked call also carries a warning built from the same reason; a cluster_delete preview names the canonical resource plural beside the kind you spelled. A cluster scoped object (a Namespace, a VolumeSnapshotContent) has no namespace in its preview. A decision read from the live cluster (a StatefulSet's retention policy, the number of granted nodes) is read again at execute; when the answer changed inside the window the call refuses 409 object-changed and needs a fresh preview.

The server marks the class and verifies nothing more: on the wire a human's answer and an agent's are the same confirm_token. The platform's own clients (marshall and the workspace agent) ask you in the terminal for every marked preview. Full Access does not answer it. A remembered answer does not answer it. A third party client is told the same in the server's instructions and should hold the call for your explicit answer. The execute leg is unchanged: the same token, the same 300 second expiry, the same typed refusals.

The token a workspace agent holds cannot run a marked call at all. The server refuses 403 workspace-human-class-refused before a preview exists. Nothing in the workspace can prove a human answered. Use the console or an MCP client signed in as yourself for those calls; every other two phase call stays open to the workspace agent.

Tool annotations

Every tools/list entry carries the standard MCP annotations object, so a client can gate by hint before it reads a description:

Hint Meaning here
readOnlyHint true on every read tool; the tool changes nothing
destructiveHint true on every write tool, since none is purely additive: vm_swap, cluster_delete, ssh_keys (revoke), dashboard_write (delete), service_call (a POST reaches your own Service), bid_write (a VM withdraw releases nodes and their disks), cluster_apply (an apply can replace a template or shrink replicas), billing_write (limit and reload change what the organization spends; the hosted pages mint a fresh Stripe session)
idempotentHint true when repeating the same arguments changes nothing more: every read, bid_write, cluster_apply
openWorldHint false on every tool; they talk to the platform only

The hint and the class are different facts: a k8s bid_write, a growing VM declaration and an ordinary cluster_apply are destructive by hint and stay the ordinary gate.

Never available through any tool, with any confirmation: minting an org API token or a member token (your human does it on the console's API Keys page); inviting or removing members, join rules and domain rules; buying Base Load; destroying the shared NFS volume (the Storage page's volume destroy). A Kubernetes PersistentVolumeClaim delete is a different object: it goes through cluster_delete under the human confirmation class. The hosted billing pages (checkout, card setup, payment portal) are URLs you complete in a browser.

Feedback rewards

feedback_write records how the platform worked for you — the cluster, scheduling, pricing, the docs, the agent tools, billing — for the National Compute team to read. Nobody replies to feedback; a question that needs a human answer is escalate. It takes body (1 to 8000 characters) and optionally category (cluster, scheduling, pricing, docs, agent_tools, billing, other), rating (1 to 5), cluster and contact_ok.

Public-research organizations cannot earn feedback rewards, even after buying credits. Their feedback is still recorded and read; the reward is skipped with reason public_research.

For eligible standard organizations, substantive feedback (500 or more characters after whitespace is collapsed) deposits 25 credits to the organization's balance when submitted. Two limits apply: one rewarded submission per member per rolling seven days, and at most one reward per $1,000 of credits your organization has loaded or received, over its lifetime, with the rewards themselves not counted. Shorter feedback is recorded and read, and earns nothing.

The response says whether the reward was granted and, if not, why: short, member_weekly, org_cap, ineligible, disabled or public_research, with next_eligible_at and org_rewards_remaining. feedback_read shows the same eligibility before you write, beside your organization's earlier submissions. A rewarded submission appears on the billing ledger as a platform grant with the note "Feedback reward".

Rewards exist to hear real experience. Padded, duplicated or generated submissions make an organization ineligible for rewards, and that determination is entirely at National Compute's discretion. An ineligible organization can still submit feedback; it earns nothing.

Limits and semantics

  • Rate limit: 300 requests per minute per member. Every 429 carries a Retry-After header and retry_after_s in the body — whether from your budget (rate-limited) or from a busy replica (server-busy, and per-tool watch-busy / service-call-busy on the long-running tools). Back off and retry.
  • Attribution: every action is yours — the console's Activity page shows tool calls "via MCP", and Kubernetes audit logs name your agent service account.
  • Access follows membership: removing a member from the org cuts their MCP access within seconds, sign-in or not.
  • Errors use the API error catalog; optimistic concurrency rides the previews. Capacity writes capture a version at preview time and check it at execute — concurrent changes surface as 409 version-mismatch. cluster_delete captures the object's uid the same way: if a same-name object replaced the previewed one, execute refuses with 409 object-changed — re-preview and look again.
  • cluster_apply manifests: at most 1.5 MB of YAML and 20 objects per call.
  • The protocol is stateless: no sessions to manage, and previews / confirmations work across reconnects.

Prefer raw REST? Everything here wraps the same documented APIs — org API tokens keep working unchanged.