Skip to content

Public Research

Public Research is one whole GPU node at a time for an individual researcher, requested from Marshall and used over ssh: no cluster to set up, no limit price to pick, no auction to wait on. A member of a public research organization asks Marshall for a node, waits in a first-come, first-served queue, holds the node for a lease, and releases it; the meter runs from the moment the node is theirs to the moment it is not. The organization an eligible .edu, .mil or .gov address receives at sign-in is one; National Compute sets the kind for any other organization by arrangement.

The node

A node is yours whole: every GPU in it, its CPUs, memory and local NVMe, with nobody else on it. Nodes come in two shapes, 8 GPUs each: 8× AMD MI355X 288 GB or 8× NVIDIA B300 288 GB. Marshall's quote states the GPU model, the GPU count and the memory per GPU of each node on offer, and the quote is the authority whenever this page and the quote differ.

The local NVMe is scratch. Releasing a node, or losing it to a preempt, rebuilds it in place and destroys every byte on it; there is no undo and no recovery. Copy results to your workspace disk, or to storage of your own, before you release; a public research organization has no shared folder. Marshall says so before the first request and again as the lease nears its end.

What a node does not have: a second node to pair with (no multi-node jobs, no RDMA or scale-out fabric between pool nodes) and no notebook server (ssh is the only way in).

Asking Marshall

Marshall is the one door: there is no console form, no API route and no MCP tool for a request, and only a member of a public research organization is offered one. Marshall quotes before it asks: the hourly figure for the node, the worst case for one full lease (lease length × GPUs × rate), the balance a request needs, and whether your job fits one node. A job that needs more than one node belongs on the market, not here. Every request ends in a confirmation card, Full Access or not, because a request starts a meter.

With two node shapes open, Marshall lists both offers, each with its rate per GPU-hour, the worst case for one full lease and the current queue, and marks one Recommended from your job's software stack: a CUDA stack points to the NVIDIA B300, a ROCm stack to the AMD MI355X; when both fit, the cheaper offer, then the shorter queue. The mark is a mark, not a choice: you pick either row, and the pick is the confirmation. A request must name one offer; one that names none is refused ("site required") with the offers listed, and one naming an offer that is not open is refused with what is open. While only one shape is open, Marshall quotes it alone and asks as before.

A request needs a balance that covers the first hour of a lease at the quoted rate (GPUs × rate), not the whole lease; past that hour the meter runs against your balance like any other charge. A public research organization starts with $100 of credits (the airdrop deposit on its ledger); once those are spent a request is refused until credits are bought on the console's Billing page (card or wire); Marshall names the amount it needs. A request is also refused while a billing hold is active.

The queue

Requests are served first come, first served, one node per request. A request is queued from the moment it is accepted; Marshall reports your place in line first and then an estimated wait. The estimate takes the nodes in the pool, the requests ahead of you and the time left on the leases running now, and assumes every lease runs its full length; a release before that moves every estimate behind it earlier, so it is an estimate and never a promise. The line itself moves only when a lease ends or an idle node is added. When a node is free for you, handover takes seconds: the platform deposits Marshall's public key on the node and hands back its address and host key. Nothing is installed and nothing is rebuilt on the way in.

State Meaning
held accepted, waiting behind another request of yours; enters the queue at the back when that one ends
queued in line; position is your place
provisioning a node is being handed over; seconds
active the node is yours and the meter is running
ended the lease is over; end_reason is released, preempted, lost, balance (the organization's balance reached zero) or admin

You hold one node at a time (currently one queued or active request per user). A request from a second Marshall session while a first is live is accepted as held, outside the line, and joins the back of the queue when the live request ends; a second request from a session that is already waiting is refused, and held requests are capped (currently three per user). The console's sessions rail shows which session is waiting and which one has a node, and the console's Public Research page shows the same per GPU type: a slot for each type that holds the session using a node of that type, with the sessions waiting for one listed beneath it, each with its place in line and estimated wait; a click on a session opens it. Each type's heading shows how many GPUs the pool currently holds for it, and the console's Grid entry shows the total across types. An empty slot offers a button naming the GPU type, "Get MI355X with Marshall" or "Get B300 with Marshall": one click opens a new Marshall session that requests a node of that type, walks you through how the pool works and, once the node is yours, opens a shell on it and shows its GPUs, with the usual confirmation before anything is requested. Under the button the slot shows the estimated wait a request made now would get, before you have asked for anything.

The lease

A lease is 4 hours (currently) and is preempted only when someone is waiting. While nobody is queued the lease runs past 4 hours and keeps billing until you release. Once another request is queued and your lease has passed its length, a preempt notice is set: 10 minutes (currently) to checkpoint, then the node is taken and rebuilt. Marshall wakes you 30 and 10 minutes before the lease length is reached and the moment a notice is set, with the time left and a reminder to copy results off the node.

Release when you are done: the end is stamped before the rebuild starts, so the recycle costs you nothing. Marshall never releases a node on its own: when a job finishes it says the meter is still running and asks you; releasing, like requesting, goes through a confirmation card. A node the platform loses (a failed node, a failed handover or rebuild) ends the lease at the earliest evidence of the failure, and billing with it. A lease also ends the moment the organization's balance reaches zero: the node is taken back, the meter stops and end_reason reads balance; add credits before requesting again. National Compute can end a lease; end_reason then reads admin.

To keep working past a lease, ask Marshall for another node from the session that holds it: the new request is held, joins the back of the queue when the lease ends, and lands on a freshly rebuilt node. Nothing carries over from the old node.

Working on the node

Marshall reaches the node with plain ssh on its public address, port 22, as the login user the request names. The one key is minted by Marshall on your workspace disk and reused for every request; the node's host key arrives with the node and is pinned before Marshall announces it, so the first connection is verified rather than trusted on first use. A shell pane in your workspace carries the same key.

Long jobs detach on the node (nohup, tmux): Marshall's own command runs are time-capped and end with the session. Copy results back with scp or rsync before release.

The first job on a node is the platform's nanoGPT starter, started on the node with one line and detached from the ssh session:

ssh <ssh_user>@<public_ip> 'curl -fsSL https://access.nationalcompute.com/first-job/nanogpt-node.sh | sh'

It trains in the stock PyTorch container for the node's GPUs (the node runs it with docker), writes its log to ~/first-job/nanogpt.out there, and ends on its own after about an hour; the node keeps billing until you release it. The starter recipe catalog lists it with cluster_kind public_research, and FIRST-JOB.md walks an agent through request, connect, train, copy back and release. Your own container runs the same way with docker run on the node. The Serve and Post-train recipes are Kubernetes manifests and do not apply here.

Wiping your workspace disk mid-lease takes the key with it; the lease keeps running, and the way back in is release and a new request.

Billing

A lease bills per GPU-hour, for every GPU in the node, at the rate shown in your Marshall quote and on the console's Billing page under Rates (label "Public Research · GPU", unit GPU; one line per node shape), from the moment the node is yours to the moment the lease ends: release, preempt, loss or an end by National Compute alike. The rate is fixed at grant for the whole lease; a change applies to leases granted afterwards and never reprices a running one. Waiting costs nothing, and the rebuild after a lease costs nothing.

Charges land on your organization's credit ledger about every five minutes, in the same credit unit as every other charge, as rows of source public_research whose entries name the node, the GPUs, the rate and the window. The Billing page shows them under the Public Research chip; the billing summary carries the month's GPU-hours and charged amount as public_research. A free rate still counts GPU-hours and writes no charge.

A request is admitted only with a balance covering the first hour of a lease, and a billing hold refuses new requests. A balance that reaches zero ends a running lease at once: the node is taken back and the meter stops (end_reason balance), so a lease never runs far past what the balance covers. When nobody is waiting, a lease runs and bills past its length with no other cap. Release when the job is done.