Public Research¶
Public Research is one whole GPU node at a time for an individual
researcher, requested from Marshall and used over ssh: no cluster to
set up, no limit price to pick, no auction to wait on. A member of a
public research organization
asks Marshall for a node, waits in a first-come, first-served queue,
holds the node for a lease, and releases it; the meter runs from the
moment the node is theirs to the moment it is not. The organization an
eligible .edu, .mil or .gov address receives at sign-in is one;
National Compute sets the kind for any other organization by arrangement.
The node¶
A node is yours whole: every GPU in it, its CPUs, memory and local NVMe, with nobody else on it. Nodes come in two shapes, 8 GPUs each: 8× AMD MI355X 288 GB or 8× NVIDIA B300 288 GB. Marshall's quote states the GPU model, the GPU count and the memory per GPU of each node on offer, and the quote is the authority whenever this page and the quote differ.
The local NVMe is scratch. Releasing a node, or losing it to a preempt, rebuilds it in place and destroys every byte on it; there is no undo and no recovery. Copy results to your workspace disk, or to storage of your own, before you release; a public research organization has no shared folder. Marshall says so before the first request and again as the lease nears its end.
What a node does not have: a second node to pair with (no multi-node
jobs, no RDMA or scale-out fabric between pool nodes) and no notebook
server (ssh is the only way in).
Asking Marshall¶
Marshall is the one door: there is no console form, no API route and no MCP tool for a request, and only a member of a public research organization is offered one. Marshall quotes before it asks: the hourly figure for the node, the worst case for one full lease (lease length × GPUs × rate), the balance a request needs, and whether your job fits one node. A job that needs more than one node belongs on the market, not here. Every request ends in a confirmation card, Full Access or not, because a request starts a meter.
With two node shapes open, Marshall lists both offers, each with its rate per GPU-hour, the worst case for one full lease and the current queue, and marks one Recommended from your job's software stack: a CUDA stack points to the NVIDIA B300, a ROCm stack to the AMD MI355X; when both fit, the cheaper offer, then the shorter queue. The mark is a mark, not a choice: you pick either row, and the pick is the confirmation. A request must name one offer; one that names none is refused ("site required") with the offers listed, and one naming an offer that is not open is refused with what is open. While only one shape is open, Marshall quotes it alone and asks as before.
A request needs a balance that covers the first hour of a lease at the
quoted rate (GPUs × rate), not the whole lease; past that hour the meter
runs against your balance like any other charge. A public research
organization starts with $100 of credits (the airdrop deposit on its
ledger); once those are spent a request is refused until credits are
bought on the console's Billing page
(card or wire); Marshall names the amount
it needs. A request is also refused while a
billing hold is active.
The queue¶
Requests are served first come, first served, one node per request. A
request is queued from the moment it is accepted; Marshall reports
your place in line first and then an estimated wait. The estimate takes
the nodes in the pool, the requests ahead of you and the time left on
the leases running now, and assumes every lease runs its full length; a
release before that moves every estimate behind it earlier, so it is an
estimate and never a promise. The line itself moves only when a lease
ends or an idle node is added. When a node is free for you,
handover takes seconds: the platform deposits Marshall's public key on
the node and hands back its address and host key. Nothing is installed
and nothing is rebuilt on the way in.
| State | Meaning |
|---|---|
held |
accepted, waiting behind another request of yours; enters the queue at the back when that one ends |
queued |
in line; position is your place |
provisioning |
a node is being handed over; seconds |
active |
the node is yours and the meter is running |
ended |
the lease is over; end_reason is released, preempted, lost, balance (the organization's balance reached zero) or admin |
You hold one node at a time (currently one queued or active request per
user). A request from a second Marshall session while a first is live is
accepted as held, outside the line, and joins the back of the queue
when the live request ends; a second request from a session that is
already waiting is refused, and held requests are capped (currently
three per user). The console's sessions rail shows which session is
waiting and which one has a node, and the console's Public Research page
shows the same per GPU type: a slot for each type that holds the session
using a node of that type, with the sessions waiting for one listed
beneath it, each with its place in line and estimated wait; a click on a
session opens it. Each type's heading shows how
many GPUs the pool currently holds for it, and the console's Grid entry
shows the total across types. An empty slot offers a button naming the
GPU type, "Get MI355X with Marshall" or "Get B300 with Marshall": one
click opens a new Marshall session that
requests a node of that type, walks you through how the pool works and,
once the node is yours, opens a shell on it and shows its GPUs, with the
usual confirmation before anything is requested. Under the button the
slot shows the estimated wait a request made now would get, before you
have asked for anything.
The lease¶
A lease is 4 hours (currently) and is preempted only when someone is waiting. While nobody is queued the lease runs past 4 hours and keeps billing until you release. Once another request is queued and your lease has passed its length, a preempt notice is set: 10 minutes (currently) to checkpoint, then the node is taken and rebuilt. Marshall wakes you 30 and 10 minutes before the lease length is reached and the moment a notice is set, with the time left and a reminder to copy results off the node.
Release when you are done: the end is stamped before the rebuild starts,
so the recycle costs you nothing. Marshall never releases a node on its
own: when a job finishes it says the meter is still running and asks you;
releasing, like requesting, goes through a confirmation card. A node the platform loses (a failed
node, a failed handover or rebuild) ends the lease at the earliest
evidence of the failure, and billing with it. A lease also ends the
moment the organization's balance reaches zero: the node is taken back,
the meter stops and end_reason reads balance; add credits before
requesting again. National Compute can end a lease; end_reason then
reads admin.
To keep working past a lease, ask Marshall for another node from the
session that holds it: the new request is held, joins the back of the
queue when the lease ends, and lands on a freshly rebuilt node. Nothing
carries over from the old node.
Working on the node¶
Marshall reaches the node with plain ssh on its public address, port
22, as the login user the request names. The one key is minted by
Marshall on your workspace disk and reused for every request; the node's
host key arrives with the node and is pinned before Marshall announces
it, so the first connection is verified rather than trusted on first
use. A shell pane in your workspace carries the same key.
Long jobs detach on the node (nohup, tmux): Marshall's own command
runs are time-capped and end with the session. Copy results back with
scp or rsync before release.
The first job on a node is the platform's nanoGPT starter, started on the node with one line and detached from the ssh session:
ssh <ssh_user>@<public_ip> 'curl -fsSL https://access.nationalcompute.com/first-job/nanogpt-node.sh | sh'
It trains in the stock PyTorch container for the node's GPUs (the node
runs it with docker), writes its log to ~/first-job/nanogpt.out
there, and ends on its own after about an hour; the node keeps billing
until you release it. The
starter recipe catalog
lists it with cluster_kind public_research, and
FIRST-JOB.md walks an agent
through request, connect, train, copy back and release. Your own
container runs the same way with docker run on the node. The Serve and
Post-train recipes are Kubernetes manifests and do not apply here.
Wiping your workspace disk mid-lease takes the key with it; the lease keeps running, and the way back in is release and a new request.
Billing¶
A lease bills per GPU-hour, for every GPU in the node, at the rate shown
in your Marshall quote and on the console's Billing page under Rates
(label "Public Research ·
Charges land on your organization's credit ledger about every five
minutes, in the same credit unit as every other charge, as rows of
source public_research whose entries name
the node, the GPUs, the rate and the window. The Billing page shows them
under the Public Research chip; the
billing summary carries the month's
GPU-hours and charged amount as public_research. A free rate still
counts GPU-hours and writes no charge.
A request is admitted only with a balance covering the first hour of a
lease, and a billing hold refuses new
requests. A balance that reaches zero ends a running lease at once: the
node is taken back and the meter stops (end_reason balance), so a
lease never runs far past what the balance covers. When nobody is
waiting, a lease runs and bills past its length with no other cap.
Release when the job is done.