Skip to content

FAQ

Allocation and base load

How does the allocation model work — is it first-come, first-served, or guaranteed reservations?

Neither: it's price-ordered. Every auction tick (roughly 10 seconds) re-runs the market over all demand, and capacity goes to the highest-value bids. Arrival order matters only as a tie-break — at equal bids, the older demand wins — and there is no queue to hold a place in. What plays the role of a guarantee is the minimum-duration protection window: a freshly granted node cannot be taken by any competing bid while its window (currently one hour on VM clusters, two hours on Kubernetes clusters) is open. Beyond that window, you keep capacity by keeping your ceiling at or above the going rate. The one guaranteed product is a base load block: a fixed block of nodes for a fixed term at a fixed rate, outside the auction.

Can a reservation be secured within ~36 hours of a request?

For burst capacity, faster than that: declared capacity is granted at the next auction tick in which supply exists and your ceiling clears — seconds to minutes when the market has room. For a committed base load block, any member of your organization buys it on the Burst capacity page where buying is open. Elsewhere your account team arranges it. A block starts as soon as it is sold and its nodes are delivered ahead of every bid on the island.

Is this on-demand — do we specify a time length?

For burst capacity, no fixed term in either direction. You hold capacity as long as you want it and release it whenever you want (scale-down is immediate; a Kubernetes node still bills its one-hour minimum). The only time-shaped elements are the protection window and the one-hour minimum hold on fresh grants. If you want a committed term, a base load block is sold for the posted terms at a fixed rate.

How do we do capacity planning under this system?

Three instruments: the market feed gives you the going rate per tick and up to a year of history, so you can see what a given ceiling would have held historically; the capacity read echoes your own position (declared, granted, pending, and whether a standing bid can still clear); and your worst-case spend is arithmetic — held GPUs × your ceiling. The practical loop is to pick the ceiling that clears at the rate you can see, then watch pending — persistent pending capacity means your ceiling is under the market or supply is short.

What are the SLAs?

Base load capacity carries the published Base Load Capacity SLA. Preemptible burst capacity has no formal SLA; what the mechanism itself guarantees there: the protection window on fresh grants, billing at the clearing rate and never above your declared ceiling, and unrecorded gaps are never billed. For other contractual availability or support commitments, contact us.

Pay for what you use

Are we only charged for compute that's running jobs?

You're charged for capacity you hold, metered at market-record boundaries — see Billing. For Kubernetes clusters that converges on "charged when running work, plus a five-minute tail" automatically: nodes join only when a queued job pulls them and, once the job ends, shed after a five-minute idle grace billed at that job's rate unless another job takes the node — so the meter tracks the queue. For VM clusters you pay for nodes while you hold them, working or idle — holding is your choice, and lowering max_gpus stops the meter.

If unused reserved capacity is pushed back into the pool, do we get paid even if nobody bids on it?

For burst capacity the situation can't arise: you never own capacity you have to resell. Released capacity (or a Kubernetes node shed for idleness) simply returns to the platform pool and your billing stops at that moment — whether anyone else bids on it afterward is the platform's problem, not yours. A base load block is different by design: you pay the fixed rate for the block whether the nodes are busy or idle, and there is no rebate for idle nodes in the block — the fixed rate and the no-preemption guarantee are the trade.

Pricing index

How are prices set — is there an index?

Pricing here is not computed from an external index at all. Prices are set by the auction: what you pay is determined by the demand competing with yours, never above your own ceiling. The platform's price index is its own market feed — the actual volume-weighted clearing rates, published per tick and as a long-run series. Other providers' list prices are not an input to the price.

Does any external price feed the index?

No — see above. No external price feeds into what you pay.

Is the strike price the minimum across available providers?

There is no strike price. Each tick's clearing outcome is set by competing demand, capped by your own declared ceiling. The market feed publishes what actually clears.

Pricing and auction mechanics

What is the exact auction structure — generalized second price, VCG, something else?

A repeated sealed-bid combinatorial auction, roughly every 10 seconds. Winner determination maximizes total value across all demand in whole-node units (multi-node groups are atomic — granted whole or not at all). Payments are second-price in nature: a winner pays for the demand it displaces — a VCG-style payment with a core adjustment — and your own cluster's losing demand never sets your price. A bid too low to ever clear the market is refused on write (bid-too-low); the price to win ladder says what clears. Practical consequences: bidding your true ceiling is the safe strategy, you never pay above it, and prices are per cluster rather than uniform — which is why the market feed publishes a volume-weighted mean with a min/max spread rather than a single quote.

My limit price seems reasonable — why is nothing clearing?

On Kubernetes, read the request's reason before guessing — kubectl describe job, or the console's Workloads page (queue events): Outbid names the limit price that wins right now, BidTooLow means the request's limit price — the job's label if set, else the cluster's limit price — can never clear the market as it stands (the market moved, or the label undercuts the cluster's price), BalanceTooLow means the org's balance cannot fund two hours of the request, PendingSupply means protected nodes hold the supply your gang needs. The other causes are the mundane ones. When the site is genuinely out of capacity, no bid clears it: the console's Workloads page shows capacity unavailable, raising the limit price does not start the job sooner, and the queue clears on its own when capacity returns (when the site has no capacity). Market-wide, the feed's ticks show null when nothing clears. The last mundane cause is no limit price set at all: no bid means $0, which never clears.

We already hold capacity — how do we flex up on short notice?

Raise max_gpus (VM) or submit more GPU jobs (Kubernetes). The new demand competes at the next tick and grants land in your existing cluster, next to what you already hold. Short notice is the norm here: there's no procurement step, just the market clearing.

How do we get money back for burst compute sitting idle over a weekend?

You don't get money back — you stop spending. Scale to zero (or just down) on Friday and billing stops with the release; re-declare on Monday and you re-enter the market at Monday's price. Two caveats to weigh: released nodes are destroyed (anything on local disk is gone), and re-entry is at market conditions, not a held price. Kubernetes clusters do the weekend version automatically — an empty queue sheds nodes after the idle grace. A base load block bills through the weekend: its fixed rate covers the block for the whole term, used or not.

Is burst capacity won at auction guaranteed to land on the same cluster as our existing reservation?

Yes. Capacity is granted to the cluster that declared the demand — a grant is never "somewhere else." A cluster trades in a single market (its site), so growth lands beside what you hold.

Isn't preemption disruptive? We can't instantly drain a node.

The mechanism gives you five layers before disruption, and visibility when it comes:

  1. The protection window — a fresh grant can't be preempted for its first hour (VM node) or first two hours (Kubernetes node), no matter what the market does.
  2. Your ceiling is the lever — you are only reclaimed when the market clears above your declared ceiling. Work that must not be interrupted gets a ceiling above the going rate, which you can watch on the market feed.
  3. You pick the victims — on scale-down, release nominations choose which nodes go first, and planned_release on every read shows exactly what a shrink would destroy before you commit it.
  4. Kubernetes requeues itself — a reclaimed job is suspended with its whole gang and requeued as a new request at your current limit price; it resumes when granted again, and with the retry-budget spec in place a reclaim never fails the Job.
  5. A faulty node is replaced — if our health checks take one of your nodes out of service, billing for it stops, the market grants a replacement that inherits the remaining protection window, and the pods on the faulty node are evicted after a short grace so they reschedule. The node's tile reads Node Fix In Progress and the request shows Replacement Pending until the new node joins (details).

Beyond that, this is genuinely preemptible capacity and the honest advice applies: checkpoint long-running work and keep durable state off node-local disk.

Access model

How do I get an account?

Sign in at nationalcompute.com with Google or with an email address that receives a sign-in link. A .edu, .mil or .gov address gets its own public research organization on the spot, with $100 of credits to get started, which requests whole GPU nodes from Marshall through Public Research; an address an existing organization's rules admit joins that organization; anything else is placed on the waitlist and emailed when a spot opens — early access for public users is available in waves (Sign in and create an account).

Do we submit containers or jobs that you run, or do we get machines we operate ourselves?

You operate it yourself, in both models. A Kubernetes tenancy is a dedicated cluster — you hold the kubeconfig and deploy your own workloads; the platform only moves nodes in and out as the market clears (Kubernetes clusters). A VM tenancy gives you nodes over ssh with your registered keys installed (VM capacity API). There is no "submit us a job and we run it" surface.

Are VM access and Kubernetes priced differently?

No — same market, same unit. VM and Kubernetes demand clear in the same per-site auction, both priced in USD per GPU-hour against your declared ceiling. What differs is the interface: VM capacity is declared explicitly, Kubernetes demand derives from your jobs.

How does the bare-metal offering compare to Kubernetes and VM access?

The Kubernetes offering is the bare-metal offering: cluster nodes run the Kubernetes agent directly on the metal, with no hypervisor layer, dedicated to your cluster while you hold them. The VM offering runs virtual machines on the same class of GPU hosts and trades that layer for ssh-level, operate-it-yourself access. Pick by how you want to drive the hardware, not by price — both clear in the same market.

Ancillary costs and storage

Are there fees for logging, monitoring, or ingress/egress on Kubernetes jobs?

No. Billing today is GPU capacity only — the clearing rate for nodes you hold (Billing). Monitoring dashboards, metrics, and logs come with the tenancy at no separate charge, and network traffic is not currently metered or billed.

How much storage does a Marshall workspace come with?

Two pieces, both included with the tenancy. Each member's workspace has its own 50 GiB disk, mounted at /data: home directory, Marshall's notes, uploads, anything you want to keep between restarts. An organization with a workspace hub of its own also has a shared folder at /org, visible to every member, capped at 200 GB. Over the cap, the folder stays readable and members can delete to free space, but new writes fail until the organization is back under it; writes return on their own a few minutes later. Nothing is deleted by the platform. Public research organizations are single-member and have no shared folder.

Can you support multi-petabyte storage, and is storage priced off the same index?

Storage is per-site: where a shared filesystem is provisioned, your capacity read's environment field carries the paths and onboarding facts, and it is not priced off the market — GPU capacity is the only market-priced resource today. For multi-petabyte requirements, talk to us about the specific site — large storage is a provisioning conversation, not a self-serve knob.