Skip to content

Burst capacity

You are talking to a market. You state what you want and the most you'll pay, and the market decides what you hold, tick by tick. The one way to book capacity outright is base load capacity: a fixed block of nodes for a fixed term at a fixed rate, outside the auction.

Declare, don't reserve

For a VM cluster, you declare two numbers with PUT /api/vm/capacity:

  • max_gpus — how many GPUs you want. Capacity is granted as whole nodes, so this must be a multiple of the gpus_per_node value the API returns (off-grid values are refused, never rounded).
  • max_price_per_gpu_hour — the most you'll pay per GPU-hour, in USD.

For a Kubernetes cluster there is even less to say: your GPU jobs create the demand by themselves. Each job is one request priced as a gang in whole nodes on the market. PUT /api/k8s/bid sets your price ceiling. See Kubernetes clusters for how that plays out.

The market clears

Capacity is auctioned in short ticks (roughly every 10 seconds). Each tick re-runs the auction over all demand in your market:

  • While the market clears at or under your ceiling, nodes are granted toward your declared capacity.
  • When it clears above your ceiling, nodes are reclaimed.
  • You are billed what the auction assesses, never above your ceiling — usually less, though during a protection window the assessed rate can be as high as your ceiling. There is no single posted price: rates are set per cluster (see below). See Billing.

Because grants follow the market, holding capacity is not guaranteed indefinitely: a ceiling under the going rate leaves declared capacity pending and can shed nodes you hold. Treat node-local data as ephemeral — a reclaimed or swapped node is destroyed.

Inside the auction

The mechanism is a repeated sealed-bid combinatorial auction:

  • Winner determination maximizes total value. Demand competes in whole-node units; multi-node groups are atomic — granted whole or not at all, never split.
  • Prices are second-price in nature. A winner pays for the demand it displaces, not its own bid (a VCG-style payment with a core adjustment), and your own cluster's losing demand never sets your price. Bidding your true ceiling is the honest, safe strategy — overbidding can't lower your rate, and you never pay above your declared ceiling. The corollary: raising your ceiling also raises the cap on what you can be assessed for nodes you already hold, so displaced demand your old ceiling was hiding re-prices immediately. Pick your true maximum once rather than probing upward in small steps — several small raises cost strictly more than one honest one.
  • Prices are per cluster, not uniform. Different clusters can be assessed different rates in the same tick — which is why the market feed publishes statistics (a volume-weighted mean and the min/max spread) rather than a single quote.
  • Arrival order is only a tie-break. There is no first-come, first-served queue: price decides, and at equal bids the older demand wins.
  • A bid can be too low to ever clear. The market refuses a declare that could never win capacity (bid-too-low), and demand priced that low simply waits. What wins is published, never posted: read the market feed and the price-to-win ladder, and bid against the going rate.

Minimum-duration protection

A freshly granted node carries a protection window (currently one hour for a VM node, two hours for a Kubernetes node), counted from the moment the node is usable: while it is open, no competing bid can take the node — the grant is a property right, not a standing auction entry. In exchange, during the window you can be assessed up to (but never above) your limit price: on a VM cluster the bid that won the grant, locked for the window even if you lower your ceiling; on Kubernetes the live limit price (below) — a raise lifts the cap, and a cut below the price the job held when its gang started voids the window. When the window lapses, the node competes normally again.

For Kubernetes clusters the window belongs to the node grant, not to the job:

  • The first hour is a minimum hold. A granted node stays in your cluster — and bills — for at least one hour, even if the job that asked for it finishes sooner; idle release starts only once that hour has run, and after it charges accrue in minimum increments of five minutes. The hold is deliberate on both sides: it gives you a stable window to work in — a job that crashes restarts on capacity you still hold instead of re-entering the auction — and it keeps sub-hour churn from misusing capacity others are waiting for.
  • Jobs come and go under it. A job that finishes inside the window leaves the window on the node, which holds for a five-minute idle grace billed at that job's rate — provided the market would still clear that rate: the grace only holds capacity at a price that clears, so a node whose job's bid could never clear is released the moment it goes idle. A job that takes the node within the grace inherits the remainder and never extends it. Once the minimum hold has run, a node left idle past the grace is released, window and all.
  • The limit price stays live. You may change a job's limit price at any time; an increase is always fine, and a decrease below the price the job held when its gang started (Provisioned=True) voids the window — while the limit price sits below that price, the job's nodes are preemptible at once. Raising it never resets the window.
  • A lost node is owed a replacement. If a protected node stops responding to your cluster (its kubelet goes dark), your cluster is owed a replacement node for the remainder of the window (wall-clock — provisioning time counts against it). The right ends early only when the market reclaims the node, when the node is released after sitting idle past the grace, or when an operator moves it out of your cluster.
  • A node we take out of service is replaced, not billed. When our health checks cordon one of your nodes, the market treats it as undelivered from that moment: billing for it stops, your job reads short by one node and the market grants a replacement that inherits the remaining window at the locked rate. The faulty node stays in your cluster, cordoned, while we investigate; your pods on it are evicted after a short grace so they reschedule onto the replacement. Once our checks clear it, the node returns to your cluster as idle capacity and is released after the idle grace unless a job takes it. A node you cordon yourself is your own decision: it keeps billing and earns no replacement.
  • Growth is unprotected. A gang is granted whole, so a job is never half-protected; a new job, or a Deployment's added replica, competes at your live ceiling until its own grant lands, and its window starts there.
  • A base load block being delivered outranks the window. When another account's base load block is owed nodes and the island has no free ones, the market takes idle nodes on the cheapest standing bids first, then busy ones — inside an open window too. Your pods get the standard notice and the job re-enters the auction as pending demand. The reservable inventory is capped per island precisely so this stays rare.

Base Load Capacity

Base load capacity is the committed product beside the market: a fixed block of GPU nodes, yours alone for a fixed term at a fixed rate. No bidding and no preemption. A block bills for the full term whether you use the nodes or not. It carries the Base Load Capacity SLA. The console's Burst capacity page shows it in the Base Load section, above the Preemptible section that holds the auction.

  • What you get. Your cluster holds at least the block's node count for the whole term. If a node in the block fails, the market delivers a replacement ahead of every bid on the island; time spent provisioning it counts against the SLA, not against your term.
  • Node specs. On the island that offers base load today, each node has 240 vCPUs (AMD EPYC, Turin), 8× AMD MI355X 288 GB OAM GPUs, 3 TB RAM, 8× 3.84 TB NVMe and a 3200 Gbps scale-out network. All nodes are fully interconnected with each other and with the on-site scalable NFS, and preemptible burst nodes share that fabric with your base load GPUs and storage. The console lists the same lines under Interconnected GPU Node Specs in the Base Load section.
  • Your jobs and the block. Submit jobs exactly as on any Kubernetes cluster; nothing in a manifest names the block. On each auction tick the market counts your cluster's jobs against the block's GPUs: jobs already running come first, then waiting jobs by priority class, oldest first within a class. A job that fits the free base load capacity is placed on the block's nodes without bidding and without a balance, within a couple of minutes of submission. A job that does not fit never holds up smaller jobs behind it; they are placed first, and the base load capacity left over still counts toward the larger job, which bids on the market only for the rest. When base load capacity frees, a job that was waiting on the market is placed on it, and a job already running on a market node is folded into the block where it stands, with no restart, so its market billing stops.
  • Small jobs on the block. Where the platform has enabled packing onto the block for your cluster, a job smaller than one node is placed on the block by GPU count. Several small jobs share one block node. A small job that does not fit right now waits for GPUs of the block to free. It never bids on the market. Whole node jobs take an empty block node first. A small job opens an empty block node when no whole node job is waiting for it. After a block node has sat empty for the idle grace, a small job that waited that long takes it. A job whose pods are each smaller than one node and together exceed one node is not placed. It waits with the reason ReservedTooSmall until you resubmit it with whole node pods or with fewer pods. The market never moves a running small job to another block node. Without a live block a small job waits with the reason ReservedBusy until a block starts.
  • Burst beyond it. Demand above the block's node count is ordinary preemptible burst capacity at your live ceiling. A job smaller than one node is never burst demand. Nothing else changes: your jobs create the demand, and Kubernetes decides which pods yield when a market node is reclaimed. The two-hour balance rule applies only to the part of a job that runs beyond the block.
  • Buying. Any member of your organization buys a block on the console's Burst capacity page: the Buy Base Load Capacity card in the Base Load section offers fixed sizes (in GPUs, always whole nodes) for the posted terms, with each term's rate per GPU-hour. Pick one cell and the card shows the whole block's price; Buy completes the purchase. A size and term with no posted rate reads sold out; a size the island cannot fit right now reads sold out until the date enough earlier blocks lapse; a size larger than the island's reservable inventory reads not available on this island. When your credit balance covers the block's whole price, nothing more is needed. Otherwise buying needs a security deposit on your credit balance, the block's price up to $10,000, and you wire the remaining amount by the next business day. The block bills across its term. A block starts immediately and appears under Your Base Load Capacity. A cluster holds one live block at a time; a second purchase for the same cluster is refused (one-per-cluster) until the first ends, and the page folds the buying card behind one line while a block is live. On an island where buying on the page is not open, the card says so and your account team arranges the block.
  • Changing it. Changes are made with your account team. A change to the rate or the size ends the current block at that moment and starts a new one with the new figures, so billing is exact on both sides; the end date and notes can be changed in place.
  • Billing. The block's rate is charged to your credit balance across the term, in the same five-minute cadence as everything else on Billing, for every node in the block — idle or busy, delivered or being replaced. Work inside the block is never metered by the market, and the block's nodes never appear in the market's clearing rates.
  • At the end of the term. There is no automatic renewal. At the end date the block's nodes stay in your cluster as ordinary market members and step down under the standard protection window from that moment, locked at your declared price — or, when you have not declared one or yours is too low to ever clear, at the lowest rate the market clears. A node running your work keeps running until the window closes and bills the locked price. An idle node returns to the market after the usual idle grace; the step down carries no minimum hold, because the term you paid for is over. When the window closes the node competes at your live ceiling on each auction tick, and a ceiling too low to ever clear releases it. A job smaller than one node does not step down with its node. At the end date it is requeued at once as a new request. That request waits with the reason ReservedBusy until your cluster holds a live block again. It never moves to a market node. Arrange a renewal with your account team before the end date if you need continuity.
  • Ending early. A live block keeps running when your balance goes negative, while the billing hold reclaims your other market capacity as usual. When the balance has stayed at or below zero for 24 hours, the block is cancelled at that moment: its nodes return to the market through the standing hold, billing for the block stops, and you can buy again once the balance is restored. The market's own capacity, protection windows and pricing are untouched by blocks you do not hold.

A bid that can never clear

A declare too low to ever clear the market is refused outright (bid-too-low): pending demand costs nothing, but a bid the market can never meet means nothing will ever arrive, so the write refuses it where you can see why. A Kubernetes cluster with no limit price effectively bids $0 — its jobs create demand, but nothing is granted until you set a real ceiling.

Market conditions can also move past a standing bid (the write-time check never re-runs). The capacity and limit price reads then report bid_too_low: true, and inside a Kubernetes cluster the verdict lands on the job itself as a BidTooLow event. What to bid is read off the market's own numbers — the market feed and the price-to-win ladder — never off a posted price. See how the bid prices the market for the Kubernetes mechanics.

Prices come back to you

The market price feed publishes what capacity actually clears at, tick by tick and as a long-run series, so you can tune your ceiling against the going rate instead of guessing.

Scaling down is always free

Lowering max_gpus needs no market's permission and always works — declare 0 to release everything. The market gates only what you acquire, never what you give back. On scale-down you can nominate which nodes go first (the release list on the capacity declaration), and released capacity simply stops billing — there is nothing to resell or wind down.