Telemetry and flags¶
metrics_read answers what the platform's own monitoring saw on the
hardware your job ran on: per GPU temperature, clocks, activity, power,
memory, throttle time and error counters, plus the host's CPU, pressure
and network. It also answers the platform's judgement on those readings
as flags, each one carrying the evidence behind it.
The tool lives on the MCP server. It is read only and single
phase, so no confirm_token is involved. The
cluster page reads serve the console's own
utilization and machine panels over REST; this tool is the agent-facing
read of the same monitoring, with the flags layer on top. For anything
these three actions do not answer, metrics_query
below runs one PromQL expression of your own over the same monitoring.
action |
Target | Answers |
|---|---|---|
node |
one node | every reading the platform holds for that node's GPUs and host |
workload |
one workload | where it ran over the window, plus each node's readings |
flags |
a node or a workload | the platform's judgement: which conditions fired, with evidence |
Arguments¶
| Argument | Values | Applies to |
|---|---|---|
action |
node, workload, flags |
required on every call |
cluster |
the cluster name (grid_whoami lists them) |
required on every call |
node |
the node's name | action=node; action=flags on a node |
kind |
Job, JobSet, MPIJob, Deployment, StatefulSet, Pod |
action=workload; action=flags on a workload |
name |
the workload's name | the same two |
namespace |
defaults to default |
workload targets |
range |
1h, 6h, 24h, 7d; defaults to 6h |
optional |
start, end |
ISO 8601 UTC; end defaults to now |
optional |
detail |
summary or series |
optional |
points |
series length after downsampling; defaults to 48, maximum 96 | optional |
action=flags takes either node or the pair (kind, name).
Anything else is a 422 bad-request.
The window¶
range names a window ending now. start and end set an explicit
window instead, and the call then ignores range. A span longer than
14 days is refused. A start at or after end is refused.
The bucket width follows from the span and the points you asked for:
the span divided by points, rounded up to a multiple of 30 seconds,
never below 30 seconds. A 6 hour window at the default 48 points gives
a step_s of 450. The answer's window echoes the resolved start,
end, step_s and the resulting points.
detail defaults to series for action=node. It defaults to
summary for action=workload and action=flags. Under summary
each series collapses to its {mean, min, max}, except temp_c,
tensor_active_pct and power_w, which keep their values.
The envelope¶
Every action answers the same outer object:
| Field | What it carries |
|---|---|
cluster, gpu_vendor, gpu_model |
the cluster and the GPU class its nodes carry |
window |
start, end, step_s and points, as resolved |
reference |
the comparison number for the GPU model, or null |
coverage |
present and missing: which metric families answered |
note |
the units and how to read this answer, in one short paragraph |
A series is
{"t0": <epoch seconds>, "step_s": n, "values": [...], "mean": x, "min": x, "max": x}.
A gap inside values is null. A series the platform holds no reading
for is null whole. A null is never a zero. coverage.missing names
every family that did not answer.
The family names are a fixed vocabulary: gpu_temp, memory_temp,
sm_clock, mem_clock, gpu_util, tensor_active, sm_active,
dram_active, power, hbm, throttle_violations, clock_events,
link_errors, ecc, xid, host_cpu, host_pressure, host_nic,
north_south, health_episodes.
action=node¶
{
"cluster": "aurora-prod", "gpu_vendor": "amd", "gpu_model": "MI355X",
"window": {"start": "2026-09-21T12:00:00Z", "end": "2026-09-21T18:00:00Z",
"step_s": 450, "points": 48},
"reference": {"gpu_model": "MI355X", "temp_p95_c": 74.5,
"note": "p95 of GPU temperature for this model over the window, across your nodes and the island's shared pool"},
"coverage": {"present": ["gpu_temp", "memory_temp", "tensor_active",
"power", "hbm", "ecc", "link_errors",
"host_pressure", "…"],
"missing": ["north_south"]},
"node": {"name": "aurora-prod-n2", "gpus": 8, "state": "ready"},
"gpus": {
"3": {
"temp_c": {"t0": 1758456000, "step_s": 450,
"values": [71.0, 83.5, 91.0, 88.2, null],
"mean": 83.4, "min": 68.0, "max": 91.0},
"peak_temp_c": 91.0, "memory_temp_c": {"…": "…"},
"tensor_active_pct": {"…": "…"}, "power_w": {"…": "…"},
"hbm_used_gib": {"mean": 241.6, "min": 238.0, "max": 244.9, "…": "…"},
"hbm_total_gib": 268.2,
"throttle": {"thermal_s": null, "power_s": null, "clock_events": null,
"…": null},
"errors": {"link_crc": null, "link_replay": null, "link_recovery": null,
"pcie_replay": 0, "pcie_recovery": 0, "ecc_sbe": 0,
"ecc_dbe": 0, "xid": null}
}
},
"host": {"cpu_busy_pct": {"…": "…"}, "nic": {"errors": 0, "drops": 0, "…": "…"},
"pressure": {"cpu": {"…": "…"}, "memory": null, "io": {"…": "…"}}},
"health_episodes": [{"kind": "thermal", "started": "2026-09-21T14:05:00Z",
"ended": "2026-09-21T14:46:00Z", "detail": "…"}]
}
gpus is keyed by the GPU's index on the node:
| Field | Reading |
|---|---|
temp_c, memory_temp_c |
GPU temperature and memory temperature, °C |
sm_clock_mhz, mem_clock_mhz |
core clock and memory clock, MHz |
util_pct, sm_active_pct, dram_active_pct |
how busy the GPU, its cores and its memory were, percent |
tensor_active_pct |
tensor core activity, percent: the reading that tracks training work |
power_w |
power draw, watts |
peak_temp_c |
the highest temperature the GPU reached in the window, °C |
hbm_used_gib, hbm_total_gib |
GPU memory in use, and the GPU's total, GiB. The total is the denominator memory_pressure measures against |
throttle |
seconds of throttle by cause over the window: thermal_s, power_s, board_limit_s, sync_boost_s, low_util_s, reliability_s; clock_events counts each clock-event kind the GPU reports, for example HW_SLOWDOWN |
errors |
counter increases over the window: link_crc, link_replay, link_recovery, pcie_replay, pcie_recovery, ecc_sbe, ecc_dbe, xid |
link_* counts the GPU-to-GPU fabric links. Those are the NVLink
counters on NVIDIA. The AMD equivalents (XGMI) are not read today, so
link_* is null on an AMD cluster.
pcie_recovery counts PCIe link recoveries over the window. It is
null where a maker has no such counter, which is NVIDIA today.
Throttle seconds and error counts are the increase over the window,
never a lifetime total. A counter the platform does not have for a GPU
maker reads null, field by field. The cluster above is an AMD
cluster, which is why every throttle timer, every link_* counter and
xid read null on it. Its PCIe and ECC counters answer.
host carries the node itself:
| Field | Reading |
|---|---|
cpu_busy_pct |
host CPU busy, percent |
load_per_core |
run queue length per core |
pressure.cpu, pressure.memory, pressure.io |
share of time the host waited on CPU, on memory, or on disk, percent |
nic.rx_mbps, nic.tx_mbps, nic.errors, nic.drops |
node network throughput, and error and drop counter increases over the window |
north_south.ingress_mbps, north_south.egress_mbps |
traffic to and from the internet |
health_episodes lists the platform's own recorded episodes for the
node inside the window, as {kind, started, ended, detail}. An open
episode reads "ended": null.
action=workload¶
{
"workload": {"kind": "Job", "namespace": "default", "name": "train-llm"},
"placement": [
{"pod": "train-llm-0", "node": "aurora-prod-n1",
"from": "2026-09-21T12:04:00Z", "to": null, "phase": "Running"},
{"pod": "train-llm-1", "node": "aurora-prod-n2",
"from": "2026-09-21T12:04:00Z", "to": null, "phase": "Running"}],
"placement_source": "history",
"nodes": {
"aurora-prod-n1": {"gpus": {"0": {"temp_c": {"mean": 68.1, "min": 61.0,
"max": 72.4}, "…": "…"}}},
"aurora-prod-n2": {"…": "…"}}
}
placement comes from the workload's own scheduling history,
intersected with the window. placement_source then reads history.
Where the platform holds no history for that window, placement comes
from the pods running now, and placement_source reads current-pods
instead. Read placement_source before you trust a placement.
nodes holds one node block per node the workload touched, at
summary detail unless you ask for series. Each node block carries
its own health_episodes, the same shape as on action=node.
action=flags¶
A one hour window over the same node, opened after the episode above had already started:
{
"window": {"start": "2026-09-21T14:20:00Z", "end": "2026-09-21T15:20:00Z",
"step_s": 90, "points": 40},
"target": {"workload": {"kind": "Job", "namespace": "default",
"name": "train-llm"}, "placement": ["…"]},
"checked": ["thermal", "utilization", "comm", "host", "memory"],
"unavailable": ["throttle_violations", "clock_events", "xid"],
"clean": false,
"flags": [
{"kind": "thermal_throttle", "severity": "high",
"node": "aurora-prod-n2", "gpu": "3", "minutes": 26,
"from": "2026-09-21T14:20:00Z", "to": "2026-09-21T14:46:00Z",
"clipped": {"start": true, "end": false},
"episode": {"from": "2026-09-21T14:05:00Z",
"to": "2026-09-21T14:46:00Z"},
"evidence": {"peak_temp_c": 91.0, "reference_p95_c": 74.5},
"summary": "GPU 3 on aurora-prod-n2 held at or above 83 °C for 26 min inside this window (peak 91 °C, reference p95 74.5 °C); the episode began 14:05 UTC, before the window"}],
"note": "this GPU maker reports no throttle timers, clock events or XID events, and no fabric link counters within link_errors; one flag is clipped at the window start"
}
Each flag object:
| Field | What it carries |
|---|---|
kind, severity |
the condition, and how serious this instance of it is |
node, gpu |
the node, and the index of the GPU the evidence came from |
from, to, minutes |
the sub-window the evidence covers, and its length |
clipped |
{start, end}: whether the evidence runs past that edge of the window |
episode |
the platform's own health episode bounds for this node and kind, or null |
evidence |
the numbers behind the flag, whichever of them the GPU maker reports: peaks, counter increases, clock events, the reference p95 |
summary |
one sentence an agent can quote |
clipped.start is true when the evidence was already abnormal in the
window's first bucket. clipped.end is true when it was still
abnormal in the last. Either way from, to and minutes cover only
the part inside the window, so the length is a floor. The answer's
note names every clipped flag.
episode carries the bounds of the platform's own
health episode for that node and kind, where one overlaps
the flag. to is null while the episode is still open. The field is
null where no episode overlaps. Quote the episode bounds when a narrow
window cuts a flag short.
The six kinds, and the rule that fires each one:
kind |
Fires when | severity |
|---|---|---|
thermal_throttle |
thermal throttle time increased, a thermal clock event landed, or the GPU held at or above 83 °C for at least 5 minutes | high at 10 minutes or longer, otherwise warn |
power_cap |
power throttle time increased, or a power clock event landed | high at 10 minutes or longer, otherwise warn |
utilization_collapse |
tensor core activity stayed below half the window's median for at least 10 minutes while the window's pods were Running | warn |
comm_errors |
a fabric link or PCIe error counter increased, an uncorrectable memory error landed, or a GPU fault event landed | high for uncorrectable memory errors and GPU fault events, otherwise warn |
host_pressure |
the host waited on CPU, memory or disk more than 20% of the time for at least 5 minutes | warn |
memory_pressure |
GPU memory in use held at or above 95% of the GPU's total for at least 5 minutes | warn |
utilization_collapse falls back to plain GPU utilization where the
platform has no tensor core reading for that GPU maker.
What a GPU maker cannot report changes what can fire. On AMD today the
platform reads no throttle timers, no clock events and no XID events.
Inside link_errors it reads the PCIe counters; the fabric link (XGMI)
counters are absent, so link_* is null and link_errors answers
only in part.
thermal_throttle on AMD fires from sustained temperature alone, at or
above 83 °C for five minutes, together with the platform's own health
episodes. power_cap cannot be detected there at all, so power is
absent from checked. comm_errors is detectable: it fires from the
ECC counters and the PCIe counters, so comm stays in checked. GPU
memory temperature answers too. ecc_sbe there carries the correctable
and deferred totals together, while ecc_dbe carries the uncorrectable
total. The unavailable list on such a cluster reads
throttle_violations, clock_events and xid.
Coverage and clean¶
A metric family is in one of three states. The difference decides what an honest answer says:
| State | Meaning |
|---|---|
coverage.present |
queried, and the platform has the readings |
coverage.missing |
queried, and nothing came back |
unavailable |
never queried: this GPU maker structurally has no such reading |
coverage rides every action. unavailable rides action=flags,
beside checked. A family the maker structurally lacks appears in
neither coverage list, because the platform never asks for it. A
family whose inputs answered only in part is coverage.present.
checked names only what the cluster's GPU maker can answer, so an
unavailable family never holds back a clean result. clean is true
when no flag fired and every family in checked answered. A checked
family that did not answer keeps clean: false. note then names it.
A healthy cluster whose maker lacks several families reads
clean: true, with those families named in note. An agent reporting
a clean result should say which families it checked and which the
hardware cannot report at all.
The reference¶
reference.temp_p95_c is the 95th percentile of GPU temperature over
the same window for the cluster's GPU model, taken across your
organization's own nodes and the island's shared pool, aggregated to
one number. It answers "is this GPU hot for its model" without a second
call. It is not a reading across every tenant on the island. The
aggregate carries no names, no counts and no locations. It is null
where the platform could not compute it. A flag then reports its own
temperature with no comparison.
Refusals¶
| Code | Status | Meaning |
|---|---|---|
bad-request |
422 | a missing or malformed argument: no cluster, an action without its target, a node or workload name outside [A-Za-z0-9][A-Za-z0-9._-]{0,120}, a window longer than 14 days or ending at or before its start, or points above 96 |
not-found |
404 | a cluster, node or workload outside your organization, or one that does not exist |
metrics-unavailable |
503 | the platform's monitoring did not answer for this cluster; retry |
cluster is required, so a call without one is a 422, never a
409. A 409 ambiguous-cluster survives in one case: your
organization holds two clusters of the same name on two islands. The
refusal lists the candidates.
What the platform does not see¶
metrics_read reads the hardware and the host. It does not see your
training loop: step time, samples or tokens per second, loss,
dataloader waits, or anything else inside your process. A clean flags
answer means the platform's own signals are clean. It means nothing
more. An explanation of a slowdown pairs these signals with the job's
own step timings.
Evidence for a slow run¶
- Where it ran.
metrics_readwithaction=workload,cluster=aurora-prod,kind=Job,name=train-llm,range=6hnames every node the job touched and when. Checkplacement_source. - What fired. The same target with
action=flagsreturns the conditions, each with its node, GPU, sub-window and numbers. A flag whoseclippedreadstrueat either end ran past the window, so widen the window or quote itsepisodebounds. - The shape over time.
action=node,node=aurora-prod-n2,detail=seriesshows the flagged node's own curves: temperature climbing while clock and tensor core activity fall with it. - The sentence the evidence supports. "GPU 3 on aurora-prod-n2 held at or above 83 °C for 41 minutes from 14:05 UTC, peaking at 91 °C against a reference p95 of 74.5 °C for this GPU model. The job's other node stayed clean over the same window."
That evidence supports one claim: this GPU was throttled. It does not by itself explain the run's wall clock. Read the job's own step timings over the same window before attributing the slowdown to the hardware or to the code.
Free-form queries: metrics_query¶
metrics_query runs one PromQL expression of your own over the same
monitoring, scoped to your organization: every cluster you own, nothing
else. It is the open end next to the fixed reads above. Use it for fleet
views, custom aggregations, or a metric metrics_read does not summarize.
Like metrics_read it is read only and single phase.
| Argument | Values | Meaning |
|---|---|---|
query |
PromQL, at most 4000 characters | required |
instant |
true or false (default) |
true returns one value per series at the window's end; false returns a series over the window |
range, start, end, points |
as in The window | the window; points sets the step |
The answer carries org, query, result_type, window, series,
series_total and truncated. A series over the window is
{labels, values} on the inclusive [start, end] grid, one slot per
step, null where nothing answered; an instant answer is
{labels, value}; a scalar answer carries value alone. At most 100
series come back. When some were dropped, or when nothing matched, a
note says so: an empty result is a red flag to check the metric name
and matchers, not an all-clear.
Labels worth knowing: cluster is the cluster name the console shows,
instance is the node as the cluster names it, and the device is gpu
on NVIDIA nodes and gpu_id on AMD nodes. GPU series are DCGM_FI_DEV_*
on NVIDIA nodes and amd_gpu_* on AMD nodes. To discover names, ask for
count by (__name__)({__name__=~"amd_gpu_.*"}).
Refusals: a PromQL error, or a window or points outside the catalog,
is a 422 carrying the store's own words; a throttled organization is a
429 with Retry-After; an unavailable store is a 503.