Skip to content

Telemetry and flags

metrics_read answers what the platform's own monitoring saw on the hardware your job ran on: per GPU temperature, clocks, activity, power, memory, throttle time and error counters, plus the host's CPU, pressure and network. It also answers the platform's judgement on those readings as flags, each one carrying the evidence behind it.

The tool lives on the MCP server. It is read only and single phase, so no confirm_token is involved. The cluster page reads serve the console's own utilization and machine panels over REST; this tool is the agent-facing read of the same monitoring, with the flags layer on top. For anything these three actions do not answer, metrics_query below runs one PromQL expression of your own over the same monitoring.

action Target Answers
node one node every reading the platform holds for that node's GPUs and host
workload one workload where it ran over the window, plus each node's readings
flags a node or a workload the platform's judgement: which conditions fired, with evidence

Arguments

Argument Values Applies to
action node, workload, flags required on every call
cluster the cluster name (grid_whoami lists them) required on every call
node the node's name action=node; action=flags on a node
kind Job, JobSet, MPIJob, Deployment, StatefulSet, Pod action=workload; action=flags on a workload
name the workload's name the same two
namespace defaults to default workload targets
range 1h, 6h, 24h, 7d; defaults to 6h optional
start, end ISO 8601 UTC; end defaults to now optional
detail summary or series optional
points series length after downsampling; defaults to 48, maximum 96 optional

action=flags takes either node or the pair (kind, name). Anything else is a 422 bad-request.

The window

range names a window ending now. start and end set an explicit window instead, and the call then ignores range. A span longer than 14 days is refused. A start at or after end is refused.

The bucket width follows from the span and the points you asked for: the span divided by points, rounded up to a multiple of 30 seconds, never below 30 seconds. A 6 hour window at the default 48 points gives a step_s of 450. The answer's window echoes the resolved start, end, step_s and the resulting points.

detail defaults to series for action=node. It defaults to summary for action=workload and action=flags. Under summary each series collapses to its {mean, min, max}, except temp_c, tensor_active_pct and power_w, which keep their values.

The envelope

Every action answers the same outer object:

Field What it carries
cluster, gpu_vendor, gpu_model the cluster and the GPU class its nodes carry
window start, end, step_s and points, as resolved
reference the comparison number for the GPU model, or null
coverage present and missing: which metric families answered
note the units and how to read this answer, in one short paragraph

A series is {"t0": <epoch seconds>, "step_s": n, "values": [...], "mean": x, "min": x, "max": x}. A gap inside values is null. A series the platform holds no reading for is null whole. A null is never a zero. coverage.missing names every family that did not answer.

The family names are a fixed vocabulary: gpu_temp, memory_temp, sm_clock, mem_clock, gpu_util, tensor_active, sm_active, dram_active, power, hbm, throttle_violations, clock_events, link_errors, ecc, xid, host_cpu, host_pressure, host_nic, north_south, health_episodes.

action=node

{
  "cluster": "aurora-prod", "gpu_vendor": "amd", "gpu_model": "MI355X",
  "window": {"start": "2026-09-21T12:00:00Z", "end": "2026-09-21T18:00:00Z",
             "step_s": 450, "points": 48},
  "reference": {"gpu_model": "MI355X", "temp_p95_c": 74.5,
                "note": "p95 of GPU temperature for this model over the window, across your nodes and the island's shared pool"},
  "coverage": {"present": ["gpu_temp", "memory_temp", "tensor_active",
                           "power", "hbm", "ecc", "link_errors",
                           "host_pressure", "…"],
               "missing": ["north_south"]},
  "node": {"name": "aurora-prod-n2", "gpus": 8, "state": "ready"},
  "gpus": {
    "3": {
      "temp_c": {"t0": 1758456000, "step_s": 450,
                 "values": [71.0, 83.5, 91.0, 88.2, null],
                 "mean": 83.4, "min": 68.0, "max": 91.0},
      "peak_temp_c": 91.0, "memory_temp_c": {"…": "…"},
      "tensor_active_pct": {"…": "…"}, "power_w": {"…": "…"},
      "hbm_used_gib": {"mean": 241.6, "min": 238.0, "max": 244.9, "…": "…"},
      "hbm_total_gib": 268.2,
      "throttle": {"thermal_s": null, "power_s": null, "clock_events": null,
                   "…": null},
      "errors": {"link_crc": null, "link_replay": null, "link_recovery": null,
                 "pcie_replay": 0, "pcie_recovery": 0, "ecc_sbe": 0,
                 "ecc_dbe": 0, "xid": null}
    }
  },
  "host": {"cpu_busy_pct": {"…": "…"}, "nic": {"errors": 0, "drops": 0, "…": "…"},
           "pressure": {"cpu": {"…": "…"}, "memory": null, "io": {"…": "…"}}},
  "health_episodes": [{"kind": "thermal", "started": "2026-09-21T14:05:00Z",
                       "ended": "2026-09-21T14:46:00Z", "detail": "…"}]
}

gpus is keyed by the GPU's index on the node:

Field Reading
temp_c, memory_temp_c GPU temperature and memory temperature, °C
sm_clock_mhz, mem_clock_mhz core clock and memory clock, MHz
util_pct, sm_active_pct, dram_active_pct how busy the GPU, its cores and its memory were, percent
tensor_active_pct tensor core activity, percent: the reading that tracks training work
power_w power draw, watts
peak_temp_c the highest temperature the GPU reached in the window, °C
hbm_used_gib, hbm_total_gib GPU memory in use, and the GPU's total, GiB. The total is the denominator memory_pressure measures against
throttle seconds of throttle by cause over the window: thermal_s, power_s, board_limit_s, sync_boost_s, low_util_s, reliability_s; clock_events counts each clock-event kind the GPU reports, for example HW_SLOWDOWN
errors counter increases over the window: link_crc, link_replay, link_recovery, pcie_replay, pcie_recovery, ecc_sbe, ecc_dbe, xid

link_* counts the GPU-to-GPU fabric links. Those are the NVLink counters on NVIDIA. The AMD equivalents (XGMI) are not read today, so link_* is null on an AMD cluster.

pcie_recovery counts PCIe link recoveries over the window. It is null where a maker has no such counter, which is NVIDIA today.

Throttle seconds and error counts are the increase over the window, never a lifetime total. A counter the platform does not have for a GPU maker reads null, field by field. The cluster above is an AMD cluster, which is why every throttle timer, every link_* counter and xid read null on it. Its PCIe and ECC counters answer.

host carries the node itself:

Field Reading
cpu_busy_pct host CPU busy, percent
load_per_core run queue length per core
pressure.cpu, pressure.memory, pressure.io share of time the host waited on CPU, on memory, or on disk, percent
nic.rx_mbps, nic.tx_mbps, nic.errors, nic.drops node network throughput, and error and drop counter increases over the window
north_south.ingress_mbps, north_south.egress_mbps traffic to and from the internet

health_episodes lists the platform's own recorded episodes for the node inside the window, as {kind, started, ended, detail}. An open episode reads "ended": null.

action=workload

{
  "workload": {"kind": "Job", "namespace": "default", "name": "train-llm"},
  "placement": [
    {"pod": "train-llm-0", "node": "aurora-prod-n1",
     "from": "2026-09-21T12:04:00Z", "to": null, "phase": "Running"},
    {"pod": "train-llm-1", "node": "aurora-prod-n2",
     "from": "2026-09-21T12:04:00Z", "to": null, "phase": "Running"}],
  "placement_source": "history",
  "nodes": {
    "aurora-prod-n1": {"gpus": {"0": {"temp_c": {"mean": 68.1, "min": 61.0,
                                      "max": 72.4}, "…": "…"}}},
    "aurora-prod-n2": {"…": "…"}}
}

placement comes from the workload's own scheduling history, intersected with the window. placement_source then reads history. Where the platform holds no history for that window, placement comes from the pods running now, and placement_source reads current-pods instead. Read placement_source before you trust a placement.

nodes holds one node block per node the workload touched, at summary detail unless you ask for series. Each node block carries its own health_episodes, the same shape as on action=node.

action=flags

A one hour window over the same node, opened after the episode above had already started:

{
  "window": {"start": "2026-09-21T14:20:00Z", "end": "2026-09-21T15:20:00Z",
             "step_s": 90, "points": 40},
  "target": {"workload": {"kind": "Job", "namespace": "default",
                          "name": "train-llm"}, "placement": ["…"]},
  "checked": ["thermal", "utilization", "comm", "host", "memory"],
  "unavailable": ["throttle_violations", "clock_events", "xid"],
  "clean": false,
  "flags": [
    {"kind": "thermal_throttle", "severity": "high",
     "node": "aurora-prod-n2", "gpu": "3", "minutes": 26,
     "from": "2026-09-21T14:20:00Z", "to": "2026-09-21T14:46:00Z",
     "clipped": {"start": true, "end": false},
     "episode": {"from": "2026-09-21T14:05:00Z",
                 "to": "2026-09-21T14:46:00Z"},
     "evidence": {"peak_temp_c": 91.0, "reference_p95_c": 74.5},
     "summary": "GPU 3 on aurora-prod-n2 held at or above 83 °C for 26 min inside this window (peak 91 °C, reference p95 74.5 °C); the episode began 14:05 UTC, before the window"}],
  "note": "this GPU maker reports no throttle timers, clock events or XID events, and no fabric link counters within link_errors; one flag is clipped at the window start"
}

Each flag object:

Field What it carries
kind, severity the condition, and how serious this instance of it is
node, gpu the node, and the index of the GPU the evidence came from
from, to, minutes the sub-window the evidence covers, and its length
clipped {start, end}: whether the evidence runs past that edge of the window
episode the platform's own health episode bounds for this node and kind, or null
evidence the numbers behind the flag, whichever of them the GPU maker reports: peaks, counter increases, clock events, the reference p95
summary one sentence an agent can quote

clipped.start is true when the evidence was already abnormal in the window's first bucket. clipped.end is true when it was still abnormal in the last. Either way from, to and minutes cover only the part inside the window, so the length is a floor. The answer's note names every clipped flag.

episode carries the bounds of the platform's own health episode for that node and kind, where one overlaps the flag. to is null while the episode is still open. The field is null where no episode overlaps. Quote the episode bounds when a narrow window cuts a flag short.

The six kinds, and the rule that fires each one:

kind Fires when severity
thermal_throttle thermal throttle time increased, a thermal clock event landed, or the GPU held at or above 83 °C for at least 5 minutes high at 10 minutes or longer, otherwise warn
power_cap power throttle time increased, or a power clock event landed high at 10 minutes or longer, otherwise warn
utilization_collapse tensor core activity stayed below half the window's median for at least 10 minutes while the window's pods were Running warn
comm_errors a fabric link or PCIe error counter increased, an uncorrectable memory error landed, or a GPU fault event landed high for uncorrectable memory errors and GPU fault events, otherwise warn
host_pressure the host waited on CPU, memory or disk more than 20% of the time for at least 5 minutes warn
memory_pressure GPU memory in use held at or above 95% of the GPU's total for at least 5 minutes warn

utilization_collapse falls back to plain GPU utilization where the platform has no tensor core reading for that GPU maker.

What a GPU maker cannot report changes what can fire. On AMD today the platform reads no throttle timers, no clock events and no XID events. Inside link_errors it reads the PCIe counters; the fabric link (XGMI) counters are absent, so link_* is null and link_errors answers only in part.

thermal_throttle on AMD fires from sustained temperature alone, at or above 83 °C for five minutes, together with the platform's own health episodes. power_cap cannot be detected there at all, so power is absent from checked. comm_errors is detectable: it fires from the ECC counters and the PCIe counters, so comm stays in checked. GPU memory temperature answers too. ecc_sbe there carries the correctable and deferred totals together, while ecc_dbe carries the uncorrectable total. The unavailable list on such a cluster reads throttle_violations, clock_events and xid.

Coverage and clean

A metric family is in one of three states. The difference decides what an honest answer says:

State Meaning
coverage.present queried, and the platform has the readings
coverage.missing queried, and nothing came back
unavailable never queried: this GPU maker structurally has no such reading

coverage rides every action. unavailable rides action=flags, beside checked. A family the maker structurally lacks appears in neither coverage list, because the platform never asks for it. A family whose inputs answered only in part is coverage.present.

checked names only what the cluster's GPU maker can answer, so an unavailable family never holds back a clean result. clean is true when no flag fired and every family in checked answered. A checked family that did not answer keeps clean: false. note then names it. A healthy cluster whose maker lacks several families reads clean: true, with those families named in note. An agent reporting a clean result should say which families it checked and which the hardware cannot report at all.

The reference

reference.temp_p95_c is the 95th percentile of GPU temperature over the same window for the cluster's GPU model, taken across your organization's own nodes and the island's shared pool, aggregated to one number. It answers "is this GPU hot for its model" without a second call. It is not a reading across every tenant on the island. The aggregate carries no names, no counts and no locations. It is null where the platform could not compute it. A flag then reports its own temperature with no comparison.

Refusals

Code Status Meaning
bad-request 422 a missing or malformed argument: no cluster, an action without its target, a node or workload name outside [A-Za-z0-9][A-Za-z0-9._-]{0,120}, a window longer than 14 days or ending at or before its start, or points above 96
not-found 404 a cluster, node or workload outside your organization, or one that does not exist
metrics-unavailable 503 the platform's monitoring did not answer for this cluster; retry

cluster is required, so a call without one is a 422, never a 409. A 409 ambiguous-cluster survives in one case: your organization holds two clusters of the same name on two islands. The refusal lists the candidates.

What the platform does not see

metrics_read reads the hardware and the host. It does not see your training loop: step time, samples or tokens per second, loss, dataloader waits, or anything else inside your process. A clean flags answer means the platform's own signals are clean. It means nothing more. An explanation of a slowdown pairs these signals with the job's own step timings.

Evidence for a slow run

  1. Where it ran. metrics_read with action=workload, cluster=aurora-prod, kind=Job, name=train-llm, range=6h names every node the job touched and when. Check placement_source.
  2. What fired. The same target with action=flags returns the conditions, each with its node, GPU, sub-window and numbers. A flag whose clipped reads true at either end ran past the window, so widen the window or quote its episode bounds.
  3. The shape over time. action=node, node=aurora-prod-n2, detail=series shows the flagged node's own curves: temperature climbing while clock and tensor core activity fall with it.
  4. The sentence the evidence supports. "GPU 3 on aurora-prod-n2 held at or above 83 °C for 41 minutes from 14:05 UTC, peaking at 91 °C against a reference p95 of 74.5 °C for this GPU model. The job's other node stayed clean over the same window."

That evidence supports one claim: this GPU was throttled. It does not by itself explain the run's wall clock. Read the job's own step timings over the same window before attributing the slowdown to the hardware or to the code.

Free-form queries: metrics_query

metrics_query runs one PromQL expression of your own over the same monitoring, scoped to your organization: every cluster you own, nothing else. It is the open end next to the fixed reads above. Use it for fleet views, custom aggregations, or a metric metrics_read does not summarize. Like metrics_read it is read only and single phase.

Argument Values Meaning
query PromQL, at most 4000 characters required
instant true or false (default) true returns one value per series at the window's end; false returns a series over the window
range, start, end, points as in The window the window; points sets the step

The answer carries org, query, result_type, window, series, series_total and truncated. A series over the window is {labels, values} on the inclusive [start, end] grid, one slot per step, null where nothing answered; an instant answer is {labels, value}; a scalar answer carries value alone. At most 100 series come back. When some were dropped, or when nothing matched, a note says so: an empty result is a red flag to check the metric name and matchers, not an all-clear.

Labels worth knowing: cluster is the cluster name the console shows, instance is the node as the cluster names it, and the device is gpu on NVIDIA nodes and gpu_id on AMD nodes. GPU series are DCGM_FI_DEV_* on NVIDIA nodes and amd_gpu_* on AMD nodes. To discover names, ask for count by (__name__)({__name__=~"amd_gpu_.*"}).

Refusals: a PromQL error, or a window or points outside the catalog, is a 422 carrying the store's own words; a throttled organization is a 429 with Retry-After; an unavailable store is a 503.