# Telemetry and flags

`metrics_read` answers what the platform's own monitoring saw on the
hardware your job ran on: per GPU temperature, clocks, activity, power,
memory, throttle time and error counters, plus the host's CPU, pressure
and network. It also answers the platform's judgement on those readings
as **flags**, each one carrying the evidence behind it.

The tool lives on the [MCP server](https://docs.nationalcompute.com/api/mcp.md). It is read only and single
phase, so no `confirm_token` is involved. The
[cluster page reads](https://docs.nationalcompute.com/api/k8s-cluster-pages.md) serve the console's own
utilization and machine panels over REST; this tool is the agent-facing
read of the same monitoring, with the flags layer on top. For anything
these three actions do not answer, [`metrics_query`](#free-form-queries-metrics_query)
below runs one PromQL expression of your own over the same monitoring.

| `action` | Target | Answers |
|---|---|---|
| `node` | one node | every reading the platform holds for that node's GPUs and host |
| `workload` | one workload | where it ran over the window, plus each node's readings |
| `flags` | a node or a workload | the platform's judgement: which conditions fired, with evidence |

## Arguments

| Argument | Values | Applies to |
|---|---|---|
| `action` | `node`, `workload`, `flags` | required on every call |
| `cluster` | the cluster name (`grid_whoami` lists them) | required on every call |
| `node` | the node's name | `action=node`; `action=flags` on a node |
| `kind` | `Job`, `JobSet`, `MPIJob`, `Deployment`, `StatefulSet`, `Pod` | `action=workload`; `action=flags` on a workload |
| `name` | the workload's name | the same two |
| `namespace` | defaults to `default` | workload targets |
| `range` | `1h`, `6h`, `24h`, `7d`; defaults to `6h` | optional |
| `start`, `end` | ISO 8601 UTC; `end` defaults to now | optional |
| `detail` | `summary` or `series` | optional |
| `points` | series length after downsampling; defaults to 48, maximum 96 | optional |

`action=flags` takes either `node` or the pair (`kind`, `name`).
Anything else is a `422 bad-request`.

## The window

`range` names a window ending now. `start` and `end` set an explicit
window instead, and the call then ignores `range`. A span longer than
14 days is refused. A `start` at or after `end` is refused.

The bucket width follows from the span and the `points` you asked for:
the span divided by `points`, rounded up to a multiple of 30 seconds,
never below 30 seconds. A 6 hour window at the default 48 points gives
a `step_s` of 450. The answer's `window` echoes the resolved `start`,
`end`, `step_s` and the resulting `points`.

`detail` defaults to `series` for `action=node`. It defaults to
`summary` for `action=workload` and `action=flags`. Under `summary`
each series collapses to its `{mean, min, max}`, except `temp_c`,
`tensor_active_pct` and `power_w`, which keep their values.

## The envelope

Every action answers the same outer object:

| Field | What it carries |
|---|---|
| `cluster`, `gpu_vendor`, `gpu_model` | the cluster and the GPU class its nodes carry |
| `window` | `start`, `end`, `step_s` and `points`, as resolved |
| `reference` | the [comparison number](#the-reference) for the GPU model, or `null` |
| `coverage` | `present` and `missing`: which metric families answered |
| `note` | the units and how to read this answer, in one short paragraph |

A series is
`{"t0": <epoch seconds>, "step_s": n, "values": [...], "mean": x, "min": x, "max": x}`.
A gap inside `values` is `null`. A series the platform holds no reading
for is `null` whole. A `null` is never a zero. `coverage.missing` names
every family that did not answer.

The family names are a fixed vocabulary: `gpu_temp`, `memory_temp`,
`sm_clock`, `mem_clock`, `gpu_util`, `tensor_active`, `sm_active`,
`dram_active`, `power`, `hbm`, `throttle_violations`, `clock_events`,
`link_errors`, `ecc`, `xid`, `host_cpu`, `host_pressure`, `host_nic`,
`north_south`, `health_episodes`.

## action=node

```json
{
  "cluster": "aurora-prod", "gpu_vendor": "amd", "gpu_model": "MI355X",
  "window": {"start": "2026-09-21T12:00:00Z", "end": "2026-09-21T18:00:00Z",
             "step_s": 450, "points": 48},
  "reference": {"gpu_model": "MI355X", "temp_p95_c": 74.5,
                "note": "p95 of GPU temperature for this model over the window, across your nodes and the island's shared pool"},
  "coverage": {"present": ["gpu_temp", "memory_temp", "tensor_active",
                           "power", "hbm", "ecc", "link_errors",
                           "host_pressure", "…"],
               "missing": ["north_south"]},
  "node": {"name": "aurora-prod-n2", "gpus": 8, "state": "ready"},
  "gpus": {
    "3": {
      "temp_c": {"t0": 1758456000, "step_s": 450,
                 "values": [71.0, 83.5, 91.0, 88.2, null],
                 "mean": 83.4, "min": 68.0, "max": 91.0},
      "peak_temp_c": 91.0, "memory_temp_c": {"…": "…"},
      "tensor_active_pct": {"…": "…"}, "power_w": {"…": "…"},
      "hbm_used_gib": {"mean": 241.6, "min": 238.0, "max": 244.9, "…": "…"},
      "hbm_total_gib": 268.2,
      "throttle": {"thermal_s": null, "power_s": null, "clock_events": null,
                   "…": null},
      "errors": {"link_crc": null, "link_replay": null, "link_recovery": null,
                 "pcie_replay": 0, "pcie_recovery": 0, "ecc_sbe": 0,
                 "ecc_dbe": 0, "xid": null}
    }
  },
  "host": {"cpu_busy_pct": {"…": "…"}, "nic": {"errors": 0, "drops": 0, "…": "…"},
           "pressure": {"cpu": {"…": "…"}, "memory": null, "io": {"…": "…"}}},
  "health_episodes": [{"kind": "thermal", "started": "2026-09-21T14:05:00Z",
                       "ended": "2026-09-21T14:46:00Z", "detail": "…"}]
}
```

`gpus` is keyed by the GPU's index on the node:

| Field | Reading |
|---|---|
| `temp_c`, `memory_temp_c` | GPU temperature and memory temperature, °C |
| `sm_clock_mhz`, `mem_clock_mhz` | core clock and memory clock, MHz |
| `util_pct`, `sm_active_pct`, `dram_active_pct` | how busy the GPU, its cores and its memory were, percent |
| `tensor_active_pct` | tensor core activity, percent: the reading that tracks training work |
| `power_w` | power draw, watts |
| `peak_temp_c` | the highest temperature the GPU reached in the window, °C |
| `hbm_used_gib`, `hbm_total_gib` | GPU memory in use, and the GPU's total, GiB. The total is the denominator `memory_pressure` measures against |
| `throttle` | seconds of throttle by cause over the window: `thermal_s`, `power_s`, `board_limit_s`, `sync_boost_s`, `low_util_s`, `reliability_s`; `clock_events` counts each clock-event kind the GPU reports, for example `HW_SLOWDOWN` |
| `errors` | counter increases over the window: `link_crc`, `link_replay`, `link_recovery`, `pcie_replay`, `pcie_recovery`, `ecc_sbe`, `ecc_dbe`, `xid` |

`link_*` counts the GPU-to-GPU fabric links. Those are the NVLink
counters on NVIDIA. The AMD equivalents (XGMI) are not read today, so
`link_*` is `null` on an AMD cluster.

`pcie_recovery` counts PCIe link recoveries over the window. It is
`null` where a maker has no such counter, which is NVIDIA today.

Throttle seconds and error counts are the **increase over the window**,
never a lifetime total. A counter the platform does not have for a GPU
maker reads `null`, field by field. The cluster above is an AMD
cluster, which is why every throttle timer, every `link_*` counter and
`xid` read `null` on it. Its PCIe and ECC counters answer.

`host` carries the node itself:

| Field | Reading |
|---|---|
| `cpu_busy_pct` | host CPU busy, percent |
| `load_per_core` | run queue length per core |
| `pressure.cpu`, `pressure.memory`, `pressure.io` | share of time the host waited on CPU, on memory, or on disk, percent |
| `nic.rx_mbps`, `nic.tx_mbps`, `nic.errors`, `nic.drops` | node network throughput, and error and drop counter increases over the window |
| `north_south.ingress_mbps`, `north_south.egress_mbps` | traffic to and from the internet |

`health_episodes` lists the platform's own recorded episodes for the
node inside the window, as `{kind, started, ended, detail}`. An open
episode reads `"ended": null`.

## action=workload

```json
{
  "workload": {"kind": "Job", "namespace": "default", "name": "train-llm"},
  "placement": [
    {"pod": "train-llm-0", "node": "aurora-prod-n1",
     "from": "2026-09-21T12:04:00Z", "to": null, "phase": "Running"},
    {"pod": "train-llm-1", "node": "aurora-prod-n2",
     "from": "2026-09-21T12:04:00Z", "to": null, "phase": "Running"}],
  "placement_source": "history",
  "nodes": {
    "aurora-prod-n1": {"gpus": {"0": {"temp_c": {"mean": 68.1, "min": 61.0,
                                      "max": 72.4}, "…": "…"}}},
    "aurora-prod-n2": {"…": "…"}}
}
```

`placement` comes from the workload's own scheduling history,
intersected with the window. `placement_source` then reads `history`.
Where the platform holds no history for that window, placement comes
from the pods running now, and `placement_source` reads `current-pods`
instead. Read `placement_source` before you trust a placement.

`nodes` holds one node block per node the workload touched, at
`summary` detail unless you ask for `series`. Each node block carries
its own `health_episodes`, the same shape as on `action=node`.

## action=flags

A one hour window over the same node, opened after the episode above had
already started:

```json
{
  "window": {"start": "2026-09-21T14:20:00Z", "end": "2026-09-21T15:20:00Z",
             "step_s": 90, "points": 40},
  "target": {"workload": {"kind": "Job", "namespace": "default",
                          "name": "train-llm"}, "placement": ["…"]},
  "checked": ["thermal", "utilization", "comm", "host", "memory"],
  "unavailable": ["throttle_violations", "clock_events", "xid"],
  "clean": false,
  "flags": [
    {"kind": "thermal_throttle", "severity": "high",
     "node": "aurora-prod-n2", "gpu": "3", "minutes": 26,
     "from": "2026-09-21T14:20:00Z", "to": "2026-09-21T14:46:00Z",
     "clipped": {"start": true, "end": false},
     "episode": {"from": "2026-09-21T14:05:00Z",
                 "to": "2026-09-21T14:46:00Z"},
     "evidence": {"peak_temp_c": 91.0, "reference_p95_c": 74.5},
     "summary": "GPU 3 on aurora-prod-n2 held at or above 83 °C for 26 min inside this window (peak 91 °C, reference p95 74.5 °C); the episode began 14:05 UTC, before the window"}],
  "note": "this GPU maker reports no throttle timers, clock events or XID events, and no fabric link counters within link_errors; one flag is clipped at the window start"
}
```

Each flag object:

| Field | What it carries |
|---|---|
| `kind`, `severity` | the condition, and how serious this instance of it is |
| `node`, `gpu` | the node, and the index of the GPU the evidence came from |
| `from`, `to`, `minutes` | the sub-window the evidence covers, and its length |
| `clipped` | `{start, end}`: whether the evidence runs past that edge of the window |
| `episode` | the platform's own health episode bounds for this node and kind, or `null` |
| `evidence` | the numbers behind the flag, whichever of them the GPU maker reports: peaks, counter increases, clock events, the reference p95 |
| `summary` | one sentence an agent can quote |

`clipped.start` is `true` when the evidence was already abnormal in the
window's first bucket. `clipped.end` is `true` when it was still
abnormal in the last. Either way `from`, `to` and `minutes` cover only
the part inside the window, so the length is a floor. The answer's
`note` names every clipped flag.

`episode` carries the bounds of the platform's own
[health episode](#actionnode) for that node and kind, where one overlaps
the flag. `to` is `null` while the episode is still open. The field is
`null` where no episode overlaps. Quote the episode bounds when a narrow
window cuts a flag short.

The six kinds, and the rule that fires each one:

| `kind` | Fires when | `severity` |
|---|---|---|
| `thermal_throttle` | thermal throttle time increased, a thermal clock event landed, or the GPU held at or above 83 °C for at least 5 minutes | `high` at 10 minutes or longer, otherwise `warn` |
| `power_cap` | power throttle time increased, or a power clock event landed | `high` at 10 minutes or longer, otherwise `warn` |
| `utilization_collapse` | tensor core activity stayed below half the window's median for at least 10 minutes while the window's pods were Running | `warn` |
| `comm_errors` | a fabric link or PCIe error counter increased, an uncorrectable memory error landed, or a GPU fault event landed | `high` for uncorrectable memory errors and GPU fault events, otherwise `warn` |
| `host_pressure` | the host waited on CPU, memory or disk more than 20% of the time for at least 5 minutes | `warn` |
| `memory_pressure` | GPU memory in use held at or above 95% of the GPU's total for at least 5 minutes | `warn` |

`utilization_collapse` falls back to plain GPU utilization where the
platform has no tensor core reading for that GPU maker.

What a GPU maker cannot report changes what can fire. On AMD today the
platform reads no throttle timers, no clock events and no XID events.
Inside `link_errors` it reads the PCIe counters; the fabric link (XGMI)
counters are absent, so `link_*` is `null` and `link_errors` answers
only in part.

`thermal_throttle` on AMD fires from sustained temperature alone, at or
above 83 °C for five minutes, together with the platform's own health
episodes. `power_cap` cannot be detected there at all, so `power` is
absent from `checked`. `comm_errors` is detectable: it fires from the
ECC counters and the PCIe counters, so `comm` stays in `checked`. GPU
memory temperature answers too. `ecc_sbe` there carries the correctable
and deferred totals together, while `ecc_dbe` carries the uncorrectable
total. The `unavailable` list on such a cluster reads
`throttle_violations`, `clock_events` and `xid`.

## Coverage and clean

A metric family is in one of three states. The difference decides what
an honest answer says:

| State | Meaning |
|---|---|
| `coverage.present` | queried, and the platform has the readings |
| `coverage.missing` | queried, and nothing came back |
| `unavailable` | never queried: this GPU maker structurally has no such reading |

`coverage` rides every action. `unavailable` rides `action=flags`,
beside `checked`. A family the maker structurally lacks appears in
neither `coverage` list, because the platform never asks for it. A
family whose inputs answered only in part is `coverage.present`.

`checked` names only what the cluster's GPU maker can answer, so an
`unavailable` family never holds back a clean result. `clean` is `true`
when no flag fired and every family in `checked` answered. A checked
family that did not answer keeps `clean: false`. `note` then names it.
A healthy cluster whose maker lacks several families reads
`clean: true`, with those families named in `note`. An agent reporting
a clean result should say which families it checked and which the
hardware cannot report at all.

## The reference

`reference.temp_p95_c` is the 95th percentile of GPU temperature over
the same window for the cluster's GPU model, taken across your
organization's own nodes and the island's shared pool, aggregated to
one number. It answers "is this GPU hot for its model" without a second
call. It is not a reading across every tenant on the island. The
aggregate carries no names, no counts and no locations. It is `null`
where the platform could not compute it. A flag then reports its own
temperature with no comparison.

## Refusals

| Code | Status | Meaning |
|---|---|---|
| `bad-request` | 422 | a missing or malformed argument: no `cluster`, an action without its target, a node or workload name outside `[A-Za-z0-9][A-Za-z0-9._-]{0,120}`, a window longer than 14 days or ending at or before its start, or `points` above 96 |
| `not-found` | 404 | a cluster, node or workload outside your organization, or one that does not exist |
| `metrics-unavailable` | 503 | the platform's monitoring did not answer for this cluster; retry |

`cluster` is required, so a call without one is a `422`, never a
`409`. A `409 ambiguous-cluster` survives in one case: your
organization holds two clusters of the same name on two islands. The
refusal lists the candidates.

## What the platform does not see

`metrics_read` reads the hardware and the host. It does not see your
training loop: step time, samples or tokens per second, loss,
dataloader waits, or anything else inside your process. A clean `flags`
answer means the platform's own signals are clean. It means nothing
more. An explanation of a slowdown pairs these signals with the job's
own step timings.

## Evidence for a slow run

1. **Where it ran.** `metrics_read` with `action=workload`,
   `cluster=aurora-prod`, `kind=Job`, `name=train-llm`, `range=6h`
   names every node the job touched and when. Check
   `placement_source`.
2. **What fired.** The same target with `action=flags` returns the
   conditions, each with its node, GPU, sub-window and numbers. A flag
   whose `clipped` reads `true` at either end ran past the window, so
   widen the window or quote its `episode` bounds.
3. **The shape over time.** `action=node`, `node=aurora-prod-n2`,
   `detail=series` shows the flagged node's own curves: temperature
   climbing while clock and tensor core activity fall with it.
4. **The sentence the evidence supports.** "GPU 3 on aurora-prod-n2
   held at or above 83 °C for 41 minutes from 14:05 UTC, peaking at
   91 °C against a reference p95 of 74.5 °C for this GPU model. The
   job's other node stayed clean over the same window."

That evidence supports one claim: this GPU was throttled. It does not
by itself explain the run's wall clock. Read the job's own step timings
over the same window before attributing the slowdown to the hardware or
to the code.

## Free-form queries: `metrics_query`

`metrics_query` runs one PromQL expression of your own over the same
monitoring, scoped to your organization: every cluster you own, nothing
else. It is the open end next to the fixed reads above. Use it for fleet
views, custom aggregations, or a metric `metrics_read` does not summarize.
Like `metrics_read` it is read only and single phase.

| Argument | Values | Meaning |
|---|---|---|
| `query` | PromQL, at most 4000 characters | required |
| `instant` | `true` or `false` (default) | `true` returns one value per series at the window's end; `false` returns a series over the window |
| `range`, `start`, `end`, `points` | as in [The window](#the-window) | the window; `points` sets the step |

The answer carries `org`, `query`, `result_type`, `window`, `series`,
`series_total` and `truncated`. A series over the window is
`{labels, values}` on the inclusive `[start, end]` grid, one slot per
step, `null` where nothing answered; an instant answer is
`{labels, value}`; a scalar answer carries `value` alone. At most 100
series come back. When some were dropped, or when nothing matched, a
`note` says so: an empty result is a red flag to check the metric name
and matchers, not an all-clear.

Labels worth knowing: `cluster` is the cluster name the console shows,
`instance` is the node as the cluster names it, and the device is `gpu`
on NVIDIA nodes and `gpu_id` on AMD nodes. GPU series are `DCGM_FI_DEV_*`
on NVIDIA nodes and `amd_gpu_*` on AMD nodes. To discover names, ask for
`count by (__name__)({__name__=~"amd_gpu_.*"})`.

Refusals: a PromQL error, or a window or `points` outside the catalog,
is a 422 carrying the store's own words; a throttled organization is a
429 with `Retry-After`; an unavailable store is a 503.
