# Kubernetes versions and maintenance

Every Kubernetes cluster runs RKE2 at a release pinned per cluster. The
platform team upgrades a cluster in two parts, the control plane first
and the workers second, inside a notice window you can defer.

```text
notice ────────▶ email to your admin contacts and a console banner;
    │            minor 14 days, patch 7 days, critical CVE 48 hours
    ▼
control plane ─▶ an etcd snapshot, then one replica at a time
    │            with a health check between replicas
    ▼
workers ───────▶ paused before the first node until you approve;
    │            the approval is recorded on the upgrade operation
    ▼
each node ─────▶ the agent restarts and one device plugin pod restarts;
                 running pods keep running
```

## The release pin

Each cluster is pinned to one RKE2 release. The pin moves only when the
platform team upgrades that cluster. Each cluster moves on its own
schedule. The platform's current release is `v1.36.2+rke2r1`
(currently; this page changes when the current release moves).
`kubectl --context <cluster> get nodes` shows the release each node runs
in its `VERSION` column.

## Upgrades

The platform team runs every upgrade per cluster. Nothing upgrades on
its own. There is no self service upgrade. An upgrade has
two parts.

### The control plane upgrade

The control plane upgrade passes three gates before anything restarts.
The target release stays in the cluster's current minor unless you
agreed to a minor upgrade. The target is never below the release the
control plane runs. The target is never below the release any worker
runs. After the gates pass, the platform takes an etcd snapshot
([etcd backups](#etcd-backups)). It then restarts one control plane
replica at a time with a health check between replicas. A control
plane replica restarts only inside an announced upgrade.

### The worker upgrade

The worker upgrade rolls the RKE2 agent across the cluster's nodes. A
skew guard refuses a worker release above the control plane's release.
On a cluster your organization holds, the roll pauses before its
first node. It runs only after your approval. The notice window sits in this
pause. Your approval is recorded on the upgrade operation. On each node
the agent restarts and one device plugin pod restarts. Running pods
keep running.

### Notice and deferral

A notice goes out before every upgrade, by email to your organization's
admin contacts and as a banner in the console. The notice length and
the deferral limit depend on the kind of upgrade. The values below are
the defaults. Your contract can set others.

| Upgrade | Notice | Deferral limit |
|---|---|---|
| minor | 14 days | 30 days; never more than one minor behind the platform's current release |
| patch | 7 days | 14 days |
| critical CVE | 48 hours | 7 days |

A deferral moves the window. It cannot move the window past the limit.
Reply to the notice to approve the window or to defer it. The
[notice templates](#upgrade-notices) show the fields every notice
carries.

## The arrival taint

Ask support for an arrival taint on a cluster, written as
`key=value:Effect`. Every node that joins the cluster carries it.
Platform daemonsets tolerate it. The platform never removes it. You
remove it once your own validation of the node is done:

```sh
kubectl --context <cluster> taint nodes <node> <key>-
```

With effect `NoSchedule`, a pod that does not tolerate the taint does
not land on the node until you remove it. A node carrying the arrival
taint is healthy to the platform. It bills while it carries the taint
([billing](https://docs.nationalcompute.com/billing.md)). No tool on the [MCP server](https://docs.nationalcompute.com/api/mcp.md) taints
or untaints a node. The removal is a `kubectl` action under your
kubeconfig.

## The registry mirror

A cluster can run an embedded registry mirror. Ask support to enable
it. With the mirror on, the cluster's nodes share pulled image layers
peer to peer. One upstream pull serves the whole cluster. A cold
cluster pulls its first copy upstream. A `latest` tag always comes from
upstream. The mirror is read only. Nothing can be pushed to it. The
platform runs no private registry and no pull through cache. Images
come from the registries your manifests name.

## etcd backups

Two schedules snapshot the cluster's etcd. The cadences and retention
below are the current values.

| Snapshot | Cadence | Kept |
|---|---|---|
| RKE2 local | every 12 hours | the last 5, on the control plane disk |
| platform | every 6 hours | up to 30 days, in cloud storage outside the cluster |

A control plane upgrade adds one more snapshot before its first replica
restarts. An etcd snapshot holds the cluster's Kubernetes objects. It
holds no bytes of the [shared volume](https://docs.nationalcompute.com/kubernetes.md#the-shared-volume)
or of node disks.

## Upgrade notices

Every upgrade notice carries the fields below. The planned form covers
a minor or a patch upgrade. The security form covers a critical CVE.
Reply to the notice to approve the window or to defer it.

A planned minor or patch upgrade:

```text
Subject: Kubernetes upgrade for <cluster>: <current release> to <target release>, <window start date>

Cluster:          <cluster>
Current release:  <current release>
Target release:   <target release>
Upgrade:          <minor | patch>
Window:           <start> to <end> <timezone>
Notice:           <14 days for a minor | 7 days for a patch> before the window
Deferral:         reply before <window start> to move the window;
                  up to <30 days for a minor | 14 days for a patch>;
                  a minor is never deferred more than one minor behind
                  the platform's current release
Approval:         reply to approve; the worker roll starts only after it
Mechanics:        https://docs.nationalcompute.com/kubernetes-versions-and-maintenance/#upgrades
Contact:          support@nationalcompute.com
```

A critical CVE upgrade:

```text
Subject: Security upgrade for <cluster>: <current release> to <target release> within 48 hours

Cluster:          <cluster>
Current release:  <current release>
Target release:   <target release>
Upgrade:          critical CVE <identifier>
Window:           <start> to <end> <timezone>, 48 hours from this notice
Deferral:         reply before <window start> to move the window; up to 7 days
Approval:         reply to approve; the worker roll starts only after it
Mechanics:        https://docs.nationalcompute.com/kubernetes-versions-and-maintenance/#upgrades
Contact:          support@nationalcompute.com
```
