# Marshall data sharing

Marshall data sharing is an organization-level setting with three tiers.
Since 1 October 2026 tier 1 is the baseline for organizations that have never set a tier, and tier 0 is one click away. The tier decides what National Compute collects from your
organization's signed-in Marshall sessions to improve Marshall and to train the model
that powers it.

[Public Marshall](https://docs.nationalcompute.com/public-marshall.md#data-retained) chat without sign-in
does not enter this program and does not store conversation text.

## Tiers

Since 1 October 2026 the baseline for an organization that has never set a
tier is **tier 1**: redacted conversation transcripts are collected. An
organization that set tier 0 itself stays at tier 0, and any member can
lower the tier at any time from Settings → Data sharing.

| Tier | Collects | Never collects | Default |
|---|---|---|---|
| 0 — Usage analytics only | Counts, tool names, timings and outcomes of Marshall turns; error classes; model usage and cost; feedback notes members write to us | Message content; tool arguments or output; files in your storage | No |
| 1 — Analytics and conversation transcripts | Everything in tier 0; conversation transcripts as the model saw them, redacted for credentials before storage | Files in your storage | Yes, for organizations that have never set a tier |
| 2 — Analytics, transcripts and training artifacts | Everything in tier 1; training artifacts your runs write, matching the patterns below, copied from the organization share, your workspaces and the cluster's shared volume | Files whose name or content looks like a credential; files still being written | |

Feedback notes — a thumbs-down note, a `/feedback` message, a task label —
are messages to National Compute and are collected at every tier.

## Training artifacts

At tier 2 the platform copies files that match these patterns. Files are
copied, never moved or modified. A file whose name or content looks like a
credential is skipped. A file written in the last fifteen minutes waits for
the next pass.

| Kind | Patterns | Minimum size |
|---|---|---|
| Checkpoints | `*.pt`, `*.pth`, `*.safetensors`, `*.ckpt`, `*.gguf`, `*.msgpack`, `*.pkl`, `*optim*.pt`; `*.bin` and `*.index.json` beside a checkpoint | 1 MiB |
| Logs | `*.log`, `events.out.tfevents.*`, `trainer_state.json`, `*.out`, `*.err` | 1 KiB |
| Configs | `*.yaml`, `*.yml`, `*.toml`, `*.cfg`, `*.ini`, `*.args`; `*.json` beside a checkpoint | 64 bytes |
| Datasets | `*.jsonl`, `*.parquet`, `*.arrow`, `*.npy`, `*.npz`, `*.tsv`, `*.csv` | 1 MiB |

Skipped by name: keys, certificates, `.env` files, kubeconfigs, SSH keys,
and anything named like a token, secret, credential or password. Skipped
by directory: `.git`, `node_modules`, `__pycache__`, virtual environments,
caches and the tool configuration directories under a home. Model files
downloaded from a public model hub are not collected: the hub's cache
directories are skipped wherever they sit, and a file whose download
record sits beside it is skipped as well. A checkpoint your run saves
into the same directory still counts.

Locations at tier 2: the organization share, each member's workspace, and
the cluster's shared volume (the volume every worker mounts). Files that
exceed 50 GiB are not copied.

## Changing the tier

Console → Settings → Data sharing. Any member of the organization can
change it; the page shows who set the current tier and when. A change takes
effect for every capture after it. An organization may also be placed at a
tier under a written agreement with National Compute, such as a research
grant; the page says so, and changes to that tier go through National
Compute under the agreement.

## Revocation and retention

Lowering the tier stops future collection. Data collected earlier is kept
under the tier in force when it was collected. Deletion requests are handled
case by case where law requires.

## Incentive

Organizations at tier 1 or 2 may receive credits on terms National Compute
publishes in the console. The amount and form may change. Abuse — padding
sessions, synthetic activity, uploading data to earn credits — makes an
organization ineligible, the determination entirely at National Compute's
discretion.

## Redaction

Transcripts and feedback notes pass through automated, best-effort
redaction before storage. It targets private keys, bearer tokens, JWTs,
credentials in URLs, secret-named fields in JSON and YAML, and Kubernetes
Secret objects. It is best effort: do not paste credentials into Marshall.

The program is governed by the
[Terms of Service](https://nationalcompute.com/documents/TOS), the
[Privacy Notice](https://nationalcompute.com/documents/PrivacyNotice) and the
[Data Processing Addendum](https://nationalcompute.com/documents/DPA).
