Marshall data sharing¶
Marshall data sharing is an organization-level setting with three tiers. Since 1 October 2026 tier 1 is the baseline for organizations that have never set a tier, and tier 0 is one click away. The tier decides what National Compute collects from your organization's signed-in Marshall sessions to improve Marshall and to train the model that powers it.
Public Marshall chat without sign-in does not enter this program and does not store conversation text.
Tiers¶
Since 1 October 2026 the baseline for an organization that has never set a tier is tier 1: redacted conversation transcripts are collected. An organization that set tier 0 itself stays at tier 0, and any member can lower the tier at any time from Settings → Data sharing.
| Tier | Collects | Never collects | Default |
|---|---|---|---|
| 0 — Usage analytics only | Counts, tool names, timings and outcomes of Marshall turns; error classes; model usage and cost; feedback notes members write to us | Message content; tool arguments or output; files in your storage | No |
| 1 — Analytics and conversation transcripts | Everything in tier 0; conversation transcripts as the model saw them, redacted for credentials before storage | Files in your storage | Yes, for organizations that have never set a tier |
| 2 — Analytics, transcripts and training artifacts | Everything in tier 1; training artifacts your runs write, matching the patterns below, copied from the organization share, your workspaces and the cluster's shared volume | Files whose name or content looks like a credential; files still being written |
Feedback notes — a thumbs-down note, a /feedback message, a task label —
are messages to National Compute and are collected at every tier.
Training artifacts¶
At tier 2 the platform copies files that match these patterns. Files are copied, never moved or modified. A file whose name or content looks like a credential is skipped. A file written in the last fifteen minutes waits for the next pass.
| Kind | Patterns | Minimum size |
|---|---|---|
| Checkpoints | *.pt, *.pth, *.safetensors, *.ckpt, *.gguf, *.msgpack, *.pkl, *optim*.pt; *.bin and *.index.json beside a checkpoint |
1 MiB |
| Logs | *.log, events.out.tfevents.*, trainer_state.json, *.out, *.err |
1 KiB |
| Configs | *.yaml, *.yml, *.toml, *.cfg, *.ini, *.args; *.json beside a checkpoint |
64 bytes |
| Datasets | *.jsonl, *.parquet, *.arrow, *.npy, *.npz, *.tsv, *.csv |
1 MiB |
Skipped by name: keys, certificates, .env files, kubeconfigs, SSH keys,
and anything named like a token, secret, credential or password. Skipped
by directory: .git, node_modules, __pycache__, virtual environments,
caches and the tool configuration directories under a home. Model files
downloaded from a public model hub are not collected: the hub's cache
directories are skipped wherever they sit, and a file whose download
record sits beside it is skipped as well. A checkpoint your run saves
into the same directory still counts.
Locations at tier 2: the organization share, each member's workspace, and the cluster's shared volume (the volume every worker mounts). Files that exceed 50 GiB are not copied.
Changing the tier¶
Console → Settings → Data sharing. Any member of the organization can change it; the page shows who set the current tier and when. A change takes effect for every capture after it. An organization may also be placed at a tier under a written agreement with National Compute, such as a research grant; the page says so, and changes to that tier go through National Compute under the agreement.
Revocation and retention¶
Lowering the tier stops future collection. Data collected earlier is kept under the tier in force when it was collected. Deletion requests are handled case by case where law requires.
Incentive¶
Organizations at tier 1 or 2 may receive credits on terms National Compute publishes in the console. The amount and form may change. Abuse — padding sessions, synthetic activity, uploading data to earn credits — makes an organization ineligible, the determination entirely at National Compute's discretion.
Redaction¶
Transcripts and feedback notes pass through automated, best-effort redaction before storage. It targets private keys, bearer tokens, JWTs, credentials in URLs, secret-named fields in JSON and YAML, and Kubernetes Secret objects. It is best effort: do not paste credentials into Marshall.
The program is governed by the Terms of Service, the Privacy Notice and the Data Processing Addendum.