Skip to content

Marshall data sharing

Marshall data sharing is an organization-level setting with three tiers. Since 1 October 2026 tier 1 is the baseline for organizations that have never set a tier, and tier 0 is one click away. The tier decides what National Compute collects from your organization's signed-in Marshall sessions to improve Marshall and to train the model that powers it.

Public Marshall chat without sign-in does not enter this program and does not store conversation text.

Tiers

Since 1 October 2026 the baseline for an organization that has never set a tier is tier 1: redacted conversation transcripts are collected. An organization that set tier 0 itself stays at tier 0, and any member can lower the tier at any time from Settings → Data sharing.

Tier Collects Never collects Default
0 — Usage analytics only Counts, tool names, timings and outcomes of Marshall turns; error classes; model usage and cost; feedback notes members write to us Message content; tool arguments or output; files in your storage No
1 — Analytics and conversation transcripts Everything in tier 0; conversation transcripts as the model saw them, redacted for credentials before storage Files in your storage Yes, for organizations that have never set a tier
2 — Analytics, transcripts and training artifacts Everything in tier 1; training artifacts your runs write, matching the patterns below, copied from the organization share, your workspaces and the cluster's shared volume Files whose name or content looks like a credential; files still being written

Feedback notes — a thumbs-down note, a /feedback message, a task label — are messages to National Compute and are collected at every tier.

Training artifacts

At tier 2 the platform copies files that match these patterns. Files are copied, never moved or modified. A file whose name or content looks like a credential is skipped. A file written in the last fifteen minutes waits for the next pass.

Kind Patterns Minimum size
Checkpoints *.pt, *.pth, *.safetensors, *.ckpt, *.gguf, *.msgpack, *.pkl, *optim*.pt; *.bin and *.index.json beside a checkpoint 1 MiB
Logs *.log, events.out.tfevents.*, trainer_state.json, *.out, *.err 1 KiB
Configs *.yaml, *.yml, *.toml, *.cfg, *.ini, *.args; *.json beside a checkpoint 64 bytes
Datasets *.jsonl, *.parquet, *.arrow, *.npy, *.npz, *.tsv, *.csv 1 MiB

Skipped by name: keys, certificates, .env files, kubeconfigs, SSH keys, and anything named like a token, secret, credential or password. Skipped by directory: .git, node_modules, __pycache__, virtual environments, caches and the tool configuration directories under a home. Model files downloaded from a public model hub are not collected: the hub's cache directories are skipped wherever they sit, and a file whose download record sits beside it is skipped as well. A checkpoint your run saves into the same directory still counts.

Locations at tier 2: the organization share, each member's workspace, and the cluster's shared volume (the volume every worker mounts). Files that exceed 50 GiB are not copied.

Changing the tier

Console → Settings → Data sharing. Any member of the organization can change it; the page shows who set the current tier and when. A change takes effect for every capture after it. An organization may also be placed at a tier under a written agreement with National Compute, such as a research grant; the page says so, and changes to that tier go through National Compute under the agreement.

Revocation and retention

Lowering the tier stops future collection. Data collected earlier is kept under the tier in force when it was collected. Deletion requests are handled case by case where law requires.

Incentive

Organizations at tier 1 or 2 may receive credits on terms National Compute publishes in the console. The amount and form may change. Abuse — padding sessions, synthetic activity, uploading data to earn credits — makes an organization ineligible, the determination entirely at National Compute's discretion.

Redaction

Transcripts and feedback notes pass through automated, best-effort redaction before storage. It targets private keys, bearer tokens, JWTs, credentials in URLs, secret-named fields in JSON and YAML, and Kubernetes Secret objects. It is best effort: do not paste credentials into Marshall.

The program is governed by the Terms of Service, the Privacy Notice and the Data Processing Addendum.