Ten vendors, read against their own terms
The model is not what protects your files. The contract is.
The same model is trained on by default in one product, contractually barred from training in a second, and processed inside a country you choose in a third. Nothing about the model changed. What changed is the paper underneath it.
Checked on 7 August 2026. Every claim below was read off the vendor's own terms, privacy notice or documentation, and every table links to the pages it came from. These documents change without announcement — the most recently updated of them changed eleven days before this page was checked. When something here stops being true, the line gets re-checked and re-dated rather than quietly reworded. If the date above is old, treat the page as old. That is what it is there for.
What actually decides where your files end up
Nobody sells you a model. They sell you an account, and the account arrives with terms attached.
Which is why "we use OpenAI" or "it runs on Claude" tells you almost nothing. The same underlying model sits behind a consumer app that trains on your input, a paid API that is contractually barred from it, and a hosting platform that will name the continent it runs on. Four questions separate them, and none of the four is about the model.
- Does it train on what you type, by default?Not whether it can be switched off. What happens if nobody switches anything.
- Where does the data physically sit?Which country, and whether you get to pick.
- How long is it kept, and can that be set to zero?Almost everyone keeps something. The question is what, and for how long.
- Is any of that a contract, or a setting?A setting can be changed by whoever holds the account. A contract cannot.
The last one is the one people miss. A privacy toggle protects you until somebody in your office clicks it back, or the vendor changes what the toggle means. A term in a signed agreement is a different kind of object.
How to read the ten
These are ten vendors a small firm actually runs into. They are listed alphabetically, and that is the only order on this page. Nothing here is ranked and nothing is scored, because which one is right depends entirely on what you are holding and what you signed — and because a ranking written today would be wrong inside a month.
Three things are worth knowing before the tables. Every vendor here that has both a consumer tier and a business tier gives different answers for each, and the consumer tier is the worse one every single time. That is not a coincidence; the free product is paid for somehow. Five of the ten publish weights you can download, which answers the question by making it yours — with a licence catch on three of them. And "not published" appears in these tables as an answer, because on a page about custody, a vendor declining to state a retention period is more informative than most of the vendors who state one.
Does it train on what you type?
By default, with nobody touching any settings. The colour repeats what the words already say; it is not a score.
| Vendor | Consumer tier | Business or API tier |
|---|---|---|
| Anthropic Claude | Opt-in. Off unless you choose to turn it on. | Contractually never. "Anthropic may not train models on Customer Content from Services." |
| DeepSeek | Yes, in effect. Group companies process user personal data for, among other things, "foundation model training and optimization". | Not published. The API terms of service — the document a business would sign — contain no clause either way. |
| Google Gemini | Yes, and people read it. A subset of chats goes to human reviewers. Google's own warning: don't enter confidential information you wouldn't want a reviewer to see. | Free tier: yes. "Google uses the content you submit… to provide, improve, and develop Google products and services." Paid tier: no. Same API. |
| Meta Llama | Not verified. Meta's consumer AI privacy page would not render for reading. | Not applicable if you self-host. Nobody sees it but you. |
| Microsoft Foundry / Azure | — no consumer tier | No. Prompts and completions "are NOT available to OpenAI or other providers" and "are NOT used by providers… to improve their models or services". |
| Mistral | Yes, by default. Your input and output are used to train, "subject to your opt-out" — which lives in account settings and which you have to go and find. | Same policy. The training clause is not scoped to a tier. Abuse retention is separately 30 days unless zero retention is switched on. |
| Nous Research Hermes | Not verified. The hosted portal's terms and privacy pages refused to load across two clients. If you run the weights yourself the question does not arise. | |
| OpenAI GPT | Could not be checked. OpenAI's consumer and enterprise privacy pages refuse automated readers — see the last section. | No, by default. "As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in…)." |
| Qwen Alibaba Model Studio | Not verified for the consumer chat product. | No. "Alibaba Cloud strictly protects your data privacy and will never use your data for model training." |
| xAI Grok | Yes, by default, with an opt-out — and two holes in it, written into the FAQ itself. See below. | Not verified. The enterprise FAQ returns a 404. |
Read from: Anthropic consumer · Anthropic commercial terms · DeepSeek privacy · DeepSeek API terms · Gemini consumer · Gemini API terms · Microsoft Foundry · Mistral privacy · OpenAI API data controls · Alibaba Model Studio · xAI FAQ. Source links open in a new tab.
Where your data physically sits
The question your insurer asks, and the one almost nobody can answer about the tools already in their office.
| Vendor | Where it is processed or stored | Can you choose? |
|---|---|---|
| Anthropic Claude | Not published. No Anthropic page stating a processing country or region was found on this pass. | Unknown |
| DeepSeek | Mainland China, stated plainly: "we directly collect, process and store your Personal Data in People's Republic of China." The API terms are governed by PRC law. | No |
| Google Vertex / Gemini Enterprise | Inputs, outputs and derived data cached in memory for 24 hours, isolated per project and not written at rest. Residency follows the region you select. | Yes. Caching can be switched off per project; the region governs residency. |
| Meta Llama, self-hosted | Wherever you run it. That is the answer, not a dodge. | Entirely |
| Microsoft Foundry / Azure | Processed in the geography you specify — unless the deployment type is Global or DataZone. Stored data stays in your own Azure tenant in that geography. For models sold by Azure deployed in the EEA, the Microsoft staff who can reach abuse data are in the EEA. | Yes — and the deployment type quietly changes the answer. See below. |
| Mistral | Providers "within the European Union" are prioritised; non-EU providers are permitted "in exceptional cases" under standard contractual clauses. No country, datacentre or cloud provider is named. | Partly |
| Nous Research Hermes | Not verified for the hosted service. Self-hosted, it is wherever you run it. | Unknown |
| OpenAI API | Data residency offered in the United States, Europe (EEA and Switzerland), the United Kingdom, Australia, Canada, Japan, India, Singapore, South Korea and the United Arab Emirates. | Yes, per project, for eligible customers. |
| Qwen Alibaba Model Studio | The region sets both the access point and the storage location: Singapore, US (Virginia), Beijing, Hong Kong, Tokyo or Frankfurt. "Your static data always remains in the selected region." | Yes, explicitly — including a scope option that data must not pass through the Chinese mainland. |
| xAI Grok | Not published. The privacy policy names no storage country. The FAQ says only that xAI is "a United States-based company". | Unknown |
Read from: OpenAI API data controls · Vertex AI data governance · Mistral privacy · Alibaba Model Studio regions · DeepSeek privacy · xAI privacy policy · Microsoft Foundry.
How long it is kept
Two different clocks run here: how long your content is stored, and how long anything derived from it is stored. The second one is longer almost everywhere.
| Vendor | Retention | Can it be set to zero? |
|---|---|---|
| Anthropic Claude, consumer | Deleted chats leave your history immediately and back-end storage within 30 days. Opt in to training and de-identified data may sit in training pipelines up to 5 years. Content flagged by trust and safety: up to 2 years, with classifier scores up to 7. Feedback you submit: 5 years. | No |
| Anthropic Claude, commercial | API inputs and outputs delete automatically within 30 days. Usage-policy violations: 2 years, with classifier scores 7 years. | Yes, on eligible APIs, arranged through sales. Safety classifier results are still retained. |
| DeepSeek | Not published as a number. "as long as necessary to provide our Services", varying by "the amount, type and sensitivity". | Not published |
| Google Gemini, consumer | Activity on: auto-deleted after 18 months. Human-reviewed chats: up to 3 years. Activity off or temporary chats: 72 hours — and even then chats are still used to respond and to protect the service, human reviewers included. | No |
| Google paid API / Vertex | Policy-violation logging "for a limited period of time", with no number given. 30 days for grounding with Search. In-memory cache 24 hours; Live session resumption stores prompt, audio and video up to 24 hours and is off by default. | Achievable — but as a checklist of things you switch off, not a contract term. |
| Microsoft Foundry / Azure | The abuse-monitoring store is per-geography, with human review gated behind flagging, secured workstations and just-in-time approval. Stateful features persist until you delete them. | Yes, by application — and you can verify it yourself from the portal or the command line. Nobody else here offers that. |
| Mistral | API abuse logs 30 rolling days. Le Chat conversations until you delete them or close the account. Fine-tuning data until deleted. Identity data 5 years after termination. | Yes, named in the privacy policy itself. |
| Nous Research Hermes | Not verified | Not verified |
| OpenAI API | Abuse-monitoring logs up to 30 days for most endpoints. Application state ranges "from none to indefinite" depending on which endpoint you use — stored assistant data persists until deleted. | Yes, subject to OpenAI's approval. Worth knowing: it forces the API's own store parameter to false even when a request asks for true. |
| Qwen Alibaba Model Studio | Not published. No retention period appears on the privacy notice or the regions page. | Not verified |
| xAI Grok | Private Chat conversations are deleted within 30 days. Normal conversations: not published as a number — "the specific retention periods depend on the nature of the information". | Not verified |
Read from: Anthropic consumer retention · Anthropic organisation retention · Anthropic zero retention · Gemini consumer · Gemini API terms · Vertex AI data governance · OpenAI API data controls · Mistral privacy · Alibaba Model Studio · xAI FAQ · Microsoft Foundry.
Can you run it yourself?
This is the only row on the page that removes the question instead of answering it. Run the weights on your own hardware and no prompt leaves your building, which makes every retention policy above irrelevant. It is also the strongest privacy position available, and it fits on one desktop graphics card.
| Family | Weights you can download | Licence |
|---|---|---|
| Anthropic Claude | No | API only |
| DeepSeek | Yes | MIT — the most permissive licence on this table |
| Google Gemini | No for the Gemini line. Gemma is a separate open-weight family, not checked on this pass. | API only |
| Meta Llama 4 | Yes | Llama 4 Community License — not open source. See below. |
| Microsoft Foundry / Azure | Not a model family — a place to run other people's | — |
| Mistral | Yes, mixed | Mistral Small 3.2 is Apache 2.0. The flagship Large 2411 is a research licence, not Apache. |
| Nous Research Hermes 4 | Yes | The 405B model card declares the Llama 3 licence and names Meta's 405B as its base, so it inherits Meta's restrictions rather than granting its own. |
| OpenAI GPT | No for the flagship line. OpenAI has released open-weight models separately; not checked on this pass. | API only for the flagship line |
| Qwen Qwen 3 | Yes | Apache 2.0 on frontier-class models |
| xAI Grok | Not checked on this pass. Older Grok weights were released; the current line was not verified. | — |
Licence text read directly from each repository: Llama 4 · Qwen 3 · DeepSeek · Hermes 4 · Mistral Small 3.2.
"Open" is doing a lot of work in three of those
Llama 4 is downloadable and it is not open source. Its licence requires a separate grant from Meta, at Meta's sole discretion, once a licensee passes 700 million monthly active users; requires "Built with Llama" to be displayed prominently; and requires any model trained on its outputs to be named beginning with "Llama". Mistral's flagship is under a research licence. Hermes inherits Meta's terms rather than writing its own.
Only Qwen, under Apache 2.0, and DeepSeek, under MIT, are unrestricted open source here. Which is worth saying plainly, because the two most permissive licences on this page are both Chinese — and DeepSeek's own hosted service is the one that stores personal data in mainland China. Those two facts sit together and they do not cancel out: the licence governs weights you run yourself, and the privacy policy governs their service. Running the weights on your own hardware is not the same decision as opening an account.
One model · three arrangements
Same model. Three different answers.
Nothing below changes the model. What changes is the paper underneath it — and that is the only thing deciding where your file ends up. The rows are ordered by how much of that decision is yours to make.
Consumer account
Decided bya setting
the free tier is paid for somehow
Business contract
Decided bywhat you signed
the same model as the row above
Your own hardware
Decided bynothing — there is no counterparty
a machine you own
The model never changed. The paper did.
This is not a recommendation, and it is not a score. Every row above is a real arrangement offered by vendors on this page, and the same company frequently offers more than one of them. Five of the ten had downloadable weights verified on this pass, three of those five carry a licence catch worth reading, and three more have published weights that this pass did not check — running them yourself removes the counterparty, it does not remove the homework. Which row you belong on depends entirely on what you are holding and what you signed.
What each one is honestly good and bad at
One sentence each, no marketing language. This is the fastest-decaying table on the page — capability moves in weeks where terms move in quarters — so weigh it lightly against everything above it.
| Vendor | Genuinely good at | Genuinely bad at |
|---|---|---|
| Anthropic Claude | Long documents, writing that reads like a person wrote it, and coding across a whole repository. | Being cheap — it is among the most expensive per token. And the five-year training-pipeline retention on its consumer tier, if you opt in, is worse than several rivals here. |
| DeepSeek | Extremely cheap, MIT-licensed weights, and reasoning quality far above its price. | The consumer service stores personal data in mainland China with no published retention period, and the API terms are silent on training. |
| Google Gemini | Very large context windows, native multimodality, and hard to beat if your firm already lives in Google Workspace. | The consumer tier is the most aggressive on this page: human reviewers, three-year retention on reviewed chats, and Google's own documentation telling you not to type anything confidential. |
| Meta Llama | The default self-hosting choice, with the deepest tooling and fine-tuning ecosystem of any downloadable family. | A licence that is not open source, with naming and attribution obligations that surprise people who assumed it was. |
| Microsoft Foundry / Azure | The clearest, most specific and most independently verifiable data-handling documentation on this page, without qualification. | Complexity. The guarantees depend on deployment type, region and whether abuse monitoring was modified, and getting one of those wrong quietly changes the answer. |
| Mistral | A European vendor with EU-leaning processing and genuinely good small models that run on modest hardware. | Training on your input is on by default and you have to go and switch it off — the most surprising default from a vendor whose pitch is European privacy. |
| Nous Research Hermes | Neutral alignment and steerability — it follows the system prompt rather than the vendor's house policy, which matters for niche business use. | A small organisation with thin published legal documentation, and it inherits Meta's licence restrictions rather than granting its own. |
| OpenAI GPT | The broadest ecosystem — tooling, SDKs, voice, images and third-party integrations all in one place. | Terms that move frequently, and a product surface large enough that "we use OpenAI" no longer tells you which retention regime you are under. |
| Qwen Alibaba | The strongest combination of capability and permissive licence available — Apache 2.0 on frontier-class models, with strong multilingual coverage. | Anyone with China-exposure concerns has to read the region and deployment-scope settings carefully, because the defaults are not the restrictive ones. |
| xAI Grok | Real-time access to X data and a looser content posture than any US rival. | The thinnest published data documentation of the US vendors: no storage country, no retention number for normal chats, and a 404 where the enterprise FAQ should be. |
Where the marketing and the terms disagree
Four cases where what a vendor is known for and what its own documents say are not the same thing. None of this is hidden. All of it is in writing, on their sites, and none of it is in the pitch.
Mistral sells European privacy and trains on your input by default
The positioning throughout is the European, sovereignty-respecting alternative. The privacy policy says input and output are used to train, "subject to your opt-out" — a setting in your account that you have to go and find. And the hosting clause is a preference rather than a guarantee: EU providers are "prioritised", non-EU providers are permitted "in exceptional cases", and no country or provider is named anywhere in it.
On Google's API, paying is the privacy switch
The free tier is marketed as an on-ramp for developers. Its terms are blunt about the price: on unpaid services, "Google uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products and services." The paid tier of the same API does not. Every prototype anyone built on a free key was training data.
xAI's opt-out has two holes, written into the FAQ itself
Turn training off and new conversations are not used. But feedback you volunteer "may be used for training purposes" even afterwards, and unauthenticated use may be retained "on an anonymous basis". Neither is hidden. Both are easy to miss in a document that opens by telling you the opt-out works.
DeepSeek's silence
The consumer policy is candid about storing data in China. The API terms of service — the document a business would actually sign — say nothing at all about whether your inputs train the model. An unanswered question in a contract is not a no.
The fifth case is Anthropic's, and it is at the top of this page.
What could not be checked
Six things are missing from this page. They are findings rather than holes, and listing them is the only thing that makes the rest of it worth reading.
- OpenAI's consumer and enterprise privacy pages.
openai.comand its help centre return an HTTP 403 to automated readers — a browser gets the page, a program does not. So ChatGPT's consumer training default, its opt-out, and Enterprise and Team retention are unverified here, and this page does not state them. The OpenAI facts above come from the developer documentation, which does answer, and are quoted from it directly. Anyone can reproduce both results. - Anthropic's processing region. No primary Anthropic page stating which country inputs are processed in was found on this pass. It is recorded as not published rather than asserted either way.
- Meta's consumer AI app. The privacy page rendered as headings with no body text, so Meta's consumer training practice and its regional opt-out are unverified. The Llama licence facts are solid and come from the licence file itself.
- Nous Research's hosted portal. Its terms and privacy pages returned HTTP 429 three times across two clients. A search result described a "Privacy Mode" opt-out under which inference payloads are not stored or trained on. That text has not been read on the page itself, so it is not stated as fact here.
- xAI's business tier. The enterprise FAQ returns a 404, so business-tier training, retention and zero retention for Grok are unverified.
- Open-weight releases from OpenAI and Google. Both exist. Neither was checked on this pass, so both are marked that way rather than described.
Every one of those could have been filled in from memory and nobody would have noticed. That is precisely why they are not.
Which of these is already in your building?
You do not need to pick one off this page. The more useful question is which of them your staff are already using, and what each of those keeps — and the answer is usually shorter than people expect and rarely the one they would have guessed.
Tell me what your firm does. I will tell you where your files are currently going, what is being kept, and what could run on hardware you own instead. No charge for that, and nothing owed afterwards.
If the answer is that you are fine as you are, that is what you will hear.
Let's talk