This is the post in the series that is worth your time even if you never speak to anyone about any of it. There is nothing to buy here and nothing at the end.
The five checks
Do this for one tool first — whichever one the most client material passes through. The ten minutes below is per tool, and doing one properly teaches you what you are looking for, which makes the rest faster than the first.
- Write down the plan, not the product Write the exact subscription name. Not "ChatGPT" — the tier. Not "we have Microsoft" — the licence. This is the single most common gap, and every other check depends on it, because vendors publish different commitments for different plans of the same product. 2 minutes
- Find the vendor's own page that answers "do you train on this" Their page. Not a blog post about their page, not a comparison article, not a summary. Search the vendor's own documentation for "train" and read the paragraph. If you cannot find such a page for a tool somebody is using, that is a finding in itself and it is arguably the most useful thing this exercise will produce. 3 minutes
- Read the carve-outs on that same page The reassuring sentence is rarely the whole document. Microsoft's enterprise data protection page is the clearest published example we found: it says prompts and responses are not used to train foundation models, and on the same page notes that web search queries go to Bing under a different agreement, that add-on agents carry their own terms, and that the specific controls vary by subscription plan. That is good disclosure. It just requires reading past the first paragraph. 2 minutes
- Retention, and whether you can delete How long is a conversation kept, and can somebody at the firm remove it. Two consequences hang off this and both are practical: what exists to be produced later, and what exists to be breached. 2 minutes
- Who else is in the chain Subprocessors, add-ons, browser extensions, anything a user can switch on themselves inside the tool. This is where an otherwise careful setup usually leaks, because the person who switched it on experienced it as a feature rather than as a vendor. 1 minute
What we ran into doing this ourselves
We tried to fetch OpenAI's own data-usage documentation to quote it in this series. Three different URLs, three HTTP 403s — their site blocks automated fetching. So we have not quoted OpenAI anywhere in these five posts, and we are not going to fill the gap with somebody's summary.
That is a small annoyance and a genuinely useful lesson. Every second-hand account of that policy we came across said roughly the same thing, which feels like corroboration and is not — it is one claim repeated, and we could not see the original. The pages we could read cleanly were on Microsoft Learn, which carries a visible date stamp on each one, and that turned out to matter more than we expected: the page we quoted had been updated the day before we read it.
So: open the vendor's page in a browser and read it yourself. It takes the same three minutes and it is the difference between knowing something and having heard it.
What you end up with
One page, five columns, one row per tool.
That last column is the one people leave off, and it is the one that makes the rest of it worth anything a year from now. These pages change. A finding without a date is a memory.
The document you have just made is not a policy and it is not compliance with anything. It is the answer to "where does our client material go", written down, which is the factual question sitting underneath every ethics opinion on this subject — and the one you would be asked first, by a client or by anybody else, on the worst possible day.
Expect at least one row you cannot complete. That row is the whole value of the exercise — a tool in daily use whose terms nobody has read is a more useful thing to have found than four tidy rows.
What to do
Do the first one this week. Take the list of tools from the first post in this series, pick the one that sees the most client material, and spend ten minutes filling in its row.
Then budget roughly ten minutes a tool for the rest, and do not do them all in one sitting unless the list is short. Five tools is an hour, and an hour is a thing that gets postponed. One row is not.
It does not have to be you. Every one of the five checks is reading and writing down, so it hands off to whoever runs operations without losing anything — what comes back is a page, and a page is a thing a partner can read in two minutes and act on. The only part worth insisting on is the last column. A source URL and the date you read it is what separates this page from a memory of a page.
What it costs to not have it is one sentence, and it is the sentence this whole series has been circling: when somebody asks where your client material goes, you either have the page or you have a guess. The first row takes ten minutes, and it is the same ten minutes whether you ever speak to anyone about any of this or not.