Where does your data go when you use AI? It depends on the product
One vendor can sell a consumer app, an enterprise API and a model in your own cloud, each on different terms. Four deployment models, compared honestly.
"Does the AI company keep our data?" is the first question a sensible director asks, and it has no one-word answer, because the same vendor typically sells several products under several sets of terms. The useful question is: which product, on which plan, deployed where?
Here are the four deployment models we design around, with what the vendors' own published documentation says at the time of writing. Policies change; verify before you rely on any of this.
1. Enterprise AI APIs
Anthropic and OpenAI both state that, by default, inputs and outputs sent to their commercial APIs are not used to train their models. Both describe short default retention periods for abuse monitoring, and both offer zero-data-retention arrangements for eligible use. The vendor is a processor under a contract; data leaves your environment for inference.
Fits: general reasoning and drafting over content you have minimised or that is not sensitive. Watch: what is sent has to be designed, and processing location may be outside the UK, which means a transfer assessment.
2. Models inside your own cloud
Models sold by Azure inside Microsoft Foundry run in your subscription and, for standard deployments, your chosen geography. Microsoft states prompts and completions are not available to OpenAI, are not used to train foundation models, and that abuse-monitoring storage can be switched off for approved customers. Claude on Amazon Bedrock offers regional endpoints for data-residency requirements.
Fits: regulated firms that want frontier models without a new data relationship: your contract, your keys, your logs. Watch: more to configure and operate than an API key.
3. AI inside the platforms you already run
Microsoft Copilot in Microsoft 365 stays within your existing Microsoft 365 commitments, is not used to train foundation models, and respects the permissions and sensitivity labels you already have.
Fits: document and email assistance in a well-governed tenant. Watch: oversharing inside the tenant, and the tenant boundary itself.
4. Self-hosted and open-weight models
No model provider receives anything. The security and operation of the entire stack becomes yours, or your partner's on your behalf.
Fits: narrow, high-sensitivity tasks where nothing may leave the environment. Watch: lower capability per pound, and an operational burden that is routinely underestimated.
What stays the same in every model
Whichever row you choose, the parts you control are identical: what data is retrievable, what is sent, what is logged, who approved what, and what can be proved afterwards. Those are architecture decisions, and they are the product. The model is a component you should be able to swap.
How to choose
Classify the data first. Then pick the row per workflow, not per company. A firm can use an enterprise API for marketing drafts and an in-tenant model for client files, through the same interface, with the same log. That is a routing decision, and it should be one.
Want help with this in your business?
Talk to Foundry — we’ll talk through your situation, no obligation.