Can You Opt Out of AI Training on Your Business Data? What to Actually Check
The training opt-out toggle in your AI tool's settings is only part of the picture. Here's what business owners should actually check before feeding customer data into AI.

A client asked me last month whether it was safe to paste customer emails into ChatGPT to draft replies. He'd found the setting that says "don't use my content to train your models" and switched it on. Job done, he thought. It isn't, quite. That toggle answers one question out of several, and it's usually not the most important one.
If you're running a business and using AI tools daily, or thinking about wiring AI into your processes properly, it's worth knowing what that opt-out actually covers and what it leaves untouched.
What the toggle actually does
On most consumer AI apps (ChatGPT, Gemini, Claude's web interface), there's a setting that stops your conversations being used to improve future versions of the model. Switch it on and your prompts shouldn't end up shaping how the model responds to someone else six months from now.
That's real and worth doing. But it's a narrow promise. It says nothing about:
- Retention. Your data can still be stored for weeks or months for abuse monitoring, even if it's never used for training.
- Human review. Some providers reserve the right to have staff or contractors look at flagged conversations, training toggle or not.
- Subprocessors. The tool you're using might sit on top of another company's infrastructure, and that infrastructure has its own policies.
- Where the data lives. "Not used for training" doesn't mean "stored in the UK" or "stored in the EU". For some businesses, that matters more than the training question.
API access usually behaves differently to the consumer app. Most major providers now say API inputs and outputs aren't used for training by default, full stop, no toggle needed. But "by default" is doing some work in that sentence. Enterprise agreements, free tiers, and beta features sometimes carry different terms. The only way to know is to read the actual data processing terms for the specific product and plan you're on, not the marketing page.
Why this matters more for businesses than individuals
If you're an individual asking an AI tool to help you plan a holiday, none of this is high stakes. If you're a business pasting client contracts, patient notes, HR complaints or financial records into a chat window, you've got obligations that don't disappear because the tool is convenient. Under UK GDPR, you're still the data controller. You still need a lawful basis for processing, and you still need to know where that data is going and who can access it.
I've written before about why AI tools that remember things need checks, and the same logic applies here. A tool that behaves helpfully doesn't automatically mean it's handling your data the way you assumed. The gap between "this feels safe" and "this is actually documented as safe" is exactly where problems start.
What to actually check before you commit
Rather than trusting a settings toggle, look for these things in the provider's actual terms:
- A data processing agreement (DPA) you can sign, not just a public privacy policy.
- Explicit retention periods, in days or months, not "as needed".
- Confirmation of whether API traffic is used for training, and whether that's the same on every plan tier.
- A list of subprocessors, so you know if your data is passing through a third party's servers too.
- Data residency options, if you need data to stay in the UK or EU for contractual or regulatory reasons.
If a provider can't produce a DPA or won't answer where data is stored, that tells you something on its own, regardless of what the training toggle says.
Where bespoke automation changes the picture
This is one of the reasons businesses come to me instead of gluing together consumer AI tools by hand. When I build automation around AI, I choose the provider, the plan, and how data flows through the system. That means I can tell you exactly which API is being called, what its training and retention terms are, and where the logs sit. There's no guessing based on a settings page that might change with the next product update.
It also means the data doesn't have to touch a general-purpose chat interface at all. A lot of what businesses actually need, summarising a ticket, drafting a reply, extracting fields from a document, can run through an API call inside your own system, with your own audit trail, rather than someone copying and pasting sensitive text into a browser tab. That's a smaller, more controllable surface area than "everyone on the team has a personal AI subscription and uses it however they like."
If you're already running processes that involve customer or staff data through AI tools and you're not sure exactly where that data goes, that's worth sorting out before it becomes a bigger problem. And if you're weighing up whether to keep patching together consumer tools or build something with proper controls around it, that's a conversation I'm happy to have. Have a look at how I approach systems integration, or just get in touch and tell me what you're trying to do.


