Local AI vs Cloud: Where Should Your Private Data Go?
The free chatbot is not a filing cabinet. A plain-language guide to which AI can touch which data.
The question isn't "is AI safe?"
Somewhere in your business right now, someone is pasting something into a chatbot. A customer complaint they want rewritten. A contract clause they want explained. A spreadsheet of last quarter's numbers they want summarized. It saves them an hour, so they'll do it again tomorrow.
Whether that's fine or a serious problem depends entirely on two things: which tool they pasted it into, and what the data was. "Is AI safe for our data?" has no yes-or-no answer. "Is this AI safe for this data?" always does.
This piece gives you the two distinctions that decide it, and a three-tier rule you can hand to your whole team.
Distinction one: consumer AI trains on you, enterprise AI doesn't
The single most important line in the AI market runs between consumer accounts and business accounts of the very same product. Same chat window, same model, radically different data terms.
On the consumer side, your conversations are, by default, raw material. All three major vendors use consumer chats to train future models unless you find the setting and switch it off:
ChatGPT Free, Plus and Pro accounts have model training on by default; you can opt out in the data controls. ChatGPT Business, Enterprise and the API are the mirror image: not used for training by default.
Claude consumer plans (Free, Pro, Max) have used chats for training by default since late 2025, with retention up to five years if you're opted in; there's a toggle in privacy settings. Claude for Work, Enterprise and the API sit under commercial terms: not used for training.
Gemini consumer apps use your activity for training, and consumer conversations can be read by human reviewers to improve the product. Gemini through Google Workspace carries enterprise protections: no training on your data, no human review.
Notice the pattern, because it's the whole story: every vendor sells a version that trains on you and a version that contractually doesn't. The free tier isn't a gift; your data is part of the price. The business tier costs money precisely because your data stops being part of the price.
So the fix for most "our staff are pasting things into ChatGPT" anxiety isn't banning AI. It's paying for the business tier of the tool your people already use, and making the consumer version off-limits for work data. That one change moves you from "our client emails may end up in a training run" to "the vendor is contractually barred from training on them."
The fine print moves. These defaults are accurate as we publish (August 2026) and each vendor's current terms are one click away above. Two gotchas worth knowing today: clicking thumbs-up or thumbs-down on a reply can opt that conversation into training even if you've opted out, and "temporary" chats are typically still retained for around 30 days for abuse screening. Assume anything typed into a consumer tool is out of your hands.
Distinction two: "won't train on it" is not "never sees it"
Enterprise terms remove the biggest risk, but be clear-eyed about what they don't change: your data still travels to someone else's servers, gets processed there, and usually sits in logs for some retention window. The vendor won't teach its next model with your payroll file, but the file still made the trip.
For most business data, that's a perfectly reasonable deal. The same trip happens every day with your email provider, your accounting software, and your cloud backups. Enterprise AI with no-training terms belongs in the same category as those: a vetted processor doing a job.
But some data can't take the trip at all. Not because the vendor is untrustworthy, but because the obligation attached to the data doesn't care about vendor promises:
- Client files you hold under an NDA or professional privilege.
- Health information and anything covered by privacy legislation.
- Records you'd have to disclose, or explain, in a breach or a lawsuit.
- Credentials, keys, and anything that unlocks something else.
For that category there's a third option, and it's more practical than most people think.
Local AI: the data never leaves the building
Modern open models run on hardware a small business can own. Not a data centre: a capable workstation or a single server. The model lives on your machine, the documents live on your machine, and the conversation between them never touches the internet. There is no privacy policy to read because there is no third party in the room.
Honest trade-offs, because there are real ones:
- Capability. The best cloud models are still smarter than the best local ones. For drafting, summarizing, extraction, classification, and question-answering over your own documents, good local models are more than enough. For frontier reasoning work, they aren't.
- Cost shape. Cloud is a subscription; local is hardware up front plus upkeep. Local wins on heavy steady use, loses on occasional use.
- Someone has to run it. Models update, disks fill, software breaks. It's less work than people fear, but it isn't zero, and it has to be someone's job.
The point of local AI isn't to replace the cloud. It's to give the most sensitive slice of your data a place to get AI's benefits without leaving your custody. We're not theorizing here: we run this exact stack ourselves, daily, on our own hardware, and the sensitive parts of our own business never leave it.
The three-tier rule
Put the two distinctions together and you get a rule simple enough to actually get followed:
Tier 1 · Truly private → local only. Client files under NDA, health and legal records, regulated data, credentials. If a leak means a lawsuit, a regulator, or a lost client, it gets processed on machines you control, or not by AI at all.
Tier 2 · Business-sensitive → enterprise cloud with no-training terms. Internal documents, financials, customer communications, strategy. Use the business tier of a major vendor, under a plan that names in writing what touches this data. Never the consumer tier.
Tier 3 · Everything else → best tool for the job. Public information, generic drafting, research, brainstorming, anything already on your website. Use whatever is most capable and convenient; this is where consumer tools are fine.
The work is deciding, once, which of your data lives in which tier and writing it down on one page. From then on nobody has to reason about privacy policies at the moment of pasting; they just have to know which tier the thing in their clipboard belongs to.
The takeaway
- Consumer AI trains on your chats by default; enterprise tiers of the same products contractually don't. Most AI privacy problems dissolve by paying for the right tier.
- "No training" still means your data travels. For data under NDA, regulation, or privilege, the only clean answer is AI that runs on hardware you control.
- Local AI is a real option for a small business today: capable enough for document work, and the private stuff stays home.
- Sort your data into three tiers once, write it on one page, and the daily decisions make themselves.
The businesses getting this right aren't the ones avoiding AI, and they aren't the ones pasting everything everywhere. They're the ones who decided, deliberately and in advance, which data gets which tool.