AI data security: where business content really goes
AI data security comes down to two questions that most vendor pages blur into one. What happens to the business content, and what happens to customers' personal data.
They have different answers and different fixes. This guide covers where content goes when someone asks a question. It also covers what the rules require, and the questions worth putting to a vendor before signing.
In this guide
What AI data security actually means
AI data security means knowing three things: what content the system may read, where that content gets processed, and whether anyone may keep it afterwards. Every other question on this topic is a version of one of those three.
This decision already gets made several times a year. Books handed to an external accountant are read in full. They work under a contract, and they do not repeat those figures to their other clients. Choosing who processes content for an AI assistant is the same kind of decision rather than a new category of risk.
Two problems that get treated as one
Business content and customers' personal data raise different questions, and merging them is why the topic feels harder than it is.
- Business content. Price lists, procedures, contract terms and internal instructions. The risk here is commercial, and it comes down to who else could read the prices.
- Customer personal data. Names, phone numbers, order histories, and in some sectors health details. The risk here is legal, and data protection law governs what may be done with it.
The first problem is settled by a contract and by control over who has access. The second needs the short list of paperwork covered further down, most of which already exists, prepared for other systems.
Where the leak usually comes from
In our experience the most common data leak in a small company has nothing to do with an assistant. An employee pastes a customer email into a free public chatbot to draft a faster reply, and nobody records that it happened.
That habit is already established in most companies before any AI project starts. It is unmonitored, undocumented, and far more exposed than a tool bought under contract. A sanctioned tool tends to suppress it, because people improvise less when something approved is available.
Where content goes when a customer asks something
Content goes to a model provider, is read to produce that answer, and does not become part of the model. Under a business agreement it is not used to train models other people use, and it is not exposed to that provider's other customers.
That is the accurate version, and it is worth understanding the mechanics rather than taking any vendor's word for it. Four steps happen between a question and an answer.
- A customer asks a question. The question reaches the assistant, not the model behind it.
- The system searches the approved content. It pulls out the few passages that look relevant to the question.
- Those passages and the question go to the model provider. This is the step people worry about, and the one worth asking about.
- The answer comes back and is shown. Under a business agreement the passages are not added to any model, and the provider keeps them only for the window the contract sets.
Most business agreements now exclude training on customer content. That holds because the contract says so, not because of anything in the technology, and terms differ between providers. Ask for the exclusion in writing, and check that it appears in the contract actually being signed.
Training, retention and logs are three different things
Three separate questions get bundled under the word "storage", and separating them makes the vendor conversation much shorter.
- Training. Whether the text is used to improve a model that other people use. Business agreements normally exclude this, so read the clause rather than trusting the summary.
- Retention. How long the provider holds the text after answering. Expect a defined window measured in days, and longer if something is flagged or a legal hold applies.
- Conversation logs. What the assistant itself keeps so its answers can be reviewed. This part belongs to the business, and it is where personal data actually accumulates.
Zero retention exists as an option, but it is requested and confirmed in writing. It is not a default and should not be assumed.
The third item gets overlooked, and it is the one entirely within the client's control. The logs in the first week show what the system really does. The guide on accurate AI answers explains what to look for.
| Question | Free public tool | Business assistant under contract |
|---|---|---|
| What is agreed | Consumer terms accepted on signup | A processing agreement with named subprocessors |
| Training on the text | Often permitted unless excluded | Excluded where the agreement says so |
| Processing location | Wherever the provider chooses | Data centres in the EU, stated in the contract |
| Retention | Set by the provider, hard to influence | Defined, and deletable on request |
| Who reads conversations | The account holder and provider staff | Named staff, with access set by the client |
| Erasing a customer's data | No agreement obliging it, and no record | A defined erasure process |
What the rules actually require
The rules do not ban AI, and most of the paperwork is the same as for any other supplier. Two things are specific to AI, and one of them applies from August 2026.
The controller is the business
The business decides why the data is processed and broadly how, which makes it the controller. The vendor is the processor, acting on written instructions. It also carries its own duties for security and for reporting when something goes wrong.
That structure is ordinary. An accountant, a payroll bureau and an email provider stand in the same relationship to a business. None of them requires a change to how it operates, and neither does this.
What is needed on paper
- A data processing agreement with the vendor, naming every subprocessor, with a transfer safeguard if any of them sits outside the EU.
- A line in the privacy notice telling customers that an automated assistant handles enquiries.
- An entry in the record of processing activities, alongside the entries for other systems.
- A route to erase one customer's conversation history when they ask for it.
- An impact assessment, where the processing is large scale, covers sensitive data, or is otherwise high risk.
Most of that copies what already exists for the CRM. Where there is no privacy notice at all, that gap exists with or without an assistant, and it is worth closing either way.
Telling people they are talking to a machine
Since 2 August 2026, the EU AI Act requires that people be told when they are interacting with an AI system. The duty applies unless that is already obvious, and it covers businesses of every size.
In practice this is a label on the chat window and an opening line from the assistant. It is the cheapest obligation on this page to meet, and the easiest to forget.
Five questions to put to any AI vendor
Ask these before signing, and ask for the answers in writing. A vendor who cannot answer all five in a single email is not ready to handle customer information.
- Which subprocessors handle our content, and in which countries? Names and locations, not a reference to the cloud.
- Is our content used to train any model, and does the contract say so? A verbal assurance here is worth very little.
- How long is our text kept after an answer is produced? Ask for a number in days. A description of good practice is not an answer.
- How do we delete one customer's history, and how do we export everything if we leave? Ask for the actual steps.
- Can we see the processing agreement before signing? A refusal settles the decision.
One more question belongs in the same email, and it is the one vendors answer least precisely. Ask who on their side can read the conversations, and whether every such access is logged. Support access is normal, and support access that leaves no record is not.
Notice that none of these five questions is about AI. They are the questions for any supplier holding customer records, which is the point of asking them.
What stays with the business
A contract covers the vendor. It does not cover what happens inside the office, and that is where most incidents start, in our experience.
- Access to the conversation history. Grant it to the people who need it, and remove it as soon as someone leaves.
- Staff pasting into public tools. Write the rule down, since this is the leak nobody records.
- Personal data in old documents. Check the old quotes and case notes, because they get indexed along with everything else.
- The limits of the assistant itself. Decide what it may never discuss, such as another customer's order.
In our experience, the last one is the item most often missed. A system connected to every order can answer a question about somebody else's order unless that boundary is set. Nothing in the technology sets that limit on its own.
Give the system less to protect
The cheapest security measure is not putting personal data into the system at all. An assistant answering questions about opening hours, services and prices has no need for a customer list.
Names, contact details and case histories come out of the indexed content. Deciding which sources belong in scope is covered step by step in the guide on preparing your content for an AI assistant.
Deciding those limits, writing them into the configuration and testing them before customers arrive is work we do for our clients. We also handle the processing agreement and the subprocessor list, and we run the assistant on an external model provider with processing in the EU. In more sensitive settings, such as an AI assistant in healthcare, we agree the boundaries in writing with the client before go-live.
Frequently asked questions
Short answers to the questions we get asked most.
Is our data used to train the AI?
It is not, provided the business agreement excludes it. That exclusion is standard practice now, but it is still a contract term worth confirming in writing. Free consumer tools are a different matter, and several of them use the text for training unless somebody actively opts out.
Where is our data processed?
That depends on the vendor, which is why it gets asked. Any vendor should state the regions in writing, and to say whether anyone outside the EU can reach the data for support.
Do we have to tell customers they are talking to AI?
Yes, that is now a legal duty. Since 2 August 2026 the EU AI Act requires that people be told when they are talking to an AI system. The duty falls away only where that is already obvious.
Is an AI assistant GDPR compliant?
No tool is compliant by itself, since compliance depends on how it is used. What is needed is a processing agreement, a line in the privacy notice and a record of the processing. Erasing one customer's history has to be possible on request.
Can the assistant see our customer database?
It sees only what it is connected to. An assistant answering questions about services needs no customer records at all, and the simplest risk reduction available is leaving them out.
Is an AI agent riskier than an assistant that only answers?
An agent carries more consequence, because it acts on systems instead of describing them. Each action gets a defined limit on what may happen without human approval. The guide on chatbot vs AI agent covers how those limits are set.
AI data security is a supplier decision before it is a technology decision. The choice is who may read the content, where they may process it and what they may keep. That is the same judgement involved in picking an accountant. Sign the agreement, leave out the personal data the system does not need, and write down who may read the conversations. Tell customers they are talking to a machine. Do those things and the assistant stops being the one to worry about.