Enterprise AI assistant security: where the data actually goes
Three things leave your network on every question: the user's prompt, the passages your system retrieved to answer it, and the model's reply. Everything else, meaning your documents, your index, your database and your permission rules, stays where you put it. Residency decides which country processes those three things. The data processing agreement decides who may do what with them, and for how long. Both need answering in writing, and neither one is implied by the other.
Every buyer asks the same question on the first call, and almost always in a form that cannot be answered: is it secure? The answerable version is narrower and much more useful. On each question a user types, which bytes leave your network, who receives them, what are they permitted to do with them, and how long do they keep them.
This article works through that boundary, the difference between residency and a data processing agreement, how GDPR Chapter V applies when the supplier sits in Türkiye, how to find the no-training commitment in a contract rather than in a brochure, what to ask about logs, and when running the model yourself is worth what it costs.
What leaves and what stays
An assistant that answers questions about your company does not contain your company. It retrieves the relevant passages at the moment of the question and asks a model to answer from them, which is why updating a document updates the answers and why every answer can carry a citation. The mechanics are set out in what an enterprise AI assistant is.
That design decides the boundary. Three payloads cross it per question, and only three.
| Data | Where it sits | Who can read it | Retention set by |
|---|---|---|---|
| The user's question text | Sent to the model provider | Provider staff under its access policy; your admins | The provider's contract |
| Retrieved passages from your documents | Sent with the question | Same as above | The provider's contract |
| The model's answer | Returned to your system | The asking user; your admins | You |
| Full document store | Your storage, in your account | Your roles | You |
| Search index and embeddings | Your storage, unless the index is managed for you | Your roles | You |
| User identity and permissions | Your directory | Your roles | You |
| Prompt and answer audit trail | Your application | Named roles, if you named them | You, deliberately or by accident |
The row buyers underestimate is the second one. A question like “what is our termination notice with this supplier” sends a few words of prompt and several hundred words of the most sensitive contract you own. Retrieval quality depends on sending enough context to answer, so the exposure is set by the design of the retrieval step, not by how careful the user was with their wording.
Two controls follow from that, and both belong in the build rather than in a policy document. Retrieval must be scoped by the asking user’s permissions before anything reaches the model, so the assistant cannot become a way around your access rules. And identifiers that the answer does not need, meaning names, account numbers, national ID numbers, can be masked at the boundary and restored in the reply. Which of the two you need is an access-control question, worked through in where your data is stored and who can reach it.
Residency and the data processing agreement answer different questions
These get merged in almost every RFP, and the merge causes real mistakes in both directions.
Residency is geography. Which country’s data centres hold the documents, run the index, and process the request. It is a hosting decision, it is yours to make, and it is usually easy to verify.
The data processing agreement is permission and accountability. What the supplier may do with the data, on whose instructions, which sub-processors are allowed, how a breach is notified, what happens at termination.
A system hosted entirely in Frankfurt with no processor agreement, no sub-processor list and no named transfer mechanism is not compliant. A system processing in a third country under executed standard contractual clauses, a processor agreement and a transfer impact assessment is documentable. Geography alone decides neither case.
The chain matters more than the endpoint. Your supplier is your processor; the model provider is their sub-processor; whoever hosts the model may be a further one. Ask for the chain in writing, ask how you are told before a link is added, and check that the no-training commitment survives every hop. A promise that stops one link short of the company actually receiving the text is not a promise.
GDPR Chapter V and a supplier in Türkiye
Transfers of personal data outside the EEA need a mechanism under Chapter V. Türkiye does not hold a European Commission adequacy decision, so the normal route is Article 46 appropriate safeguards, in practice the Commission’s standard contractual clauses, supported by a transfer impact assessment that considers the destination’s law on government access. This is routine work, and it is far easier when the data flows are described accurately, which is the real reason the architecture conversation comes before the legal one.
Two points that catch buyers out.
Remote access is a transfer. Under the prevailing regulatory view, support access from a third country counts even when the server stays in Frankfurt and nothing is copied. That does not make it unlawful. It means the same safeguard applies, over a much smaller and more controllable scope.
The UK is a separate regime. EU standard contractual clauses are not sufficient on their own for a UK controller, and Türkiye is not covered by UK adequacy regulations either. The instrument is the ICO’s International Data Transfer Agreement, or the UK Addendum bolted onto the EU clauses where a group already runs on the EU set. Our own position on both, including the fact that we have not appointed an Article 27 representative in the EU or the UK, is written out on the privacy page rather than left for you to assume. The wider version of this question, covering development teams as well as assistants, is in GDPR and nearshore development.
KVKK compliance is not GDPR compliance
Türkiye’s own regime moved towards the European model in 2024. The amendments to Law No. 6698 replaced a consent-based approach with a three-tier structure of adequacy decisions, appropriate safeguards and limited exceptions. Turkish standard contractual clauses must be used exactly as published and notified to the authority within five business days of signing.
Convergence helps. A Turkish supplier operating properly under the amended KVKK already works inside a framework whose vocabulary maps onto yours, which shortens the contract negotiation considerably.
It is still not the same thing. KVKK registration is a Turkish filing obligation and produces no GDPR artefact. A KVKK notice is not an Article 28 processor agreement. And none of it substitutes for the Chapter V mechanism, because that obligation sits on you as the exporting controller, not on the supplier. A vendor that presents its KVKK position as GDPR coverage has either misunderstood the point or is hoping that you will.
Finding the no-training commitment
“We do not train on your data” is said in every sales call and means nothing until you can point at the sentence that says it. Four checks find the real one.
What does it name? A clause covering “personal data” leaves your commercial documents outside it. You want customer content, inputs and outputs, without a carve-out for anything the provider calls aggregated or derived.
Does it cover outputs too? Some terms protect what you send and say nothing about what comes back, which is where the interesting inference sits.
What is the abuse-monitoring carve-out? Nearly every hosted provider retains requests for a short window to detect misuse, sometimes with human review on flagged content. That is honest and worth accepting on its terms: ask for the window in days, whether it can be shortened contractually, and who reads a flagged request. A provider claiming zero human access under all circumstances is either describing a self-hosted deployment or being imprecise.
Does it flow down? Your supplier’s promise is only as good as their contract with the company that actually receives the text. Ask which tier of the provider’s terms they are on. Consumer and free tiers of the same product routinely permit training.
Logs: what is written, who reads it, how long
Ask four questions of each log store, and expect three stores.
Your application audit trail should record who asked, when, which documents were retrieved, and what the assistant answered. Where the assistant acts rather than answers, it also records who approved, which is the same discipline described in designing the approval step.
Your retrieval log records which passages were fetched for which user. It is the one that proves your permission scoping actually works, and it is worth keeping longer than you think.
The provider’s request log is outside your control and holds the retrieved passages, not just the question. Get the retention in days, the region it sits in, and the roles that can read it.
Two habits separate a real audit trail from a decorative one. Set retention deliberately for each store rather than inheriting a default that keeps everything forever. And read it: schedule a quarterly review of who queried what, because a trail nobody has ever opened tells you nothing about whether the controls hold.
Getting an assistant past a German works council
If you deploy in Germany, the gate that stops the project is usually not the data protection officer. It is the works council, and its objection is rarely that data leaves the country. It is that the system can be turned into a record of how individual employees work.
That is a design question before it is a legal one, and it is the part almost nobody writes down. Law firms describe the procedure well. What a council actually asks for is a list of properties the software either has or does not have, and those properties were decided months earlier, by whoever chose what goes in the log table.
The trigger is lower than buyers expect
Section 87(1) No. 6 of the Works Constitution Act (Betriebsverfassungsgesetz, BetrVG) gives the works council co-determination over the introduction and use of technical devices designed to monitor the behaviour or performance of employees. Read literally, that sounds like it covers surveillance tools and nothing else.
It is not read literally. The Federal Labour Court has held for years that objective suitability is enough, meaning it is sufficient that the equipment is objectively likely to record information about behaviour or performance, and that the employer’s intention is irrelevant. GÖRG’s summary of the case law collects the line: subjective intent not required (BAG, 10 December 2013, 1 ABR 43/12); merely storing behaviour data the employer did not collect itself is enough, the social media channels case (13 December 2016, 1 ABR 7/15); ordinary Excel systems included (23 October 2018, 1 ABN 36/18); Office 365 included (8 March 2022, 1 ABR 20/21).
So an assistant that records who asked what and when sits inside the provision on day one, and it sits there precisely because you built the audit trail described in the previous section. The reflex that follows is the wrong one. The answer is not to log less. A system with no audit trail cannot demonstrate that its permission scoping works, and a council that cannot inspect the log design has no reason to believe anything you assert about it. Good logging creates the co-determination duty and then supplies most of the evidence that resolves it.
One decision gets quoted as though it says the opposite. In January 2024 the Hamburg Labour Court refused an injunction against employees using ChatGPT (24 BVGa 1/24, 16 January 2024), finding no Section 87(1) No. 6 right because the tool was not installed on the employer’s equipment, the employer issued no company accounts, and it could not monitor private accounts (Littler’s note on the decision). Read the facts, not the headline. Those facts are the exact inverse of an enterprise deployment, which has company accounts, company documents and server-side logs. The decision is useful for a different reason: the analysis turned on architecture, on where the software ran and who held the account. That is the whole argument of this section.
Three neighbouring provisions shape how the conversation runs.
- Section 90(1) No. 3 and 90(2) give information and consultation rights over the planning of technical installations. The Hamburg court was explicit that these are not co-determination rights. Consultation is a duty to talk, not a veto, and treating it as one wastes goodwill you will need later.
- Section 80(3) lets the works council bring in an outside expert at the employer’s cost, and since the 2021 amendments the involvement of an expert on the introduction or use of AI is treated as necessary rather than as something the employer can argue against (Reed Smith). Expect a technical reader on the other side of the table. Write the documentation for one.
- Section 95(2a) extends co-determination over personnel selection guidelines to guidelines produced with AI support, which is what makes scope creep into HR use so expensive.
Where the parties do not agree, the route is the conciliation committee under Section 87(2), whose award takes the place of the agreement. That path is slower than designing the system so the question does not arise.
The questions a council asks, and where each one is answered in the build
These are product questions with product answers. If your supplier can only answer them in policy language, the answer is no.
| What the council asks | Where it is answered | What a real answer looks like |
|---|---|---|
| Can this produce a performance figure for one named person? | Log schema and reporting layer | No endpoint, no saved query and no export groups usage by named individual, enforced in the query layer rather than by reporting convention |
| What exactly is written when I ask a question? | Application audit trail | A published field list: timestamp, user key, retrieved document ids, answer id. Not keystrokes, not idle time, not per-person latency |
| Who can read a stored prompt? | Role matrix | Named roles, countable, and every read written to a log those roles cannot edit |
| How long is it kept? | Retention configuration | A number of days per store, enforced by a scheduled deletion job that leaves a record of what it deleted |
| Can monitoring be switched on later without us knowing? | Change control | Log fields and retention live in versioned configuration, so any change is a reviewable commit rather than a setting an administrator flips |
| Does the model provider see this, and can it be used against me? | Sub-processor chain and contract | The chain, the tier, the no-training clause and the provider's own retention, as set out earlier in this article |
| Can we verify any of this ourselves? | Council access | Read-only access to the log schema and the aggregate usage report, plus test accounts for checking that retrieval respects permissions |
The first row decides most negotiations. A usage figure computed over a large group is a capacity number. The same figure computed over three people is a performance review. The workable answer is an aggregation threshold: no usage figure is produced below a minimum group size, and that minimum is a parameter both sides agree on and the code enforces. We do not publish a number for it, because there is no statutory figure to publish and inventing one would be worse than useless. What matters is that the threshold is a property of the query layer rather than a promise about how reports will be read.
Design decisions that make the negotiation easy
- Monitoring capability is absent rather than disabled. A code path that does not exist is a far stronger statement than a toggle set to off.
- Log identity is a pseudonymous user key, with the mapping to a person held in your directory and reachable only by the named roles described in where your data is stored and who can reach it.
- The append-only audit trail and the analytics store are separate systems, and the analytics store only ever receives aggregates.
- The log field list ships as part of the product and is versioned with it, so it cannot drift away from the document attached to the agreement.
- Retention is a deletion job with its own log, not a paragraph in a policy.
- Retrieval is scoped by the asking user’s permissions, and the council can test that with its own accounts rather than take it on trust.
Design decisions that make it hard
- Any leaderboard, top-users panel or adoption ranking by name. It is the fastest way to turn a discussion about capability into a discussion about intent, and it is almost never the feature anyone actually needed.
- Response time, session duration or question count recorded per user, even when nobody intends to look at it. Suitability is the test, and those fields are suitable.
- A quality or satisfaction score attached to a person rather than to an answer.
- An administrator who can read any stored prompt without that read being logged.
- Free-text prompts kept indefinitely because someone might want them for evaluation later. Undefined retention is the hardest single item to defend in the room.
- An assistant that quietly grows into drafting appraisals, ranking candidates or allocating tasks by individual behaviour. That is a different system with a different legal classification, which is the next point.
One thing worth being blunt about: several of these are not choices when the software is not yours to specify. With a per-seat product you accept the telemetry the vendor ships, and the field list is theirs to change on their release schedule. That constraint, rather than price, is often the real reason a German buyer ends up commissioning rather than subscribing, and it belongs in the comparison set out in custom AI agent versus off-the-shelf copilot.
Where the AI Act joins in
The two regimes reward the same design decision, which is convenient.
Annex III point 4(b) of the EU AI Act classifies as high risk those AI systems intended to be used to make decisions affecting work-related relationships, to allocate tasks based on individual behaviour or personal traits, or to monitor and evaluate the performance and behaviour of persons in such relationships. Obligations for Annex III systems apply from 2 August 2026, while systems caught by Article 6(1) as safety components of regulated products follow from 2 August 2027 (Article 113).
An assistant that answers questions from your own documents is not in Annex III by default. It arrives there by feature creep, one per-person metric at a time. And once a system is high risk, Article 26(7) requires a deployer who is an employer to inform workers’ representatives and the affected workers before putting it into service at the workplace, which returns you to the same room with a heavier agenda.
So the aggregation threshold is not only a concession to the council. It is also the line that keeps the system out of a regulatory category you have no reason to enter. Where the assistant acts rather than answers, the approval records become records about people too, which is why the approval step deserves the same scrutiny as the log schema, as set out in designing the approval step.
What belongs in the works agreement as a technical commitment
Your counsel drafts the Betriebsvereinbarung. What a software supplier can usefully contribute is the set of sentences that are true of the system and testable against it. Ours are these, and we put them in writing before the build rather than after.
- The field list of each log store, attached as an annex that changes only by versioned amendment.
- Retention in days per store, enforced by a scheduled job, with the deletion itself recorded.
- The named roles with read access to stored prompts, the fact that every such read is logged, and the fact that those roles cannot write to that log.
- An aggregation threshold below which no usage figure is produced, enforced in the query layer.
- A commitment that no interface, export or saved query groups usage by named individual.
- Notification before any new log field, any change of retention, or any new sub-processor in the chain.
- Read-only council access to the log schema and the aggregate usage report, plus test accounts for verifying permission scoping.
- What is deleted and what is returned on termination, and in what format.
Two honest limits. We are not German employment lawyers, we do not run the works council process on a customer’s behalf, and none of this is legal advice on your facts. And the sequencing is unforgiving: a log schema is cheap to define before the build and expensive afterwards, because changing it once an agreement exists means reopening the agreement. If a German site is in scope, put the field list on the table in the first design session, not in the week before go-live.
Self-hosting: the real cost, and when it earns it
Self-hosting removes exactly one thing, the transfer of prompt and retrieved content to an external provider. It is a real removal and sometimes the only acceptable answer.
It removes nothing else. You still choose a region, still scope retrieval by permission, still write and retain logs, still sign a processor agreement with whoever hosts the hardware unless it is your own rack. And you take on three costs that hosted use does not carry: GPU capacity sized for peak rather than average load, an evaluation harness, because quality regressions are now yours to catch, and someone who can be called when inference stops at nine in the morning.
The quality gap varies by task and is closing unevenly. Extraction, classification and summarisation over your own retrieved passages run well on open-weight models. Long multi-step reasoning still favours the frontier hosted models by a margin you will notice.
It earns its cost when a data class is legally blocked from leaving the country, when a sector rule or a customer contract forbids third-party processing, or when steady high volume makes per-request pricing exceed the fixed cost of running your own. It does not earn its cost when the reason is that self-hosting feels safer. A hosted API in a named region under a no-training agreement, with permission-scoped retrieval and honest log retention, is a defensible position, and for most companies of 20 to 250 people it is the cheaper one.
The questions to put in front of your vendor
Send these in writing and ask for written answers.
- Which payloads leave our network on each question, and which legal entity receives them?
- In which region are they processed, and in which region are they logged?
- Show me the clause stating our content is not used for training. Does it cover inputs and outputs?
- Which tier of the model provider’s terms are you on?
- Name every sub-processor in the chain, and tell me how we are notified before one is added.
- What is the provider’s log retention in days, and what is ours?
- Which named roles at your company can read a stored prompt, and is that access itself logged?
- Does retrieval respect our permission model before anything reaches the model?
- Which transfer mechanism applies to us, and may I see the executed clauses?
- What is your breach notification commitment, measured in hours?
- On termination, what is deleted, what is returned, and in what format?
A supplier who answers all eleven in writing within a week has told you a great deal about how the project itself will run. One who answers with a security page and an adjective has told you something too. For the record, we do not hold ISO 27001 or SOC 2 certification, and we would rather write that plainly here than let a badge in a footer imply otherwise.
If you want these answers for your own situation rather than in general, describe the data classes involved, where your controllers sit, and which processes you want the assistant to touch, and send that through the quote form. We will come back with the boundary drawn, the mechanism named, and the parts we cannot do written down alongside the parts we can.
Frequently asked questions
Does our data get used to train the model?
That depends on the contract you sign, not on the technology. Major providers offer enterprise terms under which customer content is not used to train or improve models, but the commitment lives in the agreement rather than on the marketing page, and the consumer tier of the same product often says the opposite. Ask for the clause, read what it actually covers, and check that it flows down from your supplier's contract to the provider they call. A verbal assurance from a salesperson is worth nothing at audit.
We host everything in Frankfurt. Does that make us GDPR compliant?
No. Residency answers where processing happens. Compliance also needs a lawful basis, an Article 28 processor agreement that defines the instruction scope, a named sub-processor list, breach notification terms, and a Chapter V transfer mechanism wherever someone outside the EEA can reach the data. Remote administrative access from a third country is treated as a transfer even though the server never moves. Frankfurt hosting narrows the problem and documents well, but it does not close it by itself.
Our supplier says it is KVKK compliant. Is that the same as GDPR compliant?
No, and treating the two as equivalent is a common and expensive error. Türkiye's Law No. 6698 was amended in 2024 into a three-tier structure of adequacy decisions, appropriate safeguards and limited exceptions, deliberately closer to the GDPR's shape. Convergence is not equivalence, and it is certainly not an adequacy decision: the European Commission has not issued one for Türkiye. A supplier operating properly under the amended KVKK is easier to contract with, but transfers from the EU still need standard contractual clauses or another Article 46 safeguard.
How long are prompts kept, and who can read them?
There are usually three answers, because there are usually three separate log stores. Your application writes its own audit trail, your retrieval layer records which passages were fetched, and the model provider keeps request logs under its own retention period. The provider's logs are the ones buyers forget, and they hold the retrieved passages rather than just the question. Ask for the retention period in days, the named roles with read access, whether that access is itself logged, and whether the logs sit in the same region as the processing.
Should we run the model on our own hardware?
Only for a reason you can state in one sentence. Self-hosting removes exactly one thing: the transfer of prompt and retrieved content to an external provider. That matters when a data class is legally blocked from leaving, when a sector rule forbids third-party processing, or when a customer contract does. It does not remove your residency decision, your access controls, your logging duties or your evaluation work, and it adds GPU capacity and someone to run it. For most companies of 20 to 250 people, a contractual no-training arrangement in a named region is cheaper and documents just as well.
Does an internal AI assistant need works council approval in Germany?
In practice yes. Section 87(1) No. 6 of the Works Constitution Act covers technical devices designed to monitor employee behaviour or performance, and the Federal Labour Court reads that as objective suitability rather than intent, so a system that records who asked what and when is inside it from day one. That is not a reason to strip out the audit trail, since without one you cannot demonstrate that permission scoping works. The productive move is to arrive with the log field list, the retention period per store, the role matrix and an aggregation threshold already written down, because those four artefacts answer most of what a council actually asks.
What can a software supplier commit to in a Betriebsvereinbarung?
Only things that are testable in the system, and your own counsel drafts the agreement itself. The commitments worth supplying are the field list of each log store as a versioned annex, retention in days per store enforced by a scheduled deletion job, the named roles that can read a stored prompt plus the rule that every such read is logged, an aggregation threshold below which no usage figure is produced, a statement that no interface or export groups usage by named individual, notice before any new log field or sub-processor, read-only council access to the log schema and aggregate report, and what is deleted or returned on termination. Define these before the build: changing a log schema after an agreement exists means reopening the agreement.
Related guides
- GDPR and a development team outside the EU: how to do it properly
- Human-in-the-loop AI: designing the approval step
- Gulf e-invoicing: what ZATCA and the UAE mandate actually require from your systems
Service page: Custom software service
Let's talk about what you need.
The 30-minute discovery call is free and carries no commitment.