Announcing US$ 3.2M pre-seed round · OneVC · Maya · Norte Ventures Read →
Back to blog

Where your data lives
when the company uses AI

TL;DR: In a well-designed context architecture, data stays inside the customer's own infrastructure — private cloud or on-premises — and only the snippet needed to produce the answer travels to the model provider on each query. The data is not copied into the model, nor used to train it. Mapping that flow explicitly is what gets a project through security review.

What actually blocks projects in regulated industries

It is not the model's capability. It is the inability to answer five questions precisely:

  1. Which information leaves our environment on each query?
  2. Does the provider store that content, and for how long?
  3. Is that content used to train models?
  4. Where is the infrastructure that processes the query physically located?
  5. Can we demonstrate all of this in an audit?

Teams that walk into the security review with a clear answer to all five move faster than teams that walk in with an impressive demo.

The map security will ask for

ElementWhere it livesLeaves the perimeter?
Source databaseCustomer environmentNo
Search index and metadataCustomer environmentNo
Query and audit logCustomer environmentNo
Retrieved context snippetRetrieved inside the environmentYes, at call time
User's questionCustomer environmentYes, at call time
Generated answerProviderReturns to the environment

What leaves on each query is the retrieved context snippet plus the user's question, never the underlying base. That design is what backs the claim that the data is not copied into the model.

Three possible topologies

Layer in the customer's environment, model via API. Covers most enterprise cases. Data and index stay inside; inference happens outside, with only the snippet sent out. It combines control over the data with access to the most capable models, without operating inference infrastructure.

Layer and model both in the customer's environment. Nothing leaves the perimeter. Requires operating GPUs, updating model versions, and owning inference availability.

Layer at the customer, model in a dedicated region. A middle ground for geographic restrictions, with processing in a defined region and a specific contract.

The choice depends on the applicable regulation and how critical the data is, and it does not have to be uniform. It is common to run critical tasks under one design and the rest under another.

What "only the necessary snippet" means

When someone asks about a specific customer's situation, the layer retrieves the set of information related to that account and sends only that to the model. The entire customer base stays where it was.

This has two effects at once: it reduces the exposure surface on each call, and it reduces cost, because processing less content means consuming fewer tokens. The opposite design — sending broad context out of caution — increases both risk and cost at the same time.

How to reduce what leaves even further

  • Precise slicing. Send the snippet that answers the question, not the whole document.
  • Pseudonymization. Replace direct identifiers with references when the task does not require identification, remapping names back into the answer inside the perimeter.
  • Prior classification. Tag fields by sensitivity and block categories that should never leave.
  • Policy by task type. Formatting and summarization usually need less identifiable data than it seems.

Applied example

A mid-sized financial institution had a project stalled for seven months in security review. The original proposal called for replicating the customer base into the vendor's environment.

The redesign kept the base where it was, inside the institution's private cloud, with the context layer running in the same environment and each user's permission applied at retrieval. For the model provider, only the context snippet for each query started traveling out, with identifiers pseudonymized when the task did not require identification.

Approval came through in six weeks, after a demo that showed, in a network capture, exactly what was sent on a real query. The discussion moved from promises on a slide to observable evidence.

What to put in the model provider's contract

  • No use of submitted content for training.
  • Content retention period, with a preference for zero retention where available.
  • Processing region.
  • Subprocessors and change-notification policy.
  • Certifications and audit rights.

It is worth requiring this in writing, because a provider's public policy can differ from the plan actually contracted.

Operational requirements of running inside the environment

Choosing to keep the layer inside your own environment means taking on responsibilities that are worth making clear from the start:

  • availability and scaling of the layer, treated as a production service;
  • dependency updates and vulnerability patching;
  • monitoring retrieval and latency;
  • secret management and credential rotation for the sources;
  • defined behavior when the model provider is unavailable, with a clear message instead of a partial answer.

The last item is the most commonly forgotten, and the one that causes the most operational incidents.

Strattum's Company Brain runs inside the customer's infrastructure for exactly this reason: the data stays in the environment, and only the necessary snippet travels on each query.

Frequently asked questions

Is running the model locally always safer?

It eliminates external traffic and shifts to the company the responsibility of operating, updating and securing the inference infrastructure. It is the right choice when regulation requires it, not by default.

Does pseudonymization always work?

It works when the task does not depend on identity. In analysis that requires a name or an ID number, the control comes from the topology and the contract instead.

What about data that already went through during testing?

It is worth surveying what was sent during the experimentation phase and checking the provider's retention policy for that period.

How do you demonstrate this in an audit?

With the flow map by query type, the retrieval log, and a capture showing the content actually sent in real calls.

Design the architecture that passes review.

Talk to a Strattum expert about where your data lives and what travels to the model.