TL;DR: In a well-designed context architecture, data stays inside the customer's own infrastructure — private cloud or on-premises — and only the snippet needed to produce the answer travels to the model provider on each query. The data is not copied into the model, nor used to train it. Mapping that flow explicitly is what gets a project through security review.
It is not the model's capability. It is the inability to answer five questions precisely:
Teams that walk into the security review with a clear answer to all five move faster than teams that walk in with an impressive demo.
| Element | Where it lives | Leaves the perimeter? |
|---|---|---|
| Source database | Customer environment | No |
| Search index and metadata | Customer environment | No |
| Query and audit log | Customer environment | No |
| Retrieved context snippet | Retrieved inside the environment | Yes, at call time |
| User's question | Customer environment | Yes, at call time |
| Generated answer | Provider | Returns to the environment |
What leaves on each query is the retrieved context snippet plus the user's question, never the underlying base. That design is what backs the claim that the data is not copied into the model.
Layer in the customer's environment, model via API. Covers most enterprise cases. Data and index stay inside; inference happens outside, with only the snippet sent out. It combines control over the data with access to the most capable models, without operating inference infrastructure.
Layer and model both in the customer's environment. Nothing leaves the perimeter. Requires operating GPUs, updating model versions, and owning inference availability.
Layer at the customer, model in a dedicated region. A middle ground for geographic restrictions, with processing in a defined region and a specific contract.
The choice depends on the applicable regulation and how critical the data is, and it does not have to be uniform. It is common to run critical tasks under one design and the rest under another.
When someone asks about a specific customer's situation, the layer retrieves the set of information related to that account and sends only that to the model. The entire customer base stays where it was.
This has two effects at once: it reduces the exposure surface on each call, and it reduces cost, because processing less content means consuming fewer tokens. The opposite design — sending broad context out of caution — increases both risk and cost at the same time.
A mid-sized financial institution had a project stalled for seven months in security review. The original proposal called for replicating the customer base into the vendor's environment.
The redesign kept the base where it was, inside the institution's private cloud, with the context layer running in the same environment and each user's permission applied at retrieval. For the model provider, only the context snippet for each query started traveling out, with identifiers pseudonymized when the task did not require identification.
Approval came through in six weeks, after a demo that showed, in a network capture, exactly what was sent on a real query. The discussion moved from promises on a slide to observable evidence.
It is worth requiring this in writing, because a provider's public policy can differ from the plan actually contracted.
Choosing to keep the layer inside your own environment means taking on responsibilities that are worth making clear from the start:
The last item is the most commonly forgotten, and the one that causes the most operational incidents.
Strattum's Company Brain runs inside the customer's infrastructure for exactly this reason: the data stays in the environment, and only the necessary snippet travels on each query.
It eliminates external traffic and shifts to the company the responsibility of operating, updating and securing the inference infrastructure. It is the right choice when regulation requires it, not by default.
It works when the task does not depend on identity. In analysis that requires a name or an ID number, the control comes from the topology and the contract instead.
It is worth surveying what was sent during the experimentation phase and checking the provider's retention policy for that period.
With the flow map by query type, the retrieval log, and a capture showing the content actually sent in real calls.
Talk to a Strattum expert about where your data lives and what travels to the model.