TL;DR: When data is not copied into the model or used for training, the integration has to retrieve the information at query time. That changes latency requirements, requires an explicit caching policy, and forces volatile data to be handled differently from stable data. In exchange, it removes the parallel sync that always ends up out of date.
Embed it. Train or fine-tune the model on company data. The knowledge sits in the weights, and updating it requires retraining.
Retrieve it. Fetch the information on every query and hand it over as context. The knowledge stays at the source, and the update is immediate.
For corporate data that changes, the second approach is what sustains a correct answer. Balance, order status, an open ticket and stock position change within hours.
A latency budget. The query now includes retrieval before inference. It is worth declaring a target per question type and measuring each stage separately.
A caching policy per data type. Internal policy and documentation change little and allow long caching. Balance and status do not. Invalidation needs to be explicit, not inherited from a single default.
Handling volatile data. For information that has to be current to the second, the layer queries the source directly instead of using an index, trading higher latency for accuracy.
Source availability. If the ERP goes down, the application needs to report the outage instead of answering with whatever was left in the index.
A copy-based design creates a second place where the data lives, with its own pipeline, refresh window and monitoring. That second place tends to drift from the source at some point, and the drift shows up as a wrong answer with no warning.
Retrieving at the source removes that entire class of problem. What remains is the search index for text content, which still needs an update strategy, but with a much smaller scope than a full replica.
| Aspect | Copy into the model | Retrieval at query time |
|---|---|---|
| Freshness | Frozen at training time | Current state |
| Per-person permission | Hard to apply | Applicable at retrieval |
| Update cost | Retraining | None |
| Data exposure | Inside the weights | A snippet in the query |
| Latency | Lower | Higher, with a retrieval step |
A retailer tried the fine-tuning route with its product catalog and promotion rules, aiming for fast answers for the store team.
The fine-tuned model answered well about the catalog from the training month. Promotions changed every week, and within 20 days the answers already diverged from the system. Every update required a new training and validation cycle.
The redesign kept the base model and started retrieving the current catalog and rule at query time, with a five-minute cache for the catalog and a direct query for price and stock. Response time went up by roughly 400 milliseconds, and the divergence from the system disappeared.
The caching decision gets simpler when each source is given a declared volatility band at modeling time:
| Band | Examples | Strategy |
|---|---|---|
| Seconds | Balance, stock, order status | Direct query to the source, no cache |
| Minutes | Support queue, delivery position | Short cache with event-based invalidation |
| Days | Customer record, catalog, commercial terms | Cache with change-based invalidation |
| Months | Internal policy, manual, signed contract | Indexing with version control |
| Immutable | Closed historical periods | Full materialization |
Without this classification, teams tend to adopt a single TTL that is too short for stable content and too long for critical data — unsatisfying on both ends.
The criterion is the data's rate of change, not how often the question is asked.
When data is not used for training, three clauses need to be in writing: no use of submitted content for training, a content retention period with a preference for zero retention, and the processing region.
It is worth confirming what applies to the contracted plan, since a provider's public policy can differ from what applies to a given account type. This check is usually requested for security review and is quick to get in writing.
A Company Brain works this way by design: data stays at the source and is retrieved at query time, with no copy into the model.
No. It remains useful for format, tone and behavior on repetitive tasks. What it does not solve is factual knowledge that changes.
In an interactive query, it usually falls somewhere between a few hundred milliseconds and a few seconds, depending on the source. Good scoping offsets part of it, since it reduces inference time.
It can be materialized and indexed without concern, updated only as new periods come in.
Talk to a Strattum expert about latency, caching and data volatility in your AI architecture.