Announcing US$ 3.2M pre-seed round · OneVC · Maya · Norte Ventures Read →
Back to blog

Data that never becomes training:
what changes in integration design

TL;DR: When data is not copied into the model or used for training, the integration has to retrieve the information at query time. That changes latency requirements, requires an explicit caching policy, and forces volatile data to be handled differently from stable data. In exchange, it removes the parallel sync that always ends up out of date.

Two ways to bring knowledge into the model

Embed it. Train or fine-tune the model on company data. The knowledge sits in the weights, and updating it requires retraining.

Retrieve it. Fetch the information on every query and hand it over as context. The knowledge stays at the source, and the update is immediate.

For corporate data that changes, the second approach is what sustains a correct answer. Balance, order status, an open ticket and stock position change within hours.

What this requires in the design

A latency budget. The query now includes retrieval before inference. It is worth declaring a target per question type and measuring each stage separately.

A caching policy per data type. Internal policy and documentation change little and allow long caching. Balance and status do not. Invalidation needs to be explicit, not inherited from a single default.

Handling volatile data. For information that has to be current to the second, the layer queries the source directly instead of using an index, trading higher latency for accuracy.

Source availability. If the ERP goes down, the application needs to report the outage instead of answering with whatever was left in the index.

The synchronization that disappears

A copy-based design creates a second place where the data lives, with its own pipeline, refresh window and monitoring. That second place tends to drift from the source at some point, and the drift shows up as a wrong answer with no warning.

Retrieving at the source removes that entire class of problem. What remains is the search index for text content, which still needs an update strategy, but with a much smaller scope than a full replica.

AspectCopy into the modelRetrieval at query time
FreshnessFrozen at training timeCurrent state
Per-person permissionHard to applyApplicable at retrieval
Update costRetrainingNone
Data exposureInside the weightsA snippet in the query
LatencyLowerHigher, with a retrieval step

Applied example

A retailer tried the fine-tuning route with its product catalog and promotion rules, aiming for fast answers for the store team.

The fine-tuned model answered well about the catalog from the training month. Promotions changed every week, and within 20 days the answers already diverged from the system. Every update required a new training and validation cycle.

The redesign kept the base model and started retrieving the current catalog and rule at query time, with a five-minute cache for the catalog and a direct query for price and stock. Response time went up by roughly 400 milliseconds, and the divergence from the system disappeared.

Volatility classification

The caching decision gets simpler when each source is given a declared volatility band at modeling time:

BandExamplesStrategy
SecondsBalance, stock, order statusDirect query to the source, no cache
MinutesSupport queue, delivery positionShort cache with event-based invalidation
DaysCustomer record, catalog, commercial termsCache with change-based invalidation
MonthsInternal policy, manual, signed contractIndexing with version control
ImmutableClosed historical periodsFull materialization

Without this classification, teams tend to adopt a single TTL that is too short for stable content and too long for critical data — unsatisfying on both ends.

Where caching actually helps

  • Policy documents and manuals, with version-based invalidation.
  • Metadata structures and the semantic dictionary, which change rarely.
  • Results of repeated queries within short windows, with a declared expiry.
  • Embeddings of stable content, recalculated only when the document changes.

The criterion is the data's rate of change, not how often the question is asked.

What to do when the source is unavailable

Retrieving at the source creates an operational dependency that needs defined behavior. Three options, in order of preference:

Report the outage. The answer states that a given source did not respond and what was left out.

Answer with a flagged cache. Use the last known value, with the date made explicit in the answer.

Decline the query. For critical questions where a partial answer is worse than no answer.

What should never happen is the silent path, where the application answers with whatever it managed to get without flagging the gap.

What this changes in the provider contract

When data is not used for training, three clauses need to be in writing: no use of submitted content for training, a content retention period with a preference for zero retention, and the processing region.

It is worth confirming what applies to the contracted plan, since a provider's public policy can differ from what applies to a given account type. This check is usually requested for security review and is quick to get in writing.

A Company Brain works this way by design: data stays at the source and is retrieved at query time, with no copy into the model.

Frequently asked questions

Does fine-tuning become useless, then?

No. It remains useful for format, tone and behavior on repetitive tasks. What it does not solve is factual knowledge that changes.

Is the extra latency noticeable?

In an interactive query, it usually falls somewhere between a few hundred milliseconds and a few seconds, depending on the source. Good scoping offsets part of it, since it reduces inference time.

What about historical data that never changes?

It can be materialized and indexed without concern, updated only as new periods come in.

Design the integration without copying data.

Talk to a Strattum expert about latency, caching and data volatility in your AI architecture.