TL;DR: Source citation and the audit trail need to be produced by the component that retrieves the data, not by the interface that displays the answer. Controls implemented in the presentation layer stop holding the moment agents, automations and integrations start consuming the platform over an API — which is exactly when query volume grows.
The first application usually has its own interface. There, it is easy to put together a side panel with the sources and log the query in that application's own backend.
Six months later, three new consumers show up: an agent running in batch, an automation triggered by a webhook, and an integration with the support tool. None of them go through the interface.
If citation and logging lived there, those three consumers now operate with no traceability. And they tend to generate more queries than the original interface did.
Every retrieval call should return, along with the content:
With that envelope, any consumer displays or logs whatever it needs. The interface decides how to present it and is not responsible for generating the information.
Beyond what goes back to the consumer, the layer logs the full query:
| Field | Use |
|---|---|
| User or agent identity | Investigation and permission review |
| Question or parameters received | Telling legitimate use apart from a workaround |
| Retrieved snippets | Reconstructing what fed the answer |
| Permission filters applied | Proving the restriction was enforced |
| Model and version, when reported | Explaining a change in behavior |
| Latency and volume sent | Cost and performance diagnostics |
The last two fields also serve operations, which helps justify the instrumentation effort to teams that see audit as pure cost.
A well-designed retrieval response separates content from provenance, so the consumer decides what to do with each part:
{
"snippets": [
{
"content": "...",
"source": {
"system": "erp",
"record": "invoice:2026-0041982",
"as_of": "2026-08-31",
"version": null
},
"retrieved_via": "structured_query",
"filters_applied": ["permission:user_1187", "customer:8842"]
}
],
"query_id": "c_9f31a2",
"sources_queried": ["erp", "crm"],
"sources_unavailable": []
}The unavailable-sources field deserves attention. Without it, a source that is down produces a silently incomplete answer — the worst possible failure mode in an enterprise setting.
Three technical reasons.
The envelope changes the contract. Adding source metadata changes the API response and forces every existing consumer to adjust.
History cannot be reconstructed. Past queries do not generate a retroactive log, so the audit trail starts from zero on the date of the change.
Retroactive permission is even worse. If the filter was never applied at retrieval, there is no record of who reached what, and the impact assessment has to be done by inference.
A logistics company built its first assistant with a sources panel on the web interface. It worked well, and the security team signed off on it.
Six months later, an automated agent started generating daily reports by consuming the same base over the API, and the support operation integrated the platform with the ticketing system.
At the next audit, the company could demonstrate traceability for 18% of queries — the slice that originated in the interface. Moving citation and logging into the retrieval layer took two months and required versioning the API so as not to break existing consumers.
The same record that serves audit also solves the most common operational problem: investigating a bad answer.
With the retrieved snippets kept, the analysis takes minutes: either the right material was not retrieved, and the problem is in the search, or it was there and the model did not use it, and the problem is in the prompt or the model.
Without that record, the investigation turns into trying to reproduce the query — with the added complication that the state of the sources has already changed.
The order that causes the least disruption has four steps.
Step 2 is what pays off fastest, because the record starts existing without depending on anyone's migration.
That is why, in a Company Brain, source citation and logging are born in retrieval, not on screen.
It couples them on purpose. That coupling stops every consumer from implementing its own partial version of the control.
The envelope is for the consumer. What goes to the model can be just the content, with the references kept on the application side to compose the final answer.
It is worth logging all of them, with a retention policy that varies by criticality, rather than deciding case by case what to log.
Talk to a Strattum expert about where your platform generates source and log today, and what changes when that moves into retrieval.