Prompt engineering answered the first question anyone asks when they start using a model: how to ask.
As agents move to the centre of how products are built, the question that decides the outcome is a different one. What information to put in front of the model, and when.
That work is context engineering.
The difference shows up once an agent leaves the prototype. Two teams can start from the same model, the same framework and the same document base and land far apart. What separates them is what each one decided to put in the context window of each call.
These are engineering decisions and can be treated as such. They have criteria, they have a cost, and they can be measured.
Retrieval-Augmented Generation was a leap forward. It gave the model a way to reach information that was not in its weights, and it solved much of the access problem.
Context engineering asks a harder question. Given a limited context window, what is the best possible combination of information for the model to execute this task, right now?
Retrieving relevant passages is one of the possible answers. It shares the window with everything else that has to be there.
In an agent running in production, a single call usually accumulates:
Each of these takes space the others no longer have. The task is to choose the right set, in the right order, at the right time.
Filling the window is the easiest decision to make and the most expensive one to maintain.
Before any retrieval happens, the model needs to know what is available.
Routing is already a context decision. Settling whether a question about revenue goes to the data warehouse, to the CRM or to the finance policy sitting in a document determines everything that follows it.
When that choice is left to the model with no description of the sources, it guesses. And it guesses again on the next run, because nothing that worked the first time was written down anywhere.
Describing which sources exist, what each one holds and when each one is the right one is context work done before the first query.
Context windows are finite, and they grow more slowly than the volume of information a company would like to put inside them.
Summarizing retrieved content, ranking results by relevance or by date, cutting redundant passages. On their own these are small choices, and their effect compounds across a full run.
Order matters more than it tends to look. The same set of ten documents produces different answers depending on which one comes first.
The same holds for anything out of date. Four versions of one policy in the same window force the model to choose between them, and it will choose with whatever criteria it has at hand.
In conversational agents, what gets stored between sessions weighs as much as what is in the current window.
That means deciding what deserves to be remembered. A correction from the user, a stated preference, an exception that keeps coming back: each of those signals can become durable context or can be lost when the conversation ends.
And it means deciding what is retrieved, and when. Memory that returns in full on every call stops working as memory and starts spending tokens without returning signal.
An agent that stores and retrieves well improves with use. It is the same mechanism that makes a new hire more productive in their second month than in their first.
A good share of the gain comes from breaking the work up.
Well-designed workflows split a complex task into steps, and each step carries the lean context it needs. Every call has less noise in it, and an error stays contained in the step where it happened.
It also makes the system observable. When an answer comes out wrong from one monolithic call with everything crammed in, it is hard to tell which piece of context caused it. In steps, you can point at it.
Systems built this way tend to be faster, cheaper to run and more predictable when something changes.
All of these decisions come back with every new agent.
When each project solves source selection, compression, ordering and memory on its own, the second agent redoes the first one’s work with different criteria. The company ends up with two versions of what is true about itself, and neither knows about the other.
That is why part of these decisions belongs to a shared layer. Which sources exist, what each one means, which version of a document is in force, who may see what: that is company context, and it can be built once.
Context engineering is the work of deciding what reaches the model. Having that work done once and reusable is what changes the economics of putting the next agent into production.
Schedule a 30-minute conversation. We show Strattum running against systems like yours.