Announcing US$ 3.2M pre-seed round · OneVC · Maya · Norte Ventures Read →
Back to blog

Why your company's AI nails the easy question
and misses the one that matters

TL;DR: Generic tasks like summarizing, editing and rewriting work well because all the information they need is inside the request. Questions about a customer, margin, risk or priority depend on internal data and business rules, and that is where quality drops. Three signals separate a lack of context from a model limitation: a generic answer, an ignored internal rule, and inconsistency across repeated queries.

Two task classes with opposite results

The tasks that work well from day one share one trait: all the information needed is inside the request. An email to edit, minutes to summarize, a text to translate. The model gets the full material and works on it.

The tasks that disappoint share the opposite trait. Asking which accounts to prioritize, whether to approve a discount, or where the biggest margin loss is happening requires information that does not fit in the request and is scattered across different systems.

The frustration shows up because the same tool displays both behaviors. Someone who tested the assistant with a summary and was impressed does not understand why the question about the pipeline came back shallow.

What exactly is missing

The data. Revenue, ticket status, product usage and contract history sit in systems the application does not query.

The definition. Active customer, at-risk account, canceled order and recurring revenue carry a specific meaning at each company. Without that declared, the model falls back on the most common market definition.

The hierarchy. When two sources disagree, someone needs to have decided which one wins for which kind of decision. That rule usually lives in the operator's head and nowhere machine-readable.

Signal 1: the answer would fit any company

The text talks about tracking metrics, prioritizing strategic customers and monitoring satisfaction. It is correct and unusable, because it names no customer, no number and no decision.

This happens when the application has no access to the relevant internal sources. The model does what it can with what it gets, and what it gets is the question and little else.

How to confirm it: ask for the same analysis while naming a specific customer. If the answer brings back no real data about that account, the problem is source access.

Signal 2: an internal rule was ignored

The answer suggests an action anyone on the team would rule out in a second, like offering a promotion to an account with blocked credit or applying a discount above the requester's authority.

The data may even be available. What is missing is the rule that says what to do with it, and it usually lives in the operator's memory or inside an old report.

How to confirm it: ask the team what criterion they would have applied. If the criterion exists and was never written anywhere the application can read, the problem is a missing business definition.

Signal 3: the same question returns different answers

Asked twice, the same question brings back different sets of accounts, numbers or recommendations. Some wording variation is expected in text generation. Variation in the underlying data cited is a different matter.

This indicates the slice of information retrieved changes between queries, usually because of a fragile search strategy or unstable ranking of results.

How to confirm it: repeat the same query five times and compare the cited data, not the wording. If the data changes, the problem is context retrieval.

What to do with the diagnosis

SignalWhere to investWhat will not fix it
Generic answerConnection to the relevant sourcesA bigger model, a more detailed prompt
Ignored internal ruleDeclaring the business definitionsA bigger model
Inconsistency across queriesSearch and modeling strategyAdjusting temperature
Confusing or poorly structured textPrompt or model choiceConnecting more sources

Only the last row justifies touching the model. The first three are context work, and the budget should go to data architecture.

Applied example

A logistics operation with 200 customers asked its internal assistant for two things on the same day. The first was to summarize the week's operations meeting notes, and the result was good, because the notes were attached to the request. The second was to list the customers with the highest churn risk for the quarter, and the answer came back generic, naming no one. There was no way to name anyone: volume lived in the TMS, complaints lived in the incident system, and the assistant did not talk to either.

A healthcare provider took the diagnosis further. After six months of underwhelming results, it was about to switch model providers. The team ran the three tests in one afternoon. A generic answer on every eligibility question, because the plan's rule base was not connected. An ignored rule on waiting periods, documented only in a PDF outside the index. Inconsistency on the provider network, where search returned different providers on each attempt.

None of the three had anything to do with the model. The switch was called off and the budget was redirected to connecting the sources and declaring the rules.

Where to start

Pick the business question that comes up most often in your team's meetings and work backwards: list which systems hold each piece of the answer, which internal definitions it requires, and who decides when two sources disagree. That map usually has between three and six items, and it is what determines whether the AI will be able to answer.

When all three signals show up together, the investment that changes the outcome is context — which is exactly what a Company Brain structures.

Frequently asked questions

Does prompt engineering fix it?

It improves the instruction and the shape of the answer. It does not fetch information the application has no access to, so it does not fix any of the three signals.

What if all three show up together?

That is the most common scenario, and it indicates the application was built without a context layer. It is worth treating as an architecture problem, not a sequence of one-off fixes.

Is there a signal that actually points to the model?

Yes: poorly structured text, a complex instruction ignored consistently, and a reasoning error with all the material already at hand.

Who should run the tests?

Someone from the business side together with someone technical. The first recognizes the ignored rule, the second identifies where the inconsistency comes from.

Diagnose your AI project before the next investment.

Talk to a Strattum expert and apply the three signals to your case.