Announcing US$ 3.2M pre-seed round · OneVC · Maya · Norte Ventures Read →
Back to blog

Switching models rarely fixes it:
what changes and what stays the same

TL;DR: Switching to a more capable model improves writing, reasoning structure, instruction-following and coding performance. It does not fix an error caused by missing information, because the model still answers from the context it receives. When the answer is wrong for lack of data or a business rule, the variable to change is the input.

What the switch actually improves

It's worth acknowledging the real gains, because they exist:

  • clearer, better-organized writing
  • stronger adherence to long instructions and requested formats
  • better performance on multi-step reasoning
  • fewer errors in code generation and review
  • in many cases, more speed and a larger context window

If your project's problem is one of these, switching is the right call.

What stays exactly the same

The set of information that reaches the model. If the application only queries the CRM, the new model will also answer using only the CRM. If the definition of an at-risk account was never declared, it will use the market's generic definition, just more eloquently.

There is a side effect here. Better argumentation raises the reader's confidence. A wrong recommendation presented with a solid structure gets through more reviews unquestioned than the same recommendation written in a confusing way.

How to know which case is yours in ten minutes

Take a bad answer that triggered the discussion about switching models and run the full-context test: assemble by hand, with no automation, the set of information an experienced person would use to answer that question. Paste all of it into the current model and ask the question again.

If the answer turns out correct, the problem is context, and no switch will fix it. If it is still bad even with all the material in hand, then the model really is the limiting factor.

This test costs the time it takes to assemble the material, and it saves months of a migration that would not have changed the outcome.

Applied example

A credit fintech used an agent for the initial triage of credit-limit increase requests. The recommendations came out inconsistent, and the team decided to migrate to the most expensive model available.

The migration took six weeks between prompt tuning, testing and cost review. The answers came out better written, and the rate of inadequate recommendations barely moved.

The full-context test, run afterward, showed the problem: the application was sending 90 days of payment history, while internal policy looked at 12 months, and also required a restriction check against a database that was not connected. With the full material, the original model got the triage right.

The fix was a data architecture fix, and it took less time than the migration that hadn't solved anything.

The hidden cost of a migration

Switching the model of an application already in production is not a one-line config change. The real work usually includes:

  • rewriting and retesting the prompts, because each model family responds differently to long instructions;
  • adjusting the output format the rest of the system consumes;
  • a new round of quality evaluation, if one exists;
  • reviewing cost and rate limits with the new provider;
  • in some cases, a new security and legal sign-off.

Six to ten weeks is a common range for this cycle in an application with real users. When the root cause was context, that entire time is spent without changing the outcome that triggered the decision.

When switching pays off

SituationSwitch modelWork on context
Correct answer, poor writingYesNo
Complex instruction ignoredYesMaybe
Generic answer about the businessNoYes
Internal rule violatedNoYes
High cost with acceptable qualityYes, to a smaller modelYes, reduces volume sent

That last row is the most interesting one for whoever has already organized their context, because the switch now runs the other way — from a bigger model to a smaller one.

The order that saves time

When an application is underperforming, the sequence that tends to resolve it fastest is this:

  1. Full-context test, to separate the two possible causes.
  2. Fix whatever was missing — a source connection, a rule definition, or a retrieval strategy.
  3. Re-evaluate with the current model, because in a good share of cases the result is already acceptable here.
  4. Compare models, with your own evaluation set and frozen context.
  5. Decide per task type, keeping different models where that makes sense.

Teams that start at step 4 spend the migration budget before knowing whether it changes anything.

The case where switching does improve context

There is one exception worth knowing. Models with a larger window let you send more material without cutting it, which can improve answers on tasks with long documents.

That fixes the symptom and keeps the cost high, because the bill tracks the volume sent on every call. It is worth it as a temporary fix while retrieval is redesigned, not as the final design.

With context organized in a Company Brain, the conversation about models tends to change direction, because it becomes possible to downgrade instead of upgrade.

Frequently asked questions

Should you always use the newest model?

It is worth testing, with your own evaluation set. The best model on public benchmarks is not necessarily the best one for your tasks.

How do you build that evaluation set?

Gather 50 to 100 real queries from the operation, including the hard cases, with the expected answer when possible. Run the same set against each candidate, with the same context.

Can you use different models per task?

Yes, and it tends to be the most economical design. Simple tasks on a smaller model, critical tasks on the larger one, with the same context layer serving both.

Find out if the problem is the model or the context.

Talk to a Strattum expert and run the full-context test on your application before deciding on the next migration.