Announcing US$ 3.2M pre-seed round · OneVC · Maya · Norte Ventures Read →
Back to blog

The sources left out of the pipeline
that carry the business rule

TL;DR: Contracts, internal policies, team spreadsheets and support history are usually left out of the analytics pipeline because they are not tabular and nobody has modeled them. That is exactly where the business rule that never became a table lives — contract exceptions, escalation criteria, discount policies. An AI application that ignores these sources is limited to what was already structured.

Why these sources get left out

Each for a different reason.

Contracts. They arrive as PDFs, many of them scanned, with specific clauses negotiated case by case. The management system stores metadata, not the clause.

Internal policies. They live on the intranet, a shared drive, or an email attachment. They have versions, and the version in force is not always the most recent one someone finds.

Team spreadsheets. They hold the rule the team actually applies, created because the system did not anticipate that case. Nobody treats them as an official source, and everyone uses them.

Support history. High-volume free text, with the explanation of why each exception was granted.

What they have in common is that they contain decisions, not just records.

What gets lost by ignoring them

A question like "can I grant this discount to this customer?" depends on three layers: the order data, which is in the ERP; the policy in force, which is in a document; and the contractual exception negotiated with that customer, which is in the contract.

With only the first layer, the answer looks complete and is wrong in every case that involves an exception. And an exception is usually exactly why the question was asked.

How to bring in each type

SourceTechnical challengeApproach
PDF contractScanning, variable layout, versioningOCR when needed, clause-level extraction, linked to the customer in the CRM
Internal policyVersioning and validity periodIndexing with an effective date and current-version flag
Team spreadsheetUnstable structure, informal ownerFormalize an owner, stabilize the columns, ingest with validation
Support historyHigh volume, noisy textSemantic indexing filtered by account and period

The link to the entity is the step that turns a loose document into usable context. A contract indexed without a link to the customer answers questions about the text, not questions about that account.

The spreadsheet case

Team spreadsheets tend to make data teams uncomfortable, because they represent a process running outside the system. Ignoring it does not make the process go away — it just keeps the rule out of the AI's reach.

The practical path has three steps: identify which spreadsheets support recurring decisions, assign an owner to each one, and stabilize the minimum structure needed for ingestion. In many cases, this also speeds up eventually migrating that rule into a system.

Applied example

A medical equipment company had a sales assistant connected to ERP and CRM. Its answers about commercial terms were right in most cases and wrong for large customers.

The reason turned out to be a spreadsheet maintained by the contracts team, holding the special terms negotiated with 40 strategic customers, including different payment terms and their own price table. The spreadsheet was updated by one person and consulted by email.

After formalizing the owner, stabilizing the columns, and ingesting the content linked to the customer code in the ERP, the assistant started answering correctly for the group that accounted for most of the revenue.

How to prioritize which ones to bring in first

Three criteria, applied together:

Frequency of use in decisions. How many times a week someone needs to check that material to answer something.

Impact of getting it wrong. What happens when that source's information is ignored in an answer.

Ingestion effort. A structured spreadsheet can be brought in within days; a scanned contract with a variable layout takes weeks.

The ideal first batch combines high frequency, high impact and low effort. It exists in almost every company, and it is usually exactly the spreadsheet nobody wanted to own.

The side benefit of bringing these sources in

Formalizing ownership, stabilizing structure and declaring validity periods improves the process even with no AI involved.

The spreadsheet that supported decisions with no owner now has one. The policy with three versions floating around now has one identified as current. The contract with a specific clause no longer depends on someone remembering it exists.

Teams that start with these sources tend to deliver two results at once: context the AI can use, and document governance the company had been putting off for years.

Bringing in these sources is part of what distinguishes a Company Brain from an analytics pipeline, because they hold a good part of the rule that governs decisions.

Frequently asked questions

Isn't this data governance work?

It is, and it becomes urgent once AI enters the picture, because the absence of these sources starts producing wrong answers instead of just an incomplete report.

Does OCR solve scanned documents?

It solves text extraction. The next job is structuring by clause and linking to the corresponding entity, which is where the value is.

Where should we start?

With the sources that support the most frequent exceptions. There are usually only a few of them, and they explain most of the errors in existing applications.

Bring those sources inside the context.

Talk to a Strattum expert about how contracts, policy and spreadsheets enter the Company Brain.