TL;DR: Contracts, internal policies, team spreadsheets and support history are usually left out of the analytics pipeline because they are not tabular and nobody has modeled them. That is exactly where the business rule that never became a table lives — contract exceptions, escalation criteria, discount policies. An AI application that ignores these sources is limited to what was already structured.
Each for a different reason.
Contracts. They arrive as PDFs, many of them scanned, with specific clauses negotiated case by case. The management system stores metadata, not the clause.
Internal policies. They live on the intranet, a shared drive, or an email attachment. They have versions, and the version in force is not always the most recent one someone finds.
Team spreadsheets. They hold the rule the team actually applies, created because the system did not anticipate that case. Nobody treats them as an official source, and everyone uses them.
Support history. High-volume free text, with the explanation of why each exception was granted.
What they have in common is that they contain decisions, not just records.
A question like "can I grant this discount to this customer?" depends on three layers: the order data, which is in the ERP; the policy in force, which is in a document; and the contractual exception negotiated with that customer, which is in the contract.
With only the first layer, the answer looks complete and is wrong in every case that involves an exception. And an exception is usually exactly why the question was asked.
| Source | Technical challenge | Approach |
|---|---|---|
| PDF contract | Scanning, variable layout, versioning | OCR when needed, clause-level extraction, linked to the customer in the CRM |
| Internal policy | Versioning and validity period | Indexing with an effective date and current-version flag |
| Team spreadsheet | Unstable structure, informal owner | Formalize an owner, stabilize the columns, ingest with validation |
| Support history | High volume, noisy text | Semantic indexing filtered by account and period |
The link to the entity is the step that turns a loose document into usable context. A contract indexed without a link to the customer answers questions about the text, not questions about that account.
Turning a document into usable context takes three operations beyond text extraction.
Identify the entity involved. Which customer, contract, project or policy the document refers to, and convert that into the key used across the systems.
Establish the validity period. Since when the condition holds, whether it replaced an earlier one, and whether it has already expired. A document with no validity period produces answers based on a revoked rule.
Preserve the permission. Whoever could see the file at the source has to remain the criterion after indexing.
All three are usually treated as implementation detail, and they are what decides whether the source becomes context or becomes noise.
Team spreadsheets tend to make data teams uncomfortable, because they represent a process running outside the system. Ignoring it does not make the process go away — it just keeps the rule out of the AI's reach.
The practical path has three steps: identify which spreadsheets support recurring decisions, assign an owner to each one, and stabilize the minimum structure needed for ingestion. In many cases, this also speeds up eventually migrating that rule into a system.
A medical equipment company had a sales assistant connected to ERP and CRM. Its answers about commercial terms were right in most cases and wrong for large customers.
The reason turned out to be a spreadsheet maintained by the contracts team, holding the special terms negotiated with 40 strategic customers, including different payment terms and their own price table. The spreadsheet was updated by one person and consulted by email.
After formalizing the owner, stabilizing the columns, and ingesting the content linked to the customer code in the ERP, the assistant started answering correctly for the group that accounted for most of the revenue.
Three criteria, applied together:
Frequency of use in decisions. How many times a week someone needs to check that material to answer something.
Impact of getting it wrong. What happens when that source's information is ignored in an answer.
Ingestion effort. A structured spreadsheet can be brought in within days; a scanned contract with a variable layout takes weeks.
The ideal first batch combines high frequency, high impact and low effort. It exists in almost every company, and it is usually exactly the spreadsheet nobody wanted to own.
Formalizing ownership, stabilizing structure and declaring validity periods improves the process even with no AI involved.
The spreadsheet that supported decisions with no owner now has one. The policy with three versions floating around now has one identified as current. The contract with a specific clause no longer depends on someone remembering it exists.
Teams that start with these sources tend to deliver two results at once: context the AI can use, and document governance the company had been putting off for years.
Bringing in these sources is part of what distinguishes a Company Brain from an analytics pipeline, because they hold a good part of the rule that governs decisions.
It is, and it becomes urgent once AI enters the picture, because the absence of these sources starts producing wrong answers instead of just an incomplete report.
It solves text extraction. The next job is structuring by clause and linking to the corresponding entity, which is where the value is.
With the sources that support the most frequent exceptions. There are usually only a few of them, and they explain most of the errors in existing applications.
Talk to a Strattum expert about how contracts, policy and spreadsheets enter the Company Brain.