Announcing US$ 3.2M pre-seed round · OneVC · Maya · Norte Ventures Read →
Back to blog

Business ontology:
how AI understands your company's vocabulary

The same company usually exists three times inside the business: as a legal entity in the ERP, as an account with a trade name in the CRM, and as an organization identified by an email domain in the helpdesk. For the people who work there, all three records are the same account. For a language model, they are three different customers, and the answer comes out with the wrong total. The model arrives fluent in the language and with no notion of that operation's vocabulary. What fills that gap is a business ontology, the layer that says what each term and each entity means in this specific company. It is one of three components of a Company Brain, the layer that knows the company: what data exists, where it lives, how it relates, and who is allowed to see what.

The business ontology defines the vocabulary before the question

It's worth separating the three pieces. The knowledge graph models entities and the links between them. The semantic layer translates the question into a query with the correct permission. Between the two sits the business ontology, which resolves meaning.

Ask three teams what a qualified lead is. Marketing counts whoever filled out the form and matches the profile. Pre-sales counts whoever accepted the meeting. Sales counts whoever has an approved budget. Without the ontology recorded somewhere, the model silently picks one of the three definitions, and the choice changes with every question.

Entity resolution decides when two records become one

Unifying records across different systems is the part that eats up the most engineering time, and it's where the architecture decision shows up most clearly.

The simplest approach compares names and merges whatever looks alike. It breaks early, because legal name, trade name and email domain rarely match. The criterion we use scores several attributes at once — tax ID, domain, address, history of engagement with the customer.

Above the threshold, the records become a single entity. Below it, they stay separate. In the middle band, the merge suggestion goes to a person to decide. We chose that middle band because a wrong merge costs more than a duplicate: it blends the history of two companies, enters the system looking legitimate, and nobody notices until someone gets billed for the wrong account. Every unification is logged with the criteria that supported it, which makes it possible to undo and audit later.

Verbalizing the connection makes the graph legible to the model

A relationship in a database is a foreign key. The model receives something like customer_id and has to guess what that means in that operation.

That's why every relationship in the business ontology carries a natural-language description, written in the company's own vocabulary. The link between a person and a company stops being a field and becomes a sentence: this person made a purchase from this company. The link between a contract and an account says which clause applies to whom.

That changes the outcome for a direct reason. The model consumes language, and it's MCP, the open protocol that connects models to data sources, that delivers that verbalized slice at the moment of the question. When the relationship arrives described, the answer can point to the path it took, and whoever reviews it can check without opening four screens.

Schema discovery comes before any ingestion

That's another architecture decision, and it cuts cost as early as the project's first month.

Connecting a source doesn't mean ingesting everything in it. The first step is mapping the schema and listing what's available. The second is sitting down with whoever knows the operation and choosing the tables that answer business questions. Most of a production database is logs, queues and auxiliary tables that never show up in a question from leadership. Ingesting that volume generates processing cost and noise in the modeling.

Documents go through the same criterion. Contracts and amendments become an entity source once they're in a queryable format, with the parties, dates and clauses extracted and linked to the account in the graph.

The business ontology is generated with a person in the loop

There's a fully automatic generation path today, where the tool reads the schema and hands back a ready-made ontology. It works for structure and fails on intent. The schema shows there's a column called status_2. It doesn't say which status counts as a closed deal in the report the board looks at.

We chose semi-automatic generation with human approval. The flow is: schema discovery, selecting the relevant tables, interviewing the business people to extract the rules that live nowhere else, and then generating the business ontology from that material.

The same rule applies after delivery. When a new connector comes in, the agent compares the new structure with what's already modeled and suggests the update. Nothing is applied without approval from someone at the company. Different versions of the ontology can run side by side to find out which one answers the team's real questions better.

The business ontology has prerequisites

The first version will be wrong somewhere. That gets corrected through use, when real questions show where the modeling doesn't match the operation.

It also depends on someone from the business with the mandate to settle a term's definition. A purely technical team delivers an ontology that sales and finance don't recognize as their own operation.

A minimal data catalog also shortens the work. A company that already knows which systems are the source of truth for each subject cuts the discovery phase in half. Anyone who doesn't have that yet can start with the inventory: which systems exist, who owns each one, and which terms each area defines differently.

Where to start your business ontology

Pick a question that today crosses three systems and takes two days to answer. List the entities it touches, the ambiguous terms that show up along the way, and the places where the same company is registered twice. That's the starting scope, and it fits into a few weeks of work.

If you want to pressure-test your own scope with the people who model this every day, reach out.

Model the vocabulary of your company.

Talk to a Strattum expert about business ontology, entity resolution and your Company Brain.