Announcing US$ 3.2M pre-seed round · OneVC · Maya · Norte Ventures Read →
Back to blog

The graph engine decision:
cost and capacity at volume

TL;DR: We put two graph engines to the same target — 200 million nodes and 100 million edges — from the same federated source and the same ontology. The hybrid engine completed the full 300 million elements in 17h 11m of ingestion, with a fixed 29.7 GiB of RAM, 184 GB of disk, and point-lookup latency between 1.2 and 3.9 ms, for US$353 a month. The in-memory engine served 249.3 million elements, the ceiling of the tested instance's 175 GiB, with point-lookup latency of 0.16 to 0.35 ms and a cost of US$1,331 a month. The cutover line we adopted sits around 5 million elements.

Why this choice is the Company Brain's last mile

The graph is the layer agents actually query. Before it there is the ontology, a graph_mapping.yaml file versioned per client, which declares which tables become nodes, which columns become edges and which properties each entity carries. The Memory Worker reads that ontology and materializes the graph, served afterwards over Cypher and MCP.

Sources can arrive two ways: through the lakehouse's clean layer, or by federation, reading the source system in place without copying it. In this campaign we used federation. The goal was not to measure lake ingestion — it was to find where each graph engine stops, what it costs to get there, and what it delivers on read.

The scenario

The base came from Databricks, under Unity Catalog, read by federation without materializing anything into raw or clean. It is 100 million users and 100 million tickets, one ticket per user, with Z-order applied on id and on ticket_id, which physically organizes the files by the lookup column. The federated reader slices the source by key ranges and issues one query per slice, with FEDERATED_SLICE_ROWS set to 250 thousand rows.

The target graph is shallow and wide: 200 million nodes, 100 million edges, a single relationship. The design is deliberate, to measure traversal rather than model complexity.

The ontology fits on one screen. User comes from the federated users table, identified by id. Ticket comes from the tickets table, identified by ticket_id. The OPENED_BY edge is built from two columns on the ticket's own row, ticket_id as source and user_id as target, with no need to query the users table.

Same source, same ontology, same pipeline, two engines with opposite memory philosophies.

The two engines, side by side

DimensionFalkorDB 4.4.1, in-memory engineNeo4j 5.26 CE, hybrid engine
Elements served249.3M, tested instance's ceiling300M, full target
Nodes and edges174.6M nodes · 74.6M edges200M nodes · 100M edges
Load time27h 08m across both ingestion phases17h 11m across both ingestion phases
Memory modelemergent, grows with the dataconfigured, fixed ceiling
Steady-state RAM175 GiB29.7 GiB
Engine disk119 GB, durability only184 GB, is the read path
Point-lookup p500.16–0.35 ms1.2–3.9 ms
Read throughput~4,100 qps~1,000 qps
p95 at 40 concurrent11.1 ms57.7 ms
Instancer7g.8xlarge, 32 vCPU, 256 GiBr7g.2xlarge, 8 vCPU, 61 GiB
Monthly infraUS$1,331US$353

The difference in served volume between the two comes from the memory line, and it is what explains everything else.

Two memory models, two different bills

Emergent memory. In the in-memory engine, RAM is the data itself. The curve, measured over 15,800 samples taken every 15 seconds, showed 0.92 GiB per million nodes in the users phase and 1.13 GiB per million node-plus-edge pairs in the tickets phase. The 0.70 GiB per million elements average is the capacity-planning ruler this benchmark established. Every meaningful growth of the graph becomes an instance decision.

Configured memory. In the hybrid engine, RAM is set before the load: 8 GiB of heap and 20 GiB of page cache, with total consumption stopping at 29.7 GiB. Going from 40 million to 200 million nodes, five times more data, memory moved by 1 GiB. What grows is disk, around 920 bytes per node, which takes the 300 million elements to 184 GB of EBS.

The measured 29.7 GiB breaks down like this: 8 GiB of reserved heap, with usage oscillating between 1.7 and 6.1 GiB, a fixed 20 GiB of page cache holding the graph's hot part, mostly indexes, a 130 MiB slice of thread stacks and native buffers, and the remainder as the JVM's own overhead. The native slice is the only one with no configurable ceiling and it swells during checkpoints, which is why the container needs to be sized with headroom over the sum of heap and page cache.

The disk line nobody budgets for: the entity registry

Both engines lean on a Postgres that holds the entity registry and the resolution index, and it is not small. In the hybrid campaign it added up to 80 GB; in the in-memory one, 73 GB, split between 29 GB of registry and 43 GB of index.

Adding engine and registry together, the hybrid load occupied about 264 GB. The transaction log came in at 1.3 GB because a retention policy was configured; without it, it would have gone past 200 GB during the load.

These two numbers tend to fall outside the initial sizing exercise and show up as a disk incident in the middle of the first big load.

Node weight drives the bill, not the element count

Bytes per element is not a constant of the engine — it is a consequence of the model. Two of our own measurements, on the same technology:

ScenarioElementsBytes per element
Lean nodes, only id and name21M~126
Mix of nodes and edges with properties249.3M~750

The difference is how much property each node carries. The effect on the AWS bill is direct. By the 0.70 GiB-per-million ruler, 150 million elements with properties call for around 105 GiB of RAM, which with headroom for querying and for the persistence fork pushes the choice into the 256 GiB class. The same 150 million with lean nodes land near 18 GiB and fit a 32 GB instance.

The modeling decision that looks harmless — putting three more attributes on every node because "it might be useful later" — is an infrastructure decision. The criterion we adopted is to carry onto the node only what traversal needs and keep the rest at the source, fetched on demand.

Ingestion costs once, reads cost every day

The 17 and 27 hours of load look alarming in isolation. They are one-off events, and the steady state is incremental sync, which processes only the changes since the last run.

Within the load, edges cost more than nodes. In the hybrid engine, the tickets phase took 10h 48m against the users phase's 6h 23m, because every edge does two index lookups, checks whether the relationship already exists, writes the record and updates the linked lists on both nodes. In the in-memory engine, the same asymmetry showed up: 3,188 rows per second on users against 1,126 in the tickets phase with edges.

Two operational notes from the campaign. The watermark used was created_at, and for mutable tables the recommendation is updated_at or direct configuration on the connector, under risk of missing an update to an existing record. And the graph write batch, MEMORY_WORKER_BATCH_SIZE, is the transaction size: it is independent of the federated read slice and is the parameter that moves total load time the most.

Point lookups, the agent's all-day job

A point lookup is the query that goes straight into a known entity, by its indexed identifier, and navigates from there to its neighbors. It is the question "who opened this ticket" and the question "what does this account have open." In support and in conversational agents, it is the dominant pattern, and its latency should not vary with the size of the graph.

ScenarioIn-memory engineHybrid engine
Lookup by indexed entity0.16 ms1.3 ms
Ticket and who opened it0.35 ms3.9 ms
User history0.21 ms3.1 ms
Total node count6.0 ms1.3 ms
Filtered list with LIMIT 3268 ms1.8 ms

Both deliver milliseconds on the point lookup, with an order of magnitude between them, and latency stayed insensitive to volume in both cases.

Under concurrency, the gap widens:

ConcurrentIn-memory p50 / p95Hybrid p50 / p95
10.35 / 0.52 ms2.8 / 6.3 ms
101.99 / 3.89 ms8.7 / 19.5 ms
4010.36 / 11.12 ms35.9 / 57.7 ms

Peak throughput landed around 4,100 queries per second on the in-memory engine and 1,000 on the hybrid. In practice, 1,000 queries per second is 2.6 billion queries a month — well above what a support operation consumes.

And there is the detail that outweighs any engine choice. The same query with an index takes 1.3 ms on the hybrid and 9,520 ms without one — a 7,300-times difference. On the in-memory engine, 0.16 ms against 10,777 ms. A point lookup is only pointed if the entity is indexed; without an index, the engine scans the entire graph and any architectural advantage disappears.

The cutover line we adopted

The platform switches between the two by configuration, via GRAPH_BACKEND, with no code change.

Up to about 5 million elements: the in-memory engine, as-is. It fits in the initial VM's memory slice, delivers sub-millisecond reads and requires no extra operation. It is where most clients sit at the start.

Above that: the hybrid engine. The in-memory engine's memory grows 0.70 GiB per million elements, which forces an instance upgrade with every growth and reaches US$1,331 a month in the 250-million range. The hybrid holds from 10 million to 300 million with the same configured RAM and scales on disk, which is cheap and elastic.

The line's validation is the US$353 a month that served 300 million elements with point-lookup latency between 1 and 4 ms, enough for support and agents. With a one-year savings plan that figure drops to around US$230, and to US$165 over three years.

The settings behind these numbers

ParameterIn-memory engineHybrid engine
Container200 GiB40 GB
Engine memory ceiling180 GiB, noeviction8 GB heap + 20 GB page cache
PersistenceAOF and snapshotscheckpoints, log with 1 GB retention
Ingestion worker12 GiB12 GiB
DuckDB limit6 GB6 GB
Graph write batch40,00040,000
Federated read slice250,000250,000

Two limits deserve attention. The container needs to sit above heap plus page cache, because the JVM uses memory beyond both. And the DuckDB limit needs to be declared, or it sizes itself off the host rather than the container.

Six lessons we carried into building the Company Brain

  1. The graph is a derived projection, not the source of truth. It is deterministically rebuilt from the source crossed with a version of the ontology, which lets you switch engines without losing anything.
  2. Node weight sets the bill. Bytes per element ranged from 126 to 750 in our measurements, and that difference changes the instance class.
  3. Read latency is the product requirement. Loading happens once, querying happens all day, and the real pattern is a point lookup on a known entity.
  4. An index is not a tuning detail. The measured difference was three to four orders of magnitude.
  5. Budget the entity registry alongside the engine. It added up to 73 and 80 GB of Postgres across the two campaigns — the same order of magnitude as the graph itself.
  6. Choosing the watermark is a correctness decision. created_at does not capture an update to an existing record, and on a mutable table that turns into stale data in the graph with no visible error.

Scope and limitations of the study

The data is synthetic and deterministic, which makes the results reproducible, and the code exercised was the production code, with the same connectors, the same federated reader and the same MemoryWorkerPipeline.

The graph tested has a single relationship and one ticket per user. Graphs with many relationships per node change both traversal cost and weight per element, so these numbers hold as a sizing ruler, not as a guarantee.

The data is clean by construction. A round with dirty data, including broken references, duplicate entities and malformed fields, is planned, as are read scenarios with more concurrent users than tested here.

Frequently asked questions

Why not always use the faster engine?

Because the faster one on reads is the more expensive one on memory. At 249.3 million elements, the in-memory engine required a US$1,331-a-month instance, against US$353 for the hybrid serving 300 million elements with latency adequate for agents.

Does switching engines require redoing the source ingestion?

No. The ontology and the pipeline stay the same, and the switch is a configuration change. What needs redoing is the graph load, which is rebuilt from the source.

Can you start small and grow without a migration?

Yes, and it is the default design. Up to about 5 million elements the graph runs embedded on the initial VM, and the engine switch happens by configuration once volume passes that.

Why federate instead of copying the data?

In this case, to isolate the measurement of the graph engine itself. When a company already has a lakehouse or warehouse, federation also avoids duplicating data that is already governed at the source.

What hasn't been measured yet?

Behavior with dirty data, and reads under concurrency above the scenarios tested here. Both are next steps.

Size the graph for your volume.

Talk to a Strattum expert about graph engine choice, memory cost and capacity for your case.