TL;DR: We put two graph engines to the same target — 200 million nodes and 100 million edges — from the same federated source and the same ontology. The hybrid engine completed the full 300 million elements in 17h 11m of ingestion, with a fixed 29.7 GiB of RAM, 184 GB of disk, and point-lookup latency between 1.2 and 3.9 ms, for US$353 a month. The in-memory engine served 249.3 million elements, the ceiling of the tested instance's 175 GiB, with point-lookup latency of 0.16 to 0.35 ms and a cost of US$1,331 a month. The cutover line we adopted sits around 5 million elements.
The graph is the layer agents actually query. Before it there is the ontology, a graph_mapping.yaml file versioned per client, which declares which tables become nodes, which columns become edges and which properties each entity carries. The Memory Worker reads that ontology and materializes the graph, served afterwards over Cypher and MCP.
Sources can arrive two ways: through the lakehouse's clean layer, or by federation, reading the source system in place without copying it. In this campaign we used federation. The goal was not to measure lake ingestion — it was to find where each graph engine stops, what it costs to get there, and what it delivers on read.
The base came from Databricks, under Unity Catalog, read by federation without materializing anything into raw or clean. It is 100 million users and 100 million tickets, one ticket per user, with Z-order applied on id and on ticket_id, which physically organizes the files by the lookup column. The federated reader slices the source by key ranges and issues one query per slice, with FEDERATED_SLICE_ROWS set to 250 thousand rows.
The target graph is shallow and wide: 200 million nodes, 100 million edges, a single relationship. The design is deliberate, to measure traversal rather than model complexity.
The ontology fits on one screen. User comes from the federated users table, identified by id. Ticket comes from the tickets table, identified by ticket_id. The OPENED_BY edge is built from two columns on the ticket's own row, ticket_id as source and user_id as target, with no need to query the users table.
Same source, same ontology, same pipeline, two engines with opposite memory philosophies.
| Dimension | FalkorDB 4.4.1, in-memory engine | Neo4j 5.26 CE, hybrid engine |
|---|---|---|
| Elements served | 249.3M, tested instance's ceiling | 300M, full target |
| Nodes and edges | 174.6M nodes · 74.6M edges | 200M nodes · 100M edges |
| Load time | 27h 08m across both ingestion phases | 17h 11m across both ingestion phases |
| Memory model | emergent, grows with the data | configured, fixed ceiling |
| Steady-state RAM | 175 GiB | 29.7 GiB |
| Engine disk | 119 GB, durability only | 184 GB, is the read path |
| Point-lookup p50 | 0.16–0.35 ms | 1.2–3.9 ms |
| Read throughput | ~4,100 qps | ~1,000 qps |
| p95 at 40 concurrent | 11.1 ms | 57.7 ms |
| Instance | r7g.8xlarge, 32 vCPU, 256 GiB | r7g.2xlarge, 8 vCPU, 61 GiB |
| Monthly infra | US$1,331 | US$353 |
The difference in served volume between the two comes from the memory line, and it is what explains everything else.
Emergent memory. In the in-memory engine, RAM is the data itself. The curve, measured over 15,800 samples taken every 15 seconds, showed 0.92 GiB per million nodes in the users phase and 1.13 GiB per million node-plus-edge pairs in the tickets phase. The 0.70 GiB per million elements average is the capacity-planning ruler this benchmark established. Every meaningful growth of the graph becomes an instance decision.
Configured memory. In the hybrid engine, RAM is set before the load: 8 GiB of heap and 20 GiB of page cache, with total consumption stopping at 29.7 GiB. Going from 40 million to 200 million nodes, five times more data, memory moved by 1 GiB. What grows is disk, around 920 bytes per node, which takes the 300 million elements to 184 GB of EBS.
The measured 29.7 GiB breaks down like this: 8 GiB of reserved heap, with usage oscillating between 1.7 and 6.1 GiB, a fixed 20 GiB of page cache holding the graph's hot part, mostly indexes, a 130 MiB slice of thread stacks and native buffers, and the remainder as the JVM's own overhead. The native slice is the only one with no configurable ceiling and it swells during checkpoints, which is why the container needs to be sized with headroom over the sum of heap and page cache.
Both engines lean on a Postgres that holds the entity registry and the resolution index, and it is not small. In the hybrid campaign it added up to 80 GB; in the in-memory one, 73 GB, split between 29 GB of registry and 43 GB of index.
Adding engine and registry together, the hybrid load occupied about 264 GB. The transaction log came in at 1.3 GB because a retention policy was configured; without it, it would have gone past 200 GB during the load.
These two numbers tend to fall outside the initial sizing exercise and show up as a disk incident in the middle of the first big load.
Bytes per element is not a constant of the engine — it is a consequence of the model. Two of our own measurements, on the same technology:
| Scenario | Elements | Bytes per element |
|---|---|---|
| Lean nodes, only id and name | 21M | ~126 |
| Mix of nodes and edges with properties | 249.3M | ~750 |
The difference is how much property each node carries. The effect on the AWS bill is direct. By the 0.70 GiB-per-million ruler, 150 million elements with properties call for around 105 GiB of RAM, which with headroom for querying and for the persistence fork pushes the choice into the 256 GiB class. The same 150 million with lean nodes land near 18 GiB and fit a 32 GB instance.
The modeling decision that looks harmless — putting three more attributes on every node because "it might be useful later" — is an infrastructure decision. The criterion we adopted is to carry onto the node only what traversal needs and keep the rest at the source, fetched on demand.
The 17 and 27 hours of load look alarming in isolation. They are one-off events, and the steady state is incremental sync, which processes only the changes since the last run.
Within the load, edges cost more than nodes. In the hybrid engine, the tickets phase took 10h 48m against the users phase's 6h 23m, because every edge does two index lookups, checks whether the relationship already exists, writes the record and updates the linked lists on both nodes. In the in-memory engine, the same asymmetry showed up: 3,188 rows per second on users against 1,126 in the tickets phase with edges.
Two operational notes from the campaign. The watermark used was created_at, and for mutable tables the recommendation is updated_at or direct configuration on the connector, under risk of missing an update to an existing record. And the graph write batch, MEMORY_WORKER_BATCH_SIZE, is the transaction size: it is independent of the federated read slice and is the parameter that moves total load time the most.
A point lookup is the query that goes straight into a known entity, by its indexed identifier, and navigates from there to its neighbors. It is the question "who opened this ticket" and the question "what does this account have open." In support and in conversational agents, it is the dominant pattern, and its latency should not vary with the size of the graph.
| Scenario | In-memory engine | Hybrid engine |
|---|---|---|
| Lookup by indexed entity | 0.16 ms | 1.3 ms |
| Ticket and who opened it | 0.35 ms | 3.9 ms |
| User history | 0.21 ms | 3.1 ms |
| Total node count | 6.0 ms | 1.3 ms |
| Filtered list with LIMIT 3 | 268 ms | 1.8 ms |
Both deliver milliseconds on the point lookup, with an order of magnitude between them, and latency stayed insensitive to volume in both cases.
Under concurrency, the gap widens:
| Concurrent | In-memory p50 / p95 | Hybrid p50 / p95 |
|---|---|---|
| 1 | 0.35 / 0.52 ms | 2.8 / 6.3 ms |
| 10 | 1.99 / 3.89 ms | 8.7 / 19.5 ms |
| 40 | 10.36 / 11.12 ms | 35.9 / 57.7 ms |
Peak throughput landed around 4,100 queries per second on the in-memory engine and 1,000 on the hybrid. In practice, 1,000 queries per second is 2.6 billion queries a month — well above what a support operation consumes.
And there is the detail that outweighs any engine choice. The same query with an index takes 1.3 ms on the hybrid and 9,520 ms without one — a 7,300-times difference. On the in-memory engine, 0.16 ms against 10,777 ms. A point lookup is only pointed if the entity is indexed; without an index, the engine scans the entire graph and any architectural advantage disappears.
The platform switches between the two by configuration, via GRAPH_BACKEND, with no code change.
Up to about 5 million elements: the in-memory engine, as-is. It fits in the initial VM's memory slice, delivers sub-millisecond reads and requires no extra operation. It is where most clients sit at the start.
Above that: the hybrid engine. The in-memory engine's memory grows 0.70 GiB per million elements, which forces an instance upgrade with every growth and reaches US$1,331 a month in the 250-million range. The hybrid holds from 10 million to 300 million with the same configured RAM and scales on disk, which is cheap and elastic.
The line's validation is the US$353 a month that served 300 million elements with point-lookup latency between 1 and 4 ms, enough for support and agents. With a one-year savings plan that figure drops to around US$230, and to US$165 over three years.
| Parameter | In-memory engine | Hybrid engine |
|---|---|---|
| Container | 200 GiB | 40 GB |
| Engine memory ceiling | 180 GiB, noeviction | 8 GB heap + 20 GB page cache |
| Persistence | AOF and snapshots | checkpoints, log with 1 GB retention |
| Ingestion worker | 12 GiB | 12 GiB |
| DuckDB limit | 6 GB | 6 GB |
| Graph write batch | 40,000 | 40,000 |
| Federated read slice | 250,000 | 250,000 |
Two limits deserve attention. The container needs to sit above heap plus page cache, because the JVM uses memory beyond both. And the DuckDB limit needs to be declared, or it sizes itself off the host rather than the container.
created_at does not capture an update to an existing record, and on a mutable table that turns into stale data in the graph with no visible error.The data is synthetic and deterministic, which makes the results reproducible, and the code exercised was the production code, with the same connectors, the same federated reader and the same MemoryWorkerPipeline.
The graph tested has a single relationship and one ticket per user. Graphs with many relationships per node change both traversal cost and weight per element, so these numbers hold as a sizing ruler, not as a guarantee.
The data is clean by construction. A round with dirty data, including broken references, duplicate entities and malformed fields, is planned, as are read scenarios with more concurrent users than tested here.
Because the faster one on reads is the more expensive one on memory. At 249.3 million elements, the in-memory engine required a US$1,331-a-month instance, against US$353 for the hybrid serving 300 million elements with latency adequate for agents.
No. The ontology and the pipeline stay the same, and the switch is a configuration change. What needs redoing is the graph load, which is rebuilt from the source.
Yes, and it is the default design. Up to about 5 million elements the graph runs embedded on the initial VM, and the engine switch happens by configuration once volume passes that.
In this case, to isolate the measurement of the graph engine itself. When a company already has a lakehouse or warehouse, federation also avoids duplicating data that is already governed at the source.
Behavior with dirty data, and reads under concurrency above the scenarios tested here. Both are next steps.
Talk to a Strattum expert about graph engine choice, memory cost and capacity for your case.