Migrating from Neo4j to KGLite¶
This page is for a developer with an existing Neo4j database and/or
neo4j-driver code who wants to evaluate or adopt KGLite. It covers
where KGLite fits, reusing tested Bolt driver code or moving to the native
Python API, how to transfer data, and where the Cypher dialect diverges.
KGLite ships a focused openCypher subset, not a Neo4j drop-in replacement. Many common reads use the same syntax; check the divergences below against your workload before adopting it.
When KGLite fits — and when it doesn’t¶
KGLite |
Neo4j |
|
|---|---|---|
Deployment |
Embedded, in-process ( |
Server deployment or embedded Java database |
Query language |
Supported Cypher dialect (see below) |
Cypher |
Storage |
|
Server store directory |
Auth |
None in-process; basic only via Bolt server |
Full RBAC |
Multi-database |
No catalog — one graph per handle or Bolt server |
Yes ( |
Clustering / routing |
No (single server) |
Causal cluster, routing |
Transactions |
Snapshot isolation + OCC; durability depends on access mode |
Server-managed ACID transactions |
Data model |
One primary type per node + optional secondary labels |
Arbitrary label sets |
KGLite fits when you want Cypher + Python ergonomics in one wheel:
analytics over a graph that fits on one machine, embedding a graph in
a Python app or notebook, shipping a queryable .kgl artifact, or
serving a read-mostly graph to LLM agents (KGLite bundles an MCP
server and a describe() schema). See the
README comparison table
for the side-by-side against other embedded graph engines, NetworkX,
rustworkx, and Neo4j Embedded.
KGLite does not fit when you need server-mode RBAC, multiple databases per instance, a causal cluster with routing, or long-lived multi-client write transactions managed by a database server.
For positioning detail see Core Concepts and the concepts index.
Two migration paths¶
Path A — reuse tested Bolt driver code¶
kglite-bolt-server is a pure-Rust binary that speaks the
Bolt v5 wire protocol. The
official Python, JavaScript, and Java drivers are regression-tested in
CI. Reconfigure the connection URL and authentication; the Java driver also
needs the server started with --neo4j-compat (see the note below). Other Bolt
v5 clients are untested and need evaluation within the documented protocol and
Cypher dialect limits:
cypher-shell connects and queries; standalone
CALL proc(),SHOW INDEXES/CONSTRAINTS/PROCEDURES, andSHOW DATABASESanswer.Neo4j Browser requires
--neo4j-compat(it readsdbms.components()for its version banner) and then populates its sidebar and schema tab fromdb.labels()/db.relationshipTypes()/db.propertyKeys()/db.schema.visualization(). Verify against your Browser version; its connect sequence varies across releases — run the server atRUST_LOG=debugto see exactly what it sends.LangChain’s
Neo4jGraphdoes NOT work unchanged: itsrefresh_schema()callsapoc.meta.data()and its graph-document writer usesapoc.merge.*, and KGLite ships no APOC (see theapoc.*row below). Construct it withrefresh_schema=Falseand supply the schema text yourself (e.g. fromdescribe()), and write through Cypher rather thanadd_graph_documents.
See the Bolt server operator guide.
cargo install kglite-bolt-server
kglite-bolt-server --graph my-graph.kgl --bind 127.0.0.1 --port 7687
Note
JVM clients need --neo4j-compat. The official Java driver refuses to talk
to a server whose handshake agent does not begin with Neo4j/, and fails at
connect time with UntrustedServerException: Server does not identify as a genuine Neo4j instance — before any query runs. The Python and JavaScript
drivers do not perform this check, so they work against the default identity.
Start the server with compatibility mode (or set
KGLITE_BOLT_NEO4J_COMPAT=1) and the agent becomes
Neo4j/5.26.0 (kglite-bolt-server/<version>), which the driver accepts while
still naming the real product:
kglite-bolt-server --graph my-graph.kgl --neo4j-compat
It is opt-in because presenting as another product is the operator’s decision. Connect a driver that enforces the check with compatibility off and the server log tells you exactly this, naming both activation routes. See “Driver identity” in the Bolt server operator guide.
Connection and authentication setup changes; query code may be reusable within the documented Bolt and Cypher limits:
# Before — against Neo4j
from neo4j import GraphDatabase
driver = GraphDatabase.driver("neo4j://prod-db:7687", auth=("neo4j", "secret"))
# After — against kglite-bolt-server
from neo4j import GraphDatabase
driver = GraphDatabase.driver("bolt://127.0.0.1:7687", auth=None)
with driver.session() as session:
result = session.run(
"MATCH (p:Person)-[:KNOWS]->(f) WHERE p.age > $min RETURN f.name",
min=30,
)
for record in result:
print(record["f.name"])
The query path is the same Cypher engine the Python API uses;
differential tests confirm row-for-row equivalence
(tests/test_bolt_server_differential.py). The official Python,
JavaScript, and Java drivers all have automated regression coverage in
CI — session and explicit-transaction lifecycle, the managed
executeWrite path, PackStream type round-trips,
Node/Relationship/Path values, Neo.* error codes, and the OCC
conflict status code (tests/conformance/). Those per-driver checks run
the managed API uncontended; that a managed transaction actually retries
a conflict is covered separately, under real contention, against the
Python driver
(tests/test_bolt_server_transactions.py::test_managed_transaction_retries_after_conflict).
Go and .NET are untested; validate them before use.
What carries over, and what does not¶
Neo4j feature |
Bolt server status |
|---|---|
|
Supported — use this |
|
Single-server routing table only; set |
Auth |
|
TLS ( |
Supported via |
Read-only enforcement |
|
Auto-commit mutations |
Not supported — wrap every write ( |
OCC on writes |
Supported — stale-snapshot commits get |
Multi-database ( |
Not supported — single graph. A |
Causal consistency / bookmarks |
Not supported — the |
Multi-statement queries ( |
Not supported — one statement per |
|
Yield |
The in-process Python API also exposes
kglite.to_neo4j(graph, uri, ...)if you want to push a KGLite graph into a real Neo4j instance (batchedUNWIND, optionalmerge=Trueupsert).
Path B — native Python (cypher() directly)¶
If you control the calling code, skip the wire protocol entirely and
call cypher() in-process. No server, no socket, no driver — the
result is a ResultView you iterate, index, or convert with
to_df=True. See Getting Started.
The same query, both ways:
# Bolt path — neo4j driver
with driver.session() as session:
rows = list(session.run(
"MATCH (p:Person) WHERE p.age > $min RETURN p.name AS name",
min=30,
))
# Native path — kglite in-process
import kglite
graph = kglite.load("my-graph.kgl")
rows = list(graph.cypher(
"MATCH (p:Person) WHERE p.age > $min RETURN p.name AS name",
params={"min": 30},
))
Note the parameter syntax difference: the driver takes **kwargs
(or a parameters= dict); cypher() takes a params= dict. The
$name placeholders in the query string are identical.
Getting data out of Neo4j into KGLite¶
There are three routes; pick by what you already have.
Route 1 — query the source database directly, bulk-load via pandas¶
Query the source database with the neo4j driver, pull rows into a
pandas DataFrame, and bulk-load with add_nodes / add_connections.
import pandas as pd
import kglite
from neo4j import GraphDatabase
src = GraphDatabase.driver("neo4j://prod-db:7687", auth=("neo4j", "secret"))
graph = kglite.KnowledgeGraph()
# Nodes
with src.session() as s:
people = pd.DataFrame([
dict(r["p"]) for r in s.run("MATCH (p:Person) RETURN p")
])
graph.add_nodes(people, node_type="Person", unique_id_field="id",
node_title_field="name")
# Relationships — return the endpoint ids, not the whole nodes
with src.session() as s:
knows = pd.DataFrame([
{"src": r["a"], "tgt": r["b"]}
for r in s.run("MATCH (a:Person)-[:KNOWS]->(b:Person) "
"RETURN a.id AS a, b.id AS b")
])
graph.add_connections(knows, connection_type="KNOWS",
source_type="Person", source_id_field="src",
target_type="Person", target_id_field="tgt")
graph.save("my-graph.kgl")
add_nodes auto-detects string vs integer ids from the column dtype
and supports a column_types= override for spatial/temporal columns;
add_connections can take a Cypher query= instead of a DataFrame.
See the data-loading guide.
Route 2 — dump to CSV, then LOAD CSV (no pandas needed)¶
KGLite runs LOAD CSV, so CSV is a direct migration route. Export with APOC
(apoc.export.csv.all('graph.csv', {}), or per-label queries for a clean
node/edge split), then adapt the import query to the restrictions below:
import kglite
graph = kglite.KnowledgeGraph()
# Nodes. Fields arrive as strings — CSV carries no types — so convert
# explicitly, as the source script already does.
graph.cypher(
"LOAD CSV WITH HEADERS FROM 'file:///export/people.csv' AS row "
"CREATE (:Person {id: toInteger(row.id), name: row.name})"
)
# Relationships, by endpoint id. Pattern properties take variables
# rather than function calls, so convert in a WITH first.
graph.cypher(
"LOAD CSV WITH HEADERS FROM 'file:///export/knows.csv' AS row "
"WITH toInteger(row.src) AS src, toInteger(row.dst) AS dst "
"MATCH (a:Person {id: src}), (b:Person {id: dst}) "
"CREATE (a)-[:KNOWS]->(b)"
)
graph.save("my-graph.kgl")
Prefer this route if you have no pandas (a Bolt client, a Rust or JVM consumer) or if you already have import scripts you would rather not rewrite. Four differences are worth knowing before you run it:
KGLite |
Neo4j |
|
|---|---|---|
Sources |
|
|
Batching |
Automatic — 1000 rows at a time for row-local pipelines, so file size does not drive memory |
|
Whole-result clauses |
An aggregate, |
Behavior depends on the execution plan and server configuration |
Position |
Must be the first clause |
Anywhere in the pipeline |
http(s):// is rejected with a message rather than a syntax error: the
engine carries no HTTP client at all (network dependencies were removed
in 0.14.x), so there is nothing to fetch a URL with. Download the file
first, or fetch it in your own code and pass the rows in as a
parameter.
Over Bolt, LOAD CSV is off by default. file:// means the
server’s filesystem, and a Bolt client is a remote caller, so serving
it ungated would publish an arbitrary-file-read primitive. In-process
callers (this Python API, the Rust library, the CLI) are allowed because
they already have the host process’s filesystem access; a Bolt client
gets nothing unless the server was started with `–allow-csv-import
Route 3 — pandas between export and load¶
Still the best fit when you want typing control, column renaming, or
cleanup in between: pd.read_csv → add_nodes / add_connections,
using the same calls as Route 1.
Cypher dialect divergence¶
KGLite’s supported surface is documented in full in
CYPHER.md;
this section lists only where it diverges from Neo4j. An opt-in live comparison
runner is available at scripts/cypher_conformance.py
(see Cypher Compatibility — Independent Differential Checks).
Data model — labels and node identity¶
Neo4j form |
KGLite status |
Workaround / note |
|---|---|---|
Arbitrary label sets |
One primary type + optional secondary labels (since 0.10.5) |
|
Retype a node by swapping labels |
Primary type is immutable via label ops |
Recreate/migrate the node under the new primary type; |
Per-row label assignment at load |
|
|
|
|
See identity note below |
Note: Neo4j docs and some older KGLite material describe KGLite as “single-label” with
labels(n)returning a string. That changed in 0.10.5 — multi-label is native andlabels(n)returns a list.
Node identity (id) — the 0.10.10 model¶
As of 0.10.10, n.id is the node’s unique identity and behaves
identically in every storage mode.
Aspect |
KGLite behaviour |
|---|---|
|
Honours |
Prefixed-id datasets (Wikidata |
Loader stores the integer as |
Lookup by string id |
|
Duplicate ids |
|
This is a breaking change from earlier releases for prefixed-id data — see the 0.10.10 CHANGELOG entry.
Missing language constructs¶
Current unsupported and partial constructs:
Neo4j construct |
KGLite status |
Workaround |
|---|---|---|
|
Supported |
Updating bodies, including nested |
|
Not supported |
Do writes in a separate top-level clause; read subqueries are supported (see below) |
Unit |
Not supported |
Body must end in |
|
Not supported |
Server batching; no in-memory analogue |
|
Not supported |
A read |
|
Not supported |
Use an explicit |
Pattern comprehensions |
Not supported |
|
Quantified path patterns |
Not supported |
Variable-length paths |
|
Supported — |
|
|
Not supported |
|
|
Not supported as a |
|
|
Not supported — rejected with the route that applies |
|
|
Parses, then rejected by name rather than approximated |
Only the type names with an exact KGLite value counterpart are accepted ( |
|
Not supported |
KGLite has no single answer for when two relationships of a type are the same one, so uniqueness would mean different things on different write paths. |
Constructs that DO work (worth confirming)¶
These forms are supported and are easy to assume missing:
MERGE ... ON CREATE SET ... ON MATCH SET— match-or-create.Variable-length paths
-[:KNOWS*1..3]->,shortestPath(...), andallShortestPaths(...).WHERE EXISTS { pattern WHERE ... }(pattern-existence), inline pattern predicates,any/all/none/single(x IN list WHERE ...).CALL { ... }read subqueries run per input row, whether or not they import outer variables. ModernCALL (p, q),CALL (*), andCALL ()scope syntax is supported; those imports remain visible acrossWITHand every set arm. The legacyCALL { WITH p ... }form remains available and requires a separate bare-variable importingWITHin each arm.UNION/UNION ALLwork inside the body, withINTERSECT/EXCEPTas KGLite extensions. Aggregating bodies preserve an outer row with a zero; non-aggregating bodies inner-join, so zero returned rows drop it. Writes, unit bodies, andIN TRANSACTIONSremain unsupported. See CYPHER.md →CALL { ... }read subqueries.Ordinary
CALL procedure(...) YIELD ...evaluates its parameters and joins its results per incoming row.cluster()is the deliberate exception: it consumes the full preceding cohort.Cypher 25
FILTER,OFFSET,NODETACH DELETE, and terminalFINISH.INSERTis supported for static node labels (&between multiple labels) and one directed static relationship type. It rejects dynamic labels/types, dynamic property maps, path assignment, colon-separated multiple labels, relationship type alternation, and undirected or variable-length edges; keepCREATEwhere one of those CREATE-only forms is required.List comprehensions
[x IN list WHERE p \| expr],reduce(...), list slicingxs[1..3], map projectionsn {.a, .b}, map literals.Map subscript
m['key']and dynamic property accessn[key]wherekeyis a variable.Window functions
row_number()/rank()/dense_rank() OVER (...),UNION/INTERSECT/EXCEPT,HAVING.
Function coverage¶
KGLite covers the common scalar / string / math / aggregation / temporal / spatial families. Rather than duplicate them, see the function tables in CYPHER.md (Built-in, String, Math, Spatial, Temporal, Timeseries, Text predicates, plus the openCypher compatibility matrix).
Current notable function differences:
Neo4j |
KGLite status |
Note / workaround |
|---|---|---|
|
Not supported, with exactly two exceptions |
No APOC library. |
|
Map form not supported |
KGLite uses |
|
Use top-level |
Geodesic (WGS84); also |
|
Map form only |
|
|
Not supported |
|
|
Not supported |
|
Calendar-aware month diffs |
Approximated (months ≈ 30 days in |
Use literal dates for exact month arithmetic — see CYPHER.md “Duration semantics” |
KGLite-specific function names include semantic search
(text_score/vector_score), timeseries (ts_*), fuzzy text
predicates (text_edit_distance, text_jaccard), and graph-algorithm
procedures (CALL pagerank/louvain/...). See CYPHER.md.
EXPLAIN / PROFILE¶
Both are supported but the shape differs from Neo4j’s plan tree:
EXPLAIN <query>returns aResultViewwith rows[step, operation, estimated_rows]— a flat, ordered step list, not a nested operator tree.PROFILE <query>executes the query (you get the real results) and attaches per-clause stats onresult.profile([clause, rows_in, rows_out, elapsed_us]).
Every cypher() call also attaches lightweight result.diagnostics
(elapsed_ms, timeout_ms, row_limit, total_rows, warnings) with no
prefix required.
Operational differences¶
Concern |
Neo4j |
KGLite |
|---|---|---|
Persistence |
Live server store directory |
A |
Backup |
|
Copy the |
Concurrency |
Server-managed sessions, ACID |
Reads parallelize (GIL released via |
Cross-process access |
Native (server) |
Embedded — use the Bolt server as the coordination point for multi-process |
Schema DDL |
|
Index DDL is supported — |
Constraint DDL |
|
Supported and enforced on every write path, including the bulk loader. Composite tuples ( |
Migrations |
Versioned migration tools |
None — you own schema evolution in Python load code |
Indexes are maintained automatically across Cypher mutations, including
CREATE/INSERT, property updates/removals, deletes, and MERGE. On disk-backed graphs
property indexes are persisted next to the store; on in-memory graphs
they live in a HashMap. See the Indexes section of
CYPHER.md.
For the transaction model (snapshot isolation, OCC, last-writer-wins, per-call cost) see Transactions and sessions; for the concurrency contract see Concurrency.
See also¶
Getting Started — install, build a graph, run Cypher.
Core Concepts — nodes, relationships, storage modes.
Transactions and sessions —
begin()/commit()/ OCC.Cypher Compatibility — Independent Differential Checks — how behavioral comparisons work.
CYPHER.md — authoritative supported Cypher reference.