Neo4j in WSL: Graph Workloads, Memory, and Service Lifecycle
Run Neo4j in WSL for graph application development, size heap and page cache against VM limits, and verify queries, imports, logs, and restart behavior.
Neo4j in WSL is useful for building graph-backed applications, exercising Cypher queries, testing migrations, and importing disposable development data on the same Linux side as an application. The database process is still bounded by the WSL distribution and its VM. A healthy Neo4j service does not make WSL an always-on graph cluster, and a local store on the distribution’s virtual disk is not an off-machine backup.
The memory model deserves special attention. Neo4j uses JVM heap for query and transaction work, a page cache for database files and indexes, and additional native/JVM memory for buffers and runtime overhead. These are separate allocations competing with the rest of Linux and the WSL VM. Treat heap plus page cache as only part of the required memory budget.
Install and control the package-owned service
Use Neo4j’s current Linux installation guide for the target Ubuntu release. Follow one package source and one supported installation path; do not layer an archive installation over files owned by the package manager. Record the edition, server version, Java runtime, distro release, configuration path, and service unit for reproducibility.
If systemd is enabled for the WSL distro, inspect the installed service before changing startup behavior:
neo4j --version
systemctl status neo4j --no-pager
sudo systemctl start neo4j
systemctl is-active neo4j
sudo journalctl --unit neo4j -n 100 --no-pager
Unit names or package commands can differ by release. Use the path from the installed package documentation if these commands do not match the actual unit. Do not start a second Neo4j process against the same data directory to work around a service error.
Before first use, follow the package’s documented initial-password procedure and keep credentials out of source control and shell history. A process that is listening is not proof that the application can authenticate, reach the intended database, or run its required query.
Budget heap and page cache together
Neo4j’s current operations manual exposes settings for JVM heap initial/max sizes and for page-cache size. The manual recommends explicit settings to control behavior and recommends keeping initial and maximum heap sizes equal to avoid unwanted full garbage-collection pauses. These values should be chosen from a workload and available memory, not copied from a larger host.
The following is a small configuration example for a disposable dataset, not a production sizing recommendation:
server.memory.heap.initial_size=512m
server.memory.heap.max_size=512m
server.memory.pagecache.size=768m
Heap and page cache do not include all memory consumed by the JVM, native allocations, thread stacks, transaction state, the Linux kernel, or other processes. In WSL, also account for the VM memory ceiling and Windows host pressure. If WSL is constrained to a small memory budget, these values may need to be lower. If they exceed what the VM can sustain, the likely symptom may be swapping or process termination rather than a clean Neo4j-specific limit error.
The Neo4j admin memory-recommendation command can provide a starting estimate for installed databases, but its output is not a substitute for measurements under the application’s query and import patterns. Run a representative workload, inspect JVM and page-cache indicators using the official monitoring path, and adjust one resource category at a time. Keep enough memory outside the configured heap and cache for transaction state and OS work.
Make graph setup and queries repeatable
Put schema constraints, indexes, and seed data in version-controlled migrations or an idempotent setup procedure. A developer’s manually edited graph is not a reliable test fixture. For repeatable development, provision a named database where the edition and deployment support it, or use an isolated disposable instance for each suite.
A Cypher smoke test should verify an application-shaped operation rather than only returning a constant:
CREATE CONSTRAINT person_id IF NOT EXISTS
FOR (p:Person) REQUIRE p.externalId IS UNIQUE;
MERGE (p:Person {externalId: 'wsl-dev-1'})
SET p.name = 'Local Test';
MATCH (p:Person {externalId: 'wsl-dev-1'})
RETURN p.externalId, p.name;
This demonstrates a uniqueness constraint and idempotent node creation. It is not a complete data model: choose labels, relationship direction, indexes, and uniqueness scope from real query patterns. Do not add high-cardinality properties as indexes without measuring their write and storage cost.
The uniqueness rule also changes concurrent write behavior: two application requests attempting to create the same external identity should converge on the same node or surface a constraint conflict that the application handles deliberately. Test both the normal upsert and the duplicate-race path. When the test suite is rerun, cleanup should remove only the fixture identified by its test namespace and should not drop a shared developer database.
Treat index creation as part of schema readiness. A migration that returns before an index is online may leave the first application query with unexpected latency or planning behavior. The setup job should wait for the database’s documented schema operation state and verify the resulting index or constraint, rather than treating a successful CREATE statement as proof that the application is ready.
Use the shell client to check a query from the same side of the WSL boundary as the application:
cypher-shell -a bolt://127.0.0.1:7687 -u neo4j 'RETURN 1 AS ready'
If the command prompts for a password, provide it through the documented interactive or approved local secret mechanism. Do not place a real password in shell history. A successful shell query is only a service-path check; run the application’s migrations, parameterized queries, and result assertions too.
Diagnose memory and import behavior
Graph imports can create large transactions, high heap usage, index-building work, and sustained disk writes. Use a disposable dataset for load testing and monitor the Neo4j logs, host/VM memory, service state, and filesystem free space together. A query that succeeds against a tiny sample says little about how a large import behaves.
Prefer bounded batches for application imports when appropriate, and validate uniqueness and relationship counts after each batch. If startup slows after adding a dataset, distinguish store recovery and index population from a service-manager problem. Read the package journal and Neo4j logs before changing memory settings or deleting files.
The page cache is especially relevant to repeated traversals over data larger than available cache. A cold query and a warm query are not comparable measurements. Record whether the data and indexes fit in cache, the query plan, result size, and host load. Avoid increasing heap simply because the process has high memory use; some memory is deliberately used outside heap, and larger heaps can change garbage collection behavior.
Plan data and lifecycle explicitly
Keep the active database directory on the Linux filesystem inside the distro unless the workload has been measured and supports another path. Accessing files through Windows can be convenient, but database engines rely on storage and locking behavior that should not be assumed to match ext4. Use a logical export or documented Neo4j backup method to move meaningful data rather than copying an active store directory.
Test shutdown by stopping the Neo4j service cleanly, then restarting and running a query. Separately document what WSL shutdown or Windows reboot does to service availability. A systemd unit can start when a distro starts, but it cannot force WSL to remain running. If important local data must survive distro loss, maintain a verified backup outside the VHDX and test restore into another instance.
Acceptance criteria for a WSL graph lab
Accept the environment when a clean install starts one package-owned server, the intended user authenticates, a migration creates the expected constraints and indexes, a representative graph query returns asserted results, and the memory settings fit within the measured WSL budget under a realistic test. Confirm that logs are accessible after both startup and controlled shutdown.
Record whether the instance is disposable, which edition and version were used, where its data lives, the heap and page-cache settings, and which backup or reset procedure applies. Do not claim clustering, failover, or production readiness from a single WSL database. Use a separately operated environment for high availability, shared access, continuous uptime, or recovery objectives.
Related:
- Java Toolchains in WSL: Separate Linux JDKs, Builds, and Windows Runtimes
- WSL 2 and cgroup v2: Resource Control Inside the Linux VM
Sources: