Apache Cassandra in WSL: Single-Node CQL and Replication Boundaries
Build a Cassandra learning environment in WSL, model partition-aware CQL, inspect node state, and understand why a single node cannot test replica resilience.
Apache Cassandra can run in WSL as a learning and application-development environment for CQL schemas, partition-key design, consistency settings, and client integration. A single local node is not a small version of a production cluster in every meaningful respect: it has one process, one machine, one WSL virtual disk, and one failure domain. Use it to test application logic and query shapes, not to claim replication or high availability.
Cassandra’s data model starts with the query contract. The partition key determines where rows are grouped, clustering columns order rows within a partition, and the keyspace chooses a replication strategy. A table that accepts inserts is not necessarily a table that can answer the application’s query efficiently. A local lab should exercise bounded partitions and the exact predicates the application will use.
Install with the matching Java and package instructions
Follow the Apache Cassandra installation guide for the Linux distribution and Java runtime supported by the selected Cassandra release. Do not copy an apt repository stanza from an older release blog or mix package sources without recording which repository owns each package. Capture the Cassandra and Java versions, distro release, and installation method as part of the project’s development setup.
If systemd is enabled in the WSL distro, inspect the package-owned service before starting it:
java -version
systemctl status cassandra --no-pager
sudo systemctl start cassandra
systemctl is-active cassandra
Some installation paths use different commands, unit names, or launch behavior. Use the current official instructions and package metadata if the unit differs. Do not start a foreground Cassandra process against a data directory that a service process already owns.
Wait for the node to finish startup before testing CQL. Cassandra initialization can take time while it creates local state and joins its configured ring. A process that exists is not sufficient evidence that the native transport is ready or that a client can authenticate and execute a query.
Prove the single-node boundary first
Use nodetool status to inspect the local cluster. The status output shows data-center and node state; a local node commonly appears with the Up/Normal state combination once it is ready. Then connect with cqlsh and identify the cluster and keyspaces:
nodetool status
cqlsh 127.0.0.1 9042
In cqlsh, inspect the cluster and keyspaces with the commands supported by the installed release. Keep the output with the test record. Do not use a successful nodetool response as proof that the application can connect through its own configured hostname, port, TLS mode, credentials, or driver protocol.
A single node can demonstrate CQL statements, partition routing within one process, client serialization, and some consistency-level behavior. It cannot demonstrate inter-node streaming, replica repair, topology changes, rack-aware placement, a coordinator surviving node loss, or network partitions between distinct hosts. With replication factor one, losing that only node means losing access to the only replica until the node and its data return.
Design a small query-shaped keyspace
For an isolated one-node exercise, SimpleStrategy with replication factor one is an explicit test configuration also used in Cassandra’s quickstart. The CQL data-definition guide warns that SimpleStrategy is generally not a wise production choice because it does not respect datacenter layouts, and recommends NetworkTopologyStrategy for production. Do not copy a SimpleStrategy keyspace into a real multi-datacenter design.
The following schema stores events for one tenant and supports a bounded time-range read:
CREATE KEYSPACE IF NOT EXISTS wsl_lab
WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 1};
CREATE TABLE IF NOT EXISTS wsl_lab.events (
tenant_id text,
event_time timestamp,
event_id text,
payload text,
PRIMARY KEY ((tenant_id), event_time, event_id)
) WITH CLUSTERING ORDER BY (event_time DESC, event_id ASC);
INSERT INTO wsl_lab.events
(tenant_id, event_time, event_id, payload)
VALUES
('demo', '2026-10-03T12:00:00+0000',
'event-1', 'local event');
SELECT event_time, event_id, payload
FROM wsl_lab.events
WHERE tenant_id = 'demo'
AND event_time >= '2026-10-03T00:00:00+0000'
AND event_time < '2026-10-04T00:00:00+0000'
LIMIT 50;
The partition key is tenant_id, so the example query targets one tenant and a bounded clustering range. That is a demonstration, not a universal model. If a tenant can produce unbounded data in one partition, consider a time bucket or another bounded partitioning strategy based on actual volume and query patterns. Cassandra does not automatically make arbitrary filtering efficient; validate access patterns against the schema before adding indexes.
The primary key also determines which rows are unique. Repeating the same tenant, timestamp, and event ID addresses the same primary-key row rather than appending a separate event. Use a stable unique event identifier or another schema element when events at the same timestamp must remain distinct.
Read consistency settings in context
Cassandra exposes consistency levels that determine how many replicas must respond to a read or write. A QUORUM on a keyspace with replication factor one is still one replica; it is not a majority spread across multiple machines. Testing a consistency setting on one node cannot demonstrate the latency, availability, or reconciliation behavior of the same setting across a replicated cluster.
For a meaningful local exercise, record the keyspace replication factor, the consistency level configured by the driver or cqlsh, the target query, and whether the expected result is present. Do not infer that a successful local write is durable across disk loss. The commit log and local storage path support server recovery mechanisms, but the WSL environment still has a single underlying disk and host lifecycle.
Use cqlsh for exploratory schema work, then move the final schema into a migration or reproducible CQL file. A developer’s interactive shell history is not a deployment plan. Run the setup against an empty disposable keyspace in CI or another clean environment to detect ordering and idempotency assumptions.
Inspect storage, compaction, and recovery state
Cassandra’s write path uses a commit log and memtables before data is flushed into SSTables. Compaction later merges SSTables according to the selected strategy and options. A local table that accepts writes can therefore still build up disk usage or compaction work. Observe node status, table metrics, free disk, and logs while running a controlled ingest, rather than extrapolating from a few manual inserts.
Keep data directories on the distro’s Linux filesystem and let the package own its process and files. Avoid relocating live commit-log or data directories to Windows-mounted paths without checking the filesystem and configuration requirements. If the service fails, preserve its logs and data before changing ownership or deleting SSTables. A guessed cleanup can destroy the only local replica.
For backups, use Cassandra’s documented snapshot and restore guidance appropriate to the cluster and release, and test the restore into a separate environment. A WSL distribution export can move the distro but does not prove that a Cassandra backup is usable, nor does it give an independent recovery point. Copying a live data directory without a defined consistency procedure is not a verified backup.
Keep WSL service lifecycle explicit
An enabled systemd unit starts Cassandra when the distro is running. It does not keep the WSL VM awake or make the node reachable after Windows sleeps, reboots, updates, or shuts down the distribution. Applications should handle a stopped local node with a bounded readiness wait and a clear failure rather than indefinitely retrying.
If a Windows-native client uses Cassandra, test from that client and validate its actual discovered address and port. Network mode, host firewall, listener configuration, and application driver settings are separate from the CQL schema. Do not expose a test database to broader interfaces just to bypass a localhost route issue.
Acceptance criteria for a local Cassandra lab
Accept the environment when a clean install uses a documented Cassandra/Java combination, exactly one package-owned node reaches Up/Normal, cqlsh executes the schema and a bounded query, and the application driver passes a representative read/write test. The keyspace replication factor, consistency level, partition design, disk location, and WSL stop behavior should all be recorded.
Call the result a single-node development lab. It does not prove replica placement, node replacement, repair, quorum availability, or recovery from machine loss. For those tests, use a multi-node environment with separate hosts or failure domains and follow the current Cassandra operations guide for its topology and maintenance procedures.
Related:
- Java Toolchains in WSL: Separate Linux JDKs, Builds, and Windows Runtimes
- WSL 2 Swap Files: Capacity, Placement, and Safe Configuration
Sources:
- Apache Cassandra: Installing Cassandra
- Apache Cassandra: Quickstart
- Apache Cassandra: CQL data definition
- Apache Cassandra: CQL data types and timestamp literal formats
- Apache Cassandra: Dynamo architecture and replication
- Apache Cassandra: Storage engine
- Apache Cassandra: nodetool status
- Apache Cassandra: Backups
- Microsoft: Use systemd to manage Linux services with WSL