Trino in WSL: A Single-Node SQL Query Lab
Run a bounded Trino coordinator and worker in WSL, configure a test catalog, query sample data, and separate local SQL tests from cluster guarantees.
Trino is a distributed SQL query engine, but its official deployment documentation also describes a single-server configuration for testing. That makes a WSL distribution useful for validating connector configuration, SQL syntax, client behavior, and local query plans. One process performing both coordinator and worker roles is still one node: it cannot test worker loss, distributed scheduling, network exchange between machines, or cluster-level capacity.
The useful development boundary is explicit. Keep the Trino installation, configuration, data directory, and temporary files on the Linux filesystem. Treat local catalogs and credentials as development material. Bind the HTTP listener only to the intended local access path, and verify that a Windows client can reach it through the active WSL network mode. A successful query against local sample data proves that a narrow vertical slice works; it does not certify production performance.
Prepare an isolated installation
Read the current Trino deployment requirements for the selected server release. The Java requirement changes over time, so do not copy a JDK version from an old article. Install a supported Linux JDK within WSL, record java -version, and test that the Trino launcher uses that runtime. Keep the CLI version compatible with the server; consult the CLI documentation for the selected release.
Extract the official server archive to a versioned location, then create a separate data directory. Trino recommends putting this directory outside the installation directory so the data can remain through upgrades. In this local lab it primarily holds logs and other node state; it is not a replacement for a durable warehouse. Avoid /mnt/c for data-intensive Linux processes unless you are intentionally measuring cross-filesystem behavior.
mkdir -p "$HOME/trino-lab/data"
df -T "$HOME/trino-lab/data"
java -version
ulimit -n
Do not keep two server versions pointed at the same data directory. Preserve the configuration and logs separately from downloaded archives. Set ownership so the Linux user that starts Trino can read the installation and write its node data. The server documentation describes file descriptor and process limits as operational requirements; a workstation limit may differ from what a real cluster needs, so record it rather than silently assuming it is production-suitable.
Configure one process as coordinator and worker
Trino’s coordinator parses and plans queries, manages workers, and accepts client requests. A worker executes tasks and processes data. For development and testing, one Trino instance can perform both roles. The coordinator’s discovery URI points back to its own HTTP endpoint. The following values illustrate the shape of a minimal single-machine configuration; preserve any version-specific files and settings from the release you installed.
# etc/node.properties
node.environment=wsl-lab
node.id=wsl-trino-local-01
node.data-dir=/home/USER/trino-lab/data
# etc/config.properties
coordinator=true
node-scheduler.include-coordinator=true
http-server.http.port=8080
discovery.uri=http://127.0.0.1:8080
Replace /home/USER with the actual Linux home directory. Keep the node ID stable across restarts of this same installation; use a different ID for another node. For larger clusters, Trino recommends keeping coordinator work off processing workers because it competes with query planning and monitoring. This local configuration deliberately trades that separation for convenience.
Configure a modest JVM heap based on the actual memory available to the WSL VM. Do not blindly paste a large production heap example from the Trino docs. The JVM needs memory for non-heap allocations and the Linux guest needs memory for the OS and filesystem cache. WSL’s global resource cap and Windows host pressure are different constraints; observe both and leave headroom.
Add a sample catalog and execute a real query
A server without a catalog has no data source to query. Trino ships a TPCH connector useful for synthetic sample data. Consult the current connector reference for supported catalog properties, then create a catalog file under the installation’s etc/catalog directory. The minimal TPCH connector configuration is:
connector.name=tpch
Start the server using the launcher shipped with the selected distribution and inspect its log output. Do not use nohup or a detached terminal until foreground startup and shutdown are understood. Confirm the listener, then use a matching Trino CLI:
ss -ltnp | grep ':8080'
./trino http://127.0.0.1:8080 --execute 'SELECT count(*) FROM tpch.tiny.nation'
The CLI speaks Trino’s HTTP-based client protocol to the coordinator. A useful acceptance test is not just opening port 8080: run a query, check its result, and inspect the server logs and query state. Then run an intentionally invalid query and verify that the client receives an error rather than a hung session. Use a bounded dataset such as tpch.tiny for repeatable checks; avoid unbounded scans on a laptop.
If you add a connector to an external database or object store, remember that connector credentials and network access are separate from SQL syntax. Keep secrets out of a committed catalog file. Validate TLS, authentication, and endpoint DNS using the same Linux-side client environment as Trino. A Windows application connecting to the coordinator tests one route; the coordinator’s own connection to an external service tests another.
For a client integration test, capture the query identifier and completion state, not just the first response. Trino’s client protocol can return results through repeated requests to the coordinator, and a client that stops fetching results can affect query completion. Test both a query that returns a small result set and a query that fails validation. Use the CLI’s batch mode with a bounded SQL file for repeatable checks, and keep output files outside the installation directory. Cancel a deliberately long query through the documented client path and verify the server reports its final state.
WSL storage, networking, and resource boundaries
Put query scratch and local data on the Linux filesystem. Microsoft documents that Linux command-line workloads often perform better there than on Windows-mounted paths. Do not benchmark Trino from /mnt/c and attribute all latency to the query engine. Record where connector data is actually stored, whether it is local or remote, and whether the test includes cold or warm filesystem cache.
The server defaults and catalog choices can expose sensitive data if a listener is broader than expected. Keep the local HTTP endpoint restricted to the test boundary and avoid exposing it to a LAN or public interface. If Windows cannot reach it, inspect the listener and WSL networking mode before binding to all interfaces. NAT mode’s localhost forwarding and mirrored networking do not behave identically.
WSL’s memory cap applies to the Linux VM, while the Java heap is only one part of that allocation. If queries fail under pressure, compare JVM logs, process RSS, Linux free memory, and Windows host usage. Increasing -Xmx can make the condition worse by starving the guest kernel and other processes. Use small datasets and one change at a time.
Diagnose and preserve the lab
For startup failures, check the selected Java runtime, malformed properties, data-directory ownership, free disk space, open-file limit, and logs. If the server starts but a client cannot connect, inspect whether the endpoint is bound to loopback or a WSL address, whether the port is occupied, and whether the client uses the expected HTTP/HTTPS endpoint. If the server registers but queries fail, inspect catalog initialization and connector errors before changing scheduler settings.
Stop the server cleanly and verify the process exits before upgrading or deleting files. Keep the configuration and sample SQL in version control, but exclude machine-specific data, logs, and passwords. The data directory should not be copied while the service is running and then treated as a query-consistent backup. Rebuild the disposable lab from its pinned artifact and text configuration when practical.
Acceptance criteria
Accept the WSL lab when the chosen Trino and Java versions are recorded, the single node starts with a separate Linux data directory, the listener is scoped to local development, the TPCH catalog loads, and a representative query returns the expected result from both the Linux CLI and any intended Windows client. Record the memory limit and confirm clean shutdown.
This setup is a SQL integration and connector-development target. It does not model multi-node query distribution, production authentication architecture, independent failure domains, workload management, or durable data-lake operations.
Related:
- Java Toolchains in WSL: Separate Linux JDKs, Builds, and Windows Runtimes
- How to Configure WSL2 Resource Limits with .wslconfig
Sources: