High Availability Deployment
EnterpriseDeploy NetStacks Controller in active-active HA with multiple instances, Valkey Sentinel, external PostgreSQL, Nginx load balancing, TLS, and backups.
Overview
NetStacks Controller supports active-active high availability, meaning every controller instance actively serves requests simultaneously. There is no primary/standby distinction among controller nodes — any instance can handle any API call, WebSocket connection, or SSH proxy session. The cluster is glued together by a shared PostgreSQL database, a shared VAULT_MASTER_KEY, and Valkey for ephemeral session coordination.
The HA architecture eliminates single points of failure across the stack:
- Controller instances — Two or more instances run behind a load balancer. If one goes down, the remaining instances continue serving all traffic.
- Valkey (session coordination) — A Redis-compatible in-memory store that tracks which controller instance owns each session, enabling cross-instance session routing. Valkey Sentinel provides automated failover for the Valkey tier itself.
- PostgreSQL — The database stores all persistent state: users, devices, credentials, audit logs, certificates, and configuration (including the JWT signing secret). In HA mode, use an external PostgreSQL cluster with replication for database-level redundancy.
- Load balancer — Nginx (or any WebSocket-aware load balancer) distributes traffic across controller instances using
least_connrouting, with health-check-based removal of unhealthy nodes.
Capabilities enabled when Valkey is configured:
- Instance registration and heartbeat — Each controller generates a random instance ID at startup and registers in Valkey with a periodic heartbeat. Stale instances are cleaned up automatically.
- Session registry — Active sessions are tracked in Valkey with the owning instance's advertise address, enabling cross-instance session lookup and routing.
- HA status — The Admin UI Dashboard includes an HA Status widget showing cluster mode, instance list, session distribution, and Valkey connectivity. It is backed by the
GET /api/admin/ha-statusendpoint.
High availability is an Enterprise Controller capability. Single-instance Controller deployments do not require Valkey or multiple controller instances — without Valkey configured, the controller simply runs in single-instance mode (mode: "single-instance").
Architecture
The following diagram shows a production HA deployment. The load balancer receives all external traffic and distributes it across the controller instances. Both controllers share the same PostgreSQL database and Valkey cluster for state coordination.
+----------------------+
| Terminal Clients |
| (Tauri desktop app) |
+----------+-----------+
|
HTTPS / WSS
|
+----------v-----------+
| Nginx Load Balancer |
| (least_conn, TLS |
| termination) |
+----+------------+----+
| |
+--------v--+ +----v--------+
|Controller | | Controller |
|Instance 1 | | Instance 2 |
| :3000 | | :3000 |
+--+----+---+ +--+----+-----+
| | | |
+--------+ +----+-----+ +--------+
| | |
+---------v--------+ +------v---------+ +--------v-------+
| PostgreSQL 16 | | Valkey Primary | | Network Devices|
| (pgvector ext) | | + Replicas | | (SSH/Telnet) |
| | | | +----------------+
| Users, devices, | | Session state, |
| credentials, | | instance reg, |
| audit logs, | | heartbeats |
| jwt secret | | |
+------------------+ +-------+--------+
|
+--------v--------+
| Valkey Sentinel |
| (3 instances) |
| Monitors |
| primary, auto- |
| promotes replica|
+-----------------+In this architecture:
- Nginx terminates TLS and distributes traffic to controller instances using
least_conn. It passes WebSocket upgrades through to the backend and proxies to the controllers over HTTPS (proxy_pass https://...). - Controller instances are stateless application servers. All persistent state lives in PostgreSQL; all ephemeral session state lives in Valkey. Scale horizontally by adding more instances.
- Valkey Primary + Replicas handle session coordination. The primary handles writes; replicas provide read scaling and failover candidates.
- Valkey Sentinel monitors the primary and automatically promotes a replica if the primary fails. Three sentinels are required for quorum.
- PostgreSQL is the single source of truth for all persistent data. Use an external PostgreSQL cluster with streaming replication for database-level HA.
The Enterprise Controller repository ships a reference HA compose file (docker-compose.production.yml) that builds local images named netstacks-controller and netstacks-nginx. The examples below mirror that shipped configuration. Substitute your own image names/registry as appropriate for your environment.
Prerequisites
Before deploying NetStacks Controller in any configuration, ensure the following requirements are met:
| Requirement | Minimum | Recommended (HA) |
|---|---|---|
| Docker Engine | 24.0+ | 25.0+ |
| Docker Compose | v2.20+ | v2.24+ |
| RAM per controller instance | 4 GB | 8 GB |
| CPU per controller instance | 2 vCPU | 4 vCPU |
| PostgreSQL | 16+ with pgvector | 16+ with pgvector, streaming replication |
| Valkey / Redis | Not required (single instance) | Valkey 8+ with Sentinel |
| Disk (database) | 20 GB | 100 GB+ SSD |
| Disk (Valkey) | 1 GB | 4 GB SSD (AOF persistence) |
| Network | Controller reaches devices on SSH/Telnet ports | Low-latency interconnect between controller instances |
| License | Enterprise Controller (Contact Sales) | Enterprise Controller (Contact Sales) |
The Enterprise Controller is provisioned through a sales engagement rather than a self-service license key. There is no per-instance seat key to copy between controllers — all instances in a cluster share one database and run as one logical deployment. Contact NetStacks to obtain the Enterprise Controller.
On AWS, use RDS for PostgreSQL with the pgvector extension (supported on RDS 16+) and ElastiCache for Valkey. On GCP, use Cloud SQL for PostgreSQL and Memorystore for Redis/Valkey. On Azure, use Azure Database for PostgreSQL Flexible Server and Azure Cache for Redis.
Single Instance Setup
Start with a single-instance deployment to verify your configuration before scaling to HA. This setup runs PostgreSQL and the controller API in a single Docker Compose stack. Valkey is optional — without it, the controller runs in single-instance mode.
Generate Secrets
Before creating the Docker Compose file, generate the required secrets. These values must be kept secure and consistent across all controller instances if you later scale to HA.
# Generate the vault master key (64 hex characters = 32 bytes)
# This encrypts all credentials, SSH CA keys, and sensitive data in the database
export VAULT_MASTER_KEY=$(openssl rand -hex 32)
echo "VAULT_MASTER_KEY=$VAULT_MASTER_KEY"
# Generate the service-token secret (used for internal service-to-service auth)
export SERVICE_TOKEN_SECRET=$(openssl rand -hex 32)
echo "SERVICE_TOKEN_SECRET=$SERVICE_TOKEN_SECRET"
# Generate the database password
export DB_PASSWORD=$(openssl rand -base64 24)
echo "DB_PASSWORD=$DB_PASSWORD"
# Save these values securely. You will need VAULT_MASTER_KEY for every
# controller instance in HA mode.The VAULT_MASTER_KEY is the root of trust for all encrypted data in NetStacks. If lost, encrypted credentials and SSH CA private keys cannot be recovered. Store it in a secrets manager (HashiCorp Vault, AWS Secrets Manager, etc.) or at minimum in a secure, backed-up location outside the Docker host.
Unlike the vault master key, the JWT signing secret is not a controller environment variable. It is stored as the auth.jwt_secret setting in PostgreSQL and configured from Admin → Settings. Because it lives in the shared database, every controller instance automatically uses the same signing key — tokens issued by one instance are accepted by all others. Change it from the default placeholder before going to production.
Docker Compose (Single Instance)
services:
postgres:
image: pgvector/pgvector:pg16
restart: unless-stopped
environment:
POSTGRES_USER: netstacks
POSTGRES_PASSWORD: ${DB_PASSWORD}
POSTGRES_DB: netstacks
volumes:
- postgres-data:/var/lib/postgresql/data
ports:
- "127.0.0.1:5432:5432"
healthcheck:
test: ["CMD-SHELL", "pg_isready -U netstacks"]
interval: 5s
timeout: 5s
retries: 5
controller:
image: netstacks-controller:latest
restart: unless-stopped
depends_on:
postgres:
condition: service_healthy
ports:
- "3000:3000"
environment:
# Database
DATABASE_URL: postgres://netstacks:${DB_PASSWORD}@postgres:5432/netstacks
# Security — generate with: openssl rand -hex 32
VAULT_MASTER_KEY: ${VAULT_MASTER_KEY}
SERVICE_TOKEN_SECRET: ${SERVICE_TOKEN_SECRET}
# Server bind (HOST defaults to 0.0.0.0, PORT defaults to 3000)
HOST: 0.0.0.0
PORT: 3000
# Instance naming (optional, shown in the Admin UI HA status)
NETSTACKS_INSTANCE_NAME: "controller-1"
# TLS — the controller always serves HTTPS. With no cert/key paths set it
# auto-generates a CA + server cert under TLS_DATA_DIR on first boot.
TLS_DATA_DIR: /data/tls
# Logging
RUST_LOG: "info,netstacks_api=info"
volumes:
- controller-data:/data
volumes:
postgres-data:
controller-data:Start the Stack
# Create a .env file with your generated secrets
cat > .env << 'EOF'
VAULT_MASTER_KEY=<your-64-hex-char-key>
SERVICE_TOKEN_SECRET=<your-64-hex-char-secret>
DB_PASSWORD=<your-database-password>
EOF
# Start the stack
docker compose up -d
# Watch the logs for the initial admin password
docker compose logs -f controllerVerify the Deployment
The basic health endpoint is unauthenticated and returns the service status and version. It is available both at the root (legacy callers) and under /api (for callers using the/api base URL):
# Basic health check (use -k for the self-signed cert). Both paths work:
curl -k https://localhost:3000/health
curl -k https://localhost:3000/api/health
# Expected response:
# { "status": "ok", "version": "0.0.14" }
# Readiness probe — checks the database (and Valkey if configured)
curl -k https://localhost:3000/ready
# { "status": "ready", "database": "connected", "valkey": "not configured" }
# Log in and get a token
curl -k -X POST https://localhost:3000/api/auth/login \
-H "Content-Type: application/json" \
-d '{"username": "admin", "password": "<password-from-logs>"}'Log in to the Admin UI at https://your-host:3000, change the default admin password, and set a strong auth.jwt_secret under Admin → Settings before exposing the controller. See Settings and User Management.
High Availability Setup
The full HA deployment adds Valkey with Sentinel for session coordination, multiple controller instances, and an Nginx load balancer. All controller instances must share the same DATABASE_URL and VAULT_MASTER_KEY. The JWT signing secret is shared automatically because it lives in the database.
Every controller instance in the cluster must use an identical VAULT_MASTER_KEY. If keys differ between instances, credentials encrypted on one instance fail to decrypt on another. Generate the key once and distribute it to all instances via your secrets management system. The auth.jwt_secret setting does not need manual distribution — all instances read it from the shared database.
Full HA Docker Compose
This mirrors the shipped docker-compose.production.yml reference: PostgreSQL with pgvector, a Valkey primary plus two replicas plus three Sentinels, two controller instances sharing a YAML anchor, and an Nginx reverse proxy that terminates TLS.
x-controller-common: &controller-common
image: netstacks-controller:latest
restart: unless-stopped
depends_on:
postgres:
condition: service_healthy
valkey-sentinel-1:
condition: service_healthy
environment: &controller-env
DATABASE_URL: postgres://netstacks:${DB_PASSWORD}@postgres:5432/netstacks
VAULT_MASTER_KEY: ${VAULT_MASTER_KEY}
SERVICE_TOKEN_SECRET: ${SERVICE_TOKEN_SECRET}
HOST: 0.0.0.0
PORT: 3000
RUST_LOG: ${RUST_LOG:-info,netstacks_api=info}
SERVE_STATIC: "false"
TLS_DATA_DIR: /data/tls
# Valkey via Sentinel — full redis:// URLs, comma-separated.
# The controller prefers VALKEY_SENTINEL_URLS over VALKEY_URL.
VALKEY_SENTINEL_URLS: "redis://valkey-sentinel-1:26379,redis://valkey-sentinel-2:26379,redis://valkey-sentinel-3:26379"
VALKEY_SENTINEL_MASTER: "mymaster"
services:
# ---------------------------------------------------------------
# PostgreSQL (use an external/managed DB for production)
# ---------------------------------------------------------------
postgres:
image: pgvector/pgvector:pg16
restart: unless-stopped
environment:
POSTGRES_USER: netstacks
POSTGRES_PASSWORD: ${DB_PASSWORD}
POSTGRES_DB: netstacks
volumes:
- postgres-data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U netstacks"]
interval: 5s
timeout: 5s
retries: 5
# ---------------------------------------------------------------
# Valkey Primary + Replicas
# ---------------------------------------------------------------
valkey-primary:
image: valkey/valkey:8-alpine
restart: unless-stopped
command: >
valkey-server
--appendonly yes
--maxmemory 512mb
--maxmemory-policy allkeys-lru
volumes:
- valkey-primary-data:/data
healthcheck:
test: ["CMD", "valkey-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
valkey-replica-1:
image: valkey/valkey:8-alpine
restart: unless-stopped
command: >
valkey-server --replicaof valkey-primary 6379 --appendonly yes
volumes:
- valkey-replica1-data:/data
depends_on:
valkey-primary:
condition: service_healthy
valkey-replica-2:
image: valkey/valkey:8-alpine
restart: unless-stopped
command: >
valkey-server --replicaof valkey-primary 6379 --appendonly yes
volumes:
- valkey-replica2-data:/data
depends_on:
valkey-primary:
condition: service_healthy
# ---------------------------------------------------------------
# Valkey Sentinels (3 for quorum). Each writes its own config and
# starts in sentinel mode — matches the shipped reference compose.
# ---------------------------------------------------------------
valkey-sentinel-1: &sentinel
image: valkey/valkey:8-alpine
restart: unless-stopped
command: >
sh -c '
echo "sentinel monitor mymaster valkey-primary 6379 2" > /tmp/sentinel.conf &&
echo "sentinel down-after-milliseconds mymaster 5000" >> /tmp/sentinel.conf &&
echo "sentinel failover-timeout mymaster 10000" >> /tmp/sentinel.conf &&
echo "sentinel parallel-syncs mymaster 1" >> /tmp/sentinel.conf &&
valkey-server /tmp/sentinel.conf --sentinel
'
depends_on:
valkey-primary:
condition: service_healthy
healthcheck:
test: ["CMD", "valkey-cli", "-p", "26379", "ping"]
interval: 5s
timeout: 3s
retries: 5
valkey-sentinel-2:
<<: *sentinel
valkey-sentinel-3:
<<: *sentinel
# ---------------------------------------------------------------
# Controller API instances (share the anchor above)
# ---------------------------------------------------------------
controller-1:
<<: *controller-common
environment:
<<: *controller-env
NETSTACKS_INSTANCE_NAME: "controller-1"
volumes:
- controller1-data:/data
controller-2:
<<: *controller-common
environment:
<<: *controller-env
NETSTACKS_INSTANCE_NAME: "controller-2"
volumes:
- controller2-data:/data
# ---------------------------------------------------------------
# Nginx reverse proxy (TLS termination)
# ---------------------------------------------------------------
nginx:
image: netstacks-nginx:latest
restart: unless-stopped
ports:
- "443:443"
volumes:
- ${TLS_CERT_PATH:-./nginx/server.pem}:/data/tls/server.pem:ro
- ${TLS_KEY_PATH:-./nginx/server-key.pem}:/data/tls/server-key.pem:ro
depends_on:
- controller-1
- controller-2
volumes:
postgres-data:
valkey-primary-data:
valkey-replica1-data:
valkey-replica2-data:
controller1-data:
controller2-data:VALKEY_SENTINEL_URLS is a comma-separated list of full redis://host:port URLs (note the redis:// prefix on each entry), and VALKEY_SENTINEL_MASTER must match the master name in your sentinel config (mymaster above). When VALKEY_SENTINEL_URLS is set it takes priority over VALKEY_URL.
Nginx Configuration
The Nginx config handles TLS termination, least_conn load balancing, and WebSocket upgrade passthrough. WebSocket locations use a long proxy_read_timeout to support long-running terminal sessions. Nginx proxies to the controllers over HTTPS, so use proxy_ssl_verify off when the backend uses the controller's auto-generated self-signed cert.
worker_processes auto;
events {
worker_connections 2048;
}
http {
# Upstream — controller instances
upstream controller_api {
least_conn;
server controller-1:3000;
server controller-2:3000;
}
# Connection upgrade map for WebSocket support
map $http_upgrade $connection_upgrade {
default upgrade;
'' close;
}
server {
listen 443 ssl;
server_name netstacks.example.net;
# TLS certificates (mounted into the container)
ssl_certificate /data/tls/server.pem;
ssl_certificate_key /data/tls/server-key.pem;
ssl_protocols TLSv1.2 TLSv1.3;
ssl_ciphers HIGH:!aNULL:!MD5;
ssl_prefer_server_ciphers on;
add_header Strict-Transport-Security "max-age=63072000; includeSubDomains" always;
add_header X-Content-Type-Options "nosniff" always;
add_header X-Frame-Options "DENY" always;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# The controller serves HTTPS internally with a self-signed cert
proxy_ssl_verify off;
# REST + admin API
location /api/ {
proxy_pass https://controller_api;
proxy_read_timeout 300s;
}
# WebSocket terminal sessions — long timeout, upgrade headers
location /ws/ {
proxy_pass https://controller_api;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_read_timeout 3600s;
proxy_send_timeout 3600s;
}
# Unauthenticated health/readiness for the load balancer
location /health {
proxy_pass https://controller_api;
access_log off;
}
location /ready {
proxy_pass https://controller_api;
access_log off;
}
location / {
proxy_pass https://controller_api;
}
}
}Start the HA Stack
# Ensure your .env file has all required secrets
cat .env
# VAULT_MASTER_KEY=<64-hex-chars>
# SERVICE_TOKEN_SECRET=<64-hex-chars>
# DB_PASSWORD=<database-password>
# Start the full stack
docker compose -f docker-compose.production.yml up -d
# Verify all containers are running
docker compose -f docker-compose.production.yml ps
# Basic health through the load balancer (status + version)
curl -k https://localhost/health
# { "status": "ok", "version": "0.0.14" }
# Verify cluster membership via the admin HA-status endpoint (requires a token)
TOKEN=$(curl -sk -X POST https://localhost/api/auth/login \
-H "Content-Type: application/json" \
-d '{"username":"admin","password":"<password>"}' | jq -r '.token')
curl -sk https://localhost/api/admin/ha-status \
-H "Authorization: Bearer $TOKEN" | jq .
# {
# "mode": "active-active",
# "this_instance": "<uuid>",
# "valkey_status": "connected",
# "instances": [ { "instance_id": "...", "advertise_addr": "https://...:3000",
# "active_sessions": 0, "is_self": true }, ... ],
# "total_sessions": 0
# }To add a third (or more) controller instance, add another service that reuses the controller-common anchor with a new name, instance name, and volume, then add the new container to the upstream block in Nginx. No Valkey or database changes are needed — additional instances register automatically via Valkey on startup.
Each instance advertises an address that peers use to route cross-instance sessions. By default this is https://{HOST}:{PORT}, which is the in-container address. If instances cannot reach each other at that address, set ADVERTISE_ADDR explicitly to a hostname/IP that other instances can reach (for example https://controller-1:3000).
External Database
For production HA deployments, use an external PostgreSQL instance (or managed service) instead of the Docker Compose PostgreSQL container. This provides database-level redundancy, automated backups, point-in-time recovery, and independent scaling.
PostgreSQL Setup
-- Connect to PostgreSQL as superuser
-- psql -h db.example.net -U postgres
-- Create the netstacks database and user
CREATE USER netstacks WITH PASSWORD 'your-secure-password';
CREATE DATABASE netstacks OWNER netstacks;
-- Connect to the netstacks database, then install pgvector
\c netstacks
CREATE EXTENSION IF NOT EXISTS vector;
-- Verify the extension is installed
SELECT extname, extversion FROM pg_extension WHERE extname = 'vector';
-- extname | extversion
-- ---------+------------
-- vector | 0.7.4
-- Grant schema permissions
GRANT ALL PRIVILEGES ON DATABASE netstacks TO netstacks;
GRANT ALL PRIVILEGES ON SCHEMA public TO netstacks;The pgvector extension is required for the AI knowledge base and semantic search features. If using a managed PostgreSQL service, verify that pgvector is available. AWS RDS (PostgreSQL 16+), Google Cloud SQL, and Azure Database for PostgreSQL Flexible Server all support pgvector. See Knowledge Base.
Connection Pooling with PgBouncer
For deployments with many concurrent sessions, add PgBouncer between the controller instances and PostgreSQL to pool database connections. Each controller instance maintains its own connection pool, so without PgBouncer, total connection count grows with the number of instances.
# pgbouncer.ini
[databases]
netstacks = host=db.example.net port=5432 dbname=netstacks
[pgbouncer]
listen_addr = 0.0.0.0
listen_port = 6432
auth_type = md5
auth_file = /etc/pgbouncer/userlist.txt
# Transaction pooling is recommended for NetStacks
pool_mode = transaction
max_client_conn = 200
default_pool_size = 25
min_pool_size = 5
reserve_pool_size = 5
# Timeouts
server_idle_timeout = 600
client_idle_timeout = 0
server_connect_timeout = 15# Add PgBouncer to your compose file
pgbouncer:
image: edoburu/pgbouncer:1.23.1-p2
restart: unless-stopped
volumes:
- ./pgbouncer.ini:/etc/pgbouncer/pgbouncer.ini:ro
- ./userlist.txt:/etc/pgbouncer/userlist.txt:ro
ports:
- "127.0.0.1:6432:6432"
depends_on:
- postgres
# Then update DATABASE_URL on each controller instance:
# DATABASE_URL: postgres://netstacks:<password>@pgbouncer:6432/netstacksPostgreSQL's default max_connections is 100. With two controller instances each using a pool of 25 connections, you need at least 50 connections plus overhead for maintenance and migrations. Raise max_connections in postgresql.conf, or use PgBouncer to multiplex many application connections over a smaller pool of database connections.
TLS Configuration
The NetStacks Controller always serves HTTPS — there is no plaintext HTTP mode and no on/off toggle. TLS behavior is driven by whether you provide a certificate and key. There are three common approaches.
Option 1: Auto-Generated Self-Signed CA (Default)
If you do not set both TLS_CERT_PATH and TLS_KEY_PATH, the controller generates a self-signed CA and server certificate on first boot, storing them under TLS_DATA_DIR (default /data/tls). The CA is reused across restarts so clients do not have to re-trust it. The certificate always includes localhost and 127.0.0.1; add more SANs (extra hostnames/IPs) with TLS_SANS. The Terminal app shows a certificate warning on first connection to an untrusted CA.
# Controller environment for auto-generated TLS
environment:
TLS_DATA_DIR: /data/tls
# Optional: extra Subject Alternative Names (comma-separated)
TLS_SANS: "netstacks.example.net,10.0.1.20"Option 2: Bring Your Own Certificate (Production)
For production, provide your own certificate and private key by setting both TLS_CERT_PATH and TLS_KEY_PATH and mounting the PEM files into the container. When both are set, the controller uses them instead of auto-generating.
# Generate a CSR and private key (if using an internal CA)
openssl req -new -newkey rsa:4096 -nodes \
-keyout netstacks.key \
-out netstacks.csr \
-subj "/CN=netstacks.example.net/O=Example Corp"
# After your CA signs the CSR you will have:
# netstacks.crt (signed certificate)
# ca-chain.crt (CA certificate chain)
# Build a full chain file:
cat netstacks.crt ca-chain.crt > fullchain.pem
# Verify the certificate
openssl x509 -in fullchain.pem -text -noout | head -20# Mount your certificates into the controller container
controller-1:
image: netstacks-controller:latest
environment:
TLS_CERT_PATH: /certs/fullchain.pem
TLS_KEY_PATH: /certs/netstacks.key
volumes:
- ./certs/fullchain.pem:/certs/fullchain.pem:ro
- ./certs/netstacks.key:/certs/netstacks.key:roOption 3: TLS Termination at Nginx (Recommended for HA)
In most HA deployments, TLS for external clients is terminated at the Nginx load balancer with a trusted (or Let's Encrypt) certificate, while the controllers keep their auto-generated self-signed certs internally. Nginx proxies to the backends with proxy_ssl_verify off.
# Obtain a Let's Encrypt certificate for the Nginx host
sudo certbot certonly --standalone \
-d netstacks.example.net \
--agree-tos --email [email protected] --non-interactive
# Certificate files:
# /etc/letsencrypt/live/netstacks.example.net/fullchain.pem
# /etc/letsencrypt/live/netstacks.example.net/privkey.pem
# Mount them where the nginx.conf expects server.pem / server-key.pem,
# or point ssl_certificate / ssl_certificate_key at the Let's Encrypt paths.
sudo systemctl enable --now certbot.timer # auto-renew
sudo certbot renew --dry-runEven with Nginx terminating TLS for clients, traffic between Nginx and the controllers remains encrypted — the controllers only speak HTTPS. External clients see the trusted Nginx certificate; internal hops use the controller's self-signed certs.
Geo-HA Considerations
For organizations that need controller availability across geographic regions, NetStacks can be deployed in a multi-region configuration. There are important trade-offs compared to a single-region HA deployment.
Managed PostgreSQL Across Regions
Use a managed PostgreSQL service with cross-region replication for database-level geo-redundancy:
- AWS RDS — Multi-AZ deployment with a cross-region read replica. In a failover scenario, promote the read replica to primary.
- Google Cloud SQL — Cross-region replicas with automatic failover enabled.
- Azure Database for PostgreSQL — Geo-redundant backup or read replicas in a secondary region.
Cross-region PostgreSQL replication is asynchronous by default, meaning there is a replication lag window during which the secondary region may serve stale data. Synchronous replication across regions eliminates this but adds significant write latency (typically 50–200ms per write). Most deployments should use asynchronous replication and accept the small consistency window.
Regional Valkey Clusters
Valkey does not natively support cross-region replication. For geo-HA, deploy an independent Valkey cluster (primary + replicas + sentinels) in each region. Each regional controller cluster uses its own Valkey cluster.
- Session state is regional — sessions created in Region A are tracked in Region A's Valkey cluster and served by Region A's controller instances.
- If Region A fails, sessions on Region A controllers are lost. Users reconnect to Region B via DNS failover and establish new sessions. Persistent data (users, devices, credentials) is available in Region B via the database replica.
DNS-Based Routing
Use DNS routing to direct users to the nearest or healthiest region. Health checks should poll the unauthenticated /health endpoint:
# Example: AWS Route 53 health-checked failover
# Primary record (us-east-1)
netstacks.example.net A 203.0.113.10 (failover: PRIMARY, health-check: /health)
# Secondary record (eu-west-1)
netstacks.example.net A 198.51.100.20 (failover: SECONDARY, health-check: /health)
# Alternative: latency-based routing (active-active geo)
netstacks.example.net A 203.0.113.10 (region: us-east-1, health-check: enabled)
netstacks.example.net A 198.51.100.20 (region: eu-west-1, health-check: enabled)Session Locality Limitations
- Sessions are not portable across regions. A terminal session opened in Region A cannot be seamlessly migrated to Region B. If Region A fails, the user must establish a new session in Region B.
- Session sharing is regional. Shared terminal sessions (a senior engineer watching a junior engineer's session) only work when both users connect to the same regional cluster. See Session Sharing.
- Device proximity matters. Route users to the controller region closest to their target devices. An SSH session proxied through a distant region adds round-trip latency to every keystroke.
For most organizations, an active-passive geo configuration is simpler and more reliable than active-active. Run the full HA stack in a primary region, with a warm standby in a secondary region (database replica + pre-deployed but stopped controller instances). Failover is a DNS change plus starting the standby controllers.
Environment Variable Reference
Reference of environment variables read by the NetStacks Controller. Variables marked Required must be set for the controller to start.
Core Configuration
| Variable | Description | Default | Required |
|---|---|---|---|
DATABASE_URL | PostgreSQL connection string, e.g. postgres://user:pass@host:5432/netstacks | — | Yes |
VAULT_MASTER_KEY | 64 hex character (32 byte) key for encryption of credentials and secrets. Must be identical across all instances | — | Yes |
SERVICE_TOKEN_SECRET | Secret used to sign internal service tokens | — | Yes |
HOST | Bind address for the API server | 0.0.0.0 | No |
PORT | Port the controller API (HTTPS) listens on | 3000 | No |
NETSTACKS_INSTANCE_NAME | Human-readable name shown in the Admin UI HA status | hostname | No |
ADVERTISE_ADDR | Address other instances use to reach this one. Override when behind a load balancer | https://{HOST}:{PORT} | No |
RUST_LOG | Log level filter, e.g. info,netstacks_api=info | info | No |
SERVE_STATIC | Whether the controller serves the bundled UI assets. Set true to serve them from the controller; leave unset (or any non-true value) when a separate proxy serves them | false | No |
RECORDING_DIR | Directory for session recordings | /var/lib/netstacks/recordings | No |
The JWT signing secret is not a controller environment variable. It is the auth.jwt_secret setting in PostgreSQL (HS256), configured from Admin → Settings. Because it lives in the shared database, all instances use the same value automatically.
TLS Configuration
| Variable | Description | Default |
|---|---|---|
TLS_CERT_PATH | Path to a PEM certificate. Set together with TLS_KEY_PATH to bring your own cert | — (auto-generate) |
TLS_KEY_PATH | Path to the PEM private key matching the certificate | — (auto-generate) |
TLS_DATA_DIR | Directory where the auto-generated CA and server cert are stored and reused across restarts | /data/tls |
TLS_SANS | Comma-separated extra Subject Alternative Names added to the auto-generated cert (in addition to localhost and detected IPs) | empty |
Valkey / HA Configuration
| Variable | Description | Default |
|---|---|---|
VALKEY_SENTINEL_URLS | Comma-separated full Sentinel URLs for automated failover, e.g. redis://s1:26379,redis://s2:26379,redis://s3:26379. Takes priority over VALKEY_URL | None (HA disabled) |
VALKEY_SENTINEL_MASTER | Sentinel master name for primary discovery. Must match your sentinel config | mymaster |
VALKEY_URL | Direct Valkey connection URL, e.g. redis://valkey:6379. Used only when Sentinel URLs are not set | None (HA disabled) |
If neither VALKEY_SENTINEL_URLS nor VALKEY_URL is set, the controller logs that it is running without HA support and operates in single-instance mode.
SSH Proxy Configuration
| Variable | Description | Default |
|---|---|---|
SSH_PROXY_ENABLED | Enable the bastion-style SSH proxy listener (true / 1 to enable) | false |
SSH_PROXY_ADDR | Bind address for the SSH proxy listener | 0.0.0.0:2222 |
Health Checks & Monitoring
The controller exposes health and status endpoints for load balancer checks, monitoring systems, and the Admin UI Dashboard.
Health Endpoints
| Endpoint | Auth | Purpose & response |
|---|---|---|
GET /healthGET /api/health | No | Lightweight liveness check. Returns { "status": "ok", "version": "..." }. Ideal for load balancer health checks. |
GET /readyGET /api/ready | No | Readiness probe. Tests the database (and Valkey if configured) and returns status, database, and valkey fields. |
GET /api/admin/health | Yes (admin) | Detailed system health with per-component status and latencies (database, SSH proxy, etc.). |
GET /api/admin/ha-status | Yes (admin) | Cluster mode, instance list, Valkey status, and session distribution. Backs the Dashboard HA Status widget. |
Basic Health Response
curl -k https://netstacks.example.net/health{
"status": "ok",
"version": "0.0.14"
}Readiness Response
{
"status": "ready",
"database": "connected",
"valkey": "connected"
}HA Status Response (admin)
{
"mode": "active-active",
"this_instance": "3f9c2d8a-1b6e-4a77-9c0b-2e5f1d8a4b7c",
"valkey_status": "connected",
"instances": [
{
"instance_id": "3f9c2d8a-1b6e-4a77-9c0b-2e5f1d8a4b7c",
"advertise_addr": "https://controller-1:3000",
"active_sessions": 12,
"is_self": true
},
{
"instance_id": "a1b2c3d4-e5f6-7890-abcd-ef0123456789",
"advertise_addr": "https://controller-2:3000",
"active_sessions": 9,
"is_self": false
}
],
"total_sessions": 21
}mode is "active-active" when more than one instance is registered and "single-instance" otherwise. valkey_status is one of "connected", "disconnected", or "not configured".
Monitoring with Prometheus
The controller does not expose a native Prometheus metrics endpoint. Use a JSON exporter to scrape /api/admin/ha-status (with a bearer token) or the unauthenticated /ready endpoint and convert the fields to metrics.
# json_exporter config (config.yml) — scrape /api/admin/ha-status
modules:
netstacks_ha:
headers:
Authorization: "Bearer ${NETSTACKS_TOKEN}"
metrics:
- name: netstacks_cluster_instances
path: '{ len(.instances) }'
help: "Number of registered controller instances"
- name: netstacks_sessions_total
path: '{ .total_sessions }'
help: "Total active sessions across the cluster"
- name: netstacks_valkey_connected
path: '{ .valkey_status }'
help: "Valkey connection status"
values:
connected: 1
disconnected: 0Alerting Script
#!/bin/bash
# Poll the unauthenticated readiness endpoint from a monitoring host
RESPONSE=$(curl -sk https://netstacks.example.net/ready)
STATUS=$(echo "$RESPONSE" | jq -r '.status')
DB=$(echo "$RESPONSE" | jq -r '.database')
if [ "$STATUS" != "ready" ]; then
echo "CRITICAL: controller not ready (db=$DB)"
exit 2
fi
echo "OK: ready (db=$DB)"
exit 0Configure your load balancer to poll /health every 10 seconds with a 5-second timeout, and remove an instance from the pool after 3 consecutive failures. The endpoint is lightweight and unauthenticated, making it safe for frequent polling.
Backup & Restore
PostgreSQL is the single source of truth for all persistent data in NetStacks (including the JWT signing secret). Regular database backups are essential. Valkey data is ephemeral (session state, heartbeats) and does not need backing up — it rebuilds automatically when instances restart.
Manual Database Backup
# Backup the full database (compressed)
pg_dump -h db.example.net -U netstacks -d netstacks \
--format=custom \
--compress=9 \
--file=netstacks-backup-$(date +%Y%m%d-%H%M%S).dump
# Backup only the schema (for documentation or migration reference)
pg_dump -h db.example.net -U netstacks -d netstacks \
--schema-only \
--file=netstacks-schema-$(date +%Y%m%d).sqlAutomated Backups
Add an automated backup container to your compose stack that runs daily backups with configurable retention.
db-backup:
image: prodrigestivill/postgres-backup-local:16
restart: unless-stopped
depends_on:
postgres:
condition: service_healthy
environment:
POSTGRES_HOST: postgres
POSTGRES_DB: netstacks
POSTGRES_USER: netstacks
POSTGRES_PASSWORD: ${DB_PASSWORD}
POSTGRES_EXTRA_OPTS: "--format=custom --compress=9"
SCHEDULE: "0 2 * * *" # daily at 02:00
BACKUP_KEEP_DAYS: 7
BACKUP_KEEP_WEEKS: 4
BACKUP_KEEP_MONTHS: 6
volumes:
- db-backups:/backups
volumes:
db-backups:Restore from Backup
# Stop the controller instances first
docker compose -f docker-compose.production.yml stop controller-1 controller-2
# Restore from a custom-format backup
pg_restore -h db.example.net -U netstacks -d netstacks \
--clean --if-exists --no-owner --no-privileges \
netstacks-backup-20260410-020000.dump
# Verify the restore
psql -h db.example.net -U netstacks -d netstacks \
-c "SELECT count(*) FROM users;"
# Restart the controller instances
docker compose -f docker-compose.production.yml start controller-1 controller-2Database backups contain encrypted credential data. To read those credentials after a restore, you must use the same VAULT_MASTER_KEY that was active when the data was encrypted. A backup restored with a different vault key will start successfully, but all encrypted credentials will be unreadable. Always back up the vault master key separately from the database.
Upgrading
NetStacks Controller supports rolling upgrades in HA mode. The controller applies database migrations on startup, so no manual migration step is needed.
Rolling Upgrade (HA Mode)
Upgrade one controller instance at a time to maintain availability. The load balancer health check removes the restarting instance and adds it back once /health returns 200.
# Pull (or rebuild) the new controller image first
docker compose -f docker-compose.production.yml pull controller-1 controller-2
# Step 1: upgrade controller-1
docker compose -f docker-compose.production.yml up -d --no-deps controller-1
# Wait for controller-1 to report healthy through the load balancer
until curl -sk https://localhost/health | grep -q '"status":"ok"'; do
echo "Waiting for controller-1..."
sleep 5
done
echo "controller-1 is healthy"
# Step 2: upgrade controller-2
docker compose -f docker-compose.production.yml up -d --no-deps controller-2
until curl -sk https://localhost/health | grep -q '"status":"ok"'; do
echo "Waiting for controller-2..."
sleep 5
done
echo "Upgrade complete — verify cluster membership:"
# Confirm both instances are registered (admin token required)
curl -sk https://localhost/api/admin/ha-status \
-H "Authorization: Bearer $TOKEN" | jq '.instances | length'Database migrations are applied automatically when a controller starts with a new version. Upgrade one instance at a time in a rolling fashion so the cluster always has a healthy member serving traffic.
Rollback
If an upgrade introduces issues, roll back by pinning the controller image to the previous tag and restarting.
# Pin the previous version (edit the compose file or use an env override)
# image: netstacks-controller:0.0.13
docker compose -f docker-compose.production.yml up -d --no-deps controller-1
sleep 10
docker compose -f docker-compose.production.yml up -d --no-deps controller-2
# Verify
curl -sk https://localhost/health | jq .If the new version applied migrations that alter existing tables (dropping columns, changing types), rolling back to the previous version may cause errors because the old code expects the old schema. Always read the release notes and take a database backup before major upgrades.
For zero-risk upgrades, deploy the new version as a separate set of controller instances pointing at the same database, verify health on the green stack, then switch the load balancer to the green instances. If issues arise, switch back to blue instantly.
Questions & Answers
- Does NetStacks Controller run active-active or active-passive?
- Active-active. Every controller instance serves all traffic simultaneously; there is no primary controller. The HA status endpoint reports
mode: "active-active"once more than one instance is registered in Valkey. - What is the controller health check URL?
GET /health(alsoGET /api/health), unauthenticated, returning{ "status": "ok", "version": "..." }. For a readiness probe that tests the database and Valkey, use/ready. There is no/api/v1/*path.- Do I need Valkey for a single controller?
- No. If neither
VALKEY_SENTINEL_URLSnorVALKEY_URLis set, the controller runs in single-instance mode. Valkey is required only when running multiple instances that must coordinate sessions. - Which secrets must match across instances?
VAULT_MASTER_KEYmust be identical on every instance.DATABASE_URLmust point all instances at the same database. The JWT signing secret (auth.jwt_secret) is stored in that shared database, so it is automatically consistent.- How do I set the JWT signing secret?
- It is not an environment variable. Set the
auth.jwt_secretsetting from Admin → Settings. Replace the default placeholder before production. - How is TLS enabled?
- The controller always serves HTTPS. If you do not set both
TLS_CERT_PATHandTLS_KEY_PATH, it auto-generates a CA and server cert underTLS_DATA_DIR(default/data/tls). Add extra SANs withTLS_SANS. There is noTLS_ENABLEDtoggle. - What is the correct Sentinel environment variable?
VALKEY_SENTINEL_URLS(plural) — a comma-separated list of fullredis://host:portURLs — together withVALKEY_SENTINEL_MASTER(the master name, defaultmymaster).- Where do I see cluster health in the UI?
- The Admin UI Dashboard includes an HA Status widget showing mode, instances, Valkey connectivity, and session distribution, backed by
GET /api/admin/ha-status.
Related Documentation
- Installation — install the Controller and Terminal app.
- Requirements — supported platforms and sizing.
- Settings — set
auth.jwt_secretand other controller settings. - Authentication — LDAP, OIDC, and password auth for the cluster.
- User Management and Roles & Permissions — control access across the deployment.
- Audit Logs and Activity Monitor — observe cluster activity.
- Credential Vault — how the
VAULT_MASTER_KEYprotects stored secrets. - Session Sharing — behavior and regional limitations across instances.
- Knowledge Base — the pgvector-backed feature that requires the extension.