NetStacksNetStacks

High Availability Deployment

Enterprise

Deploy NetStacks Controller in active-active HA with multiple instances, Valkey Sentinel, external PostgreSQL, Nginx load balancing, TLS, and backups.

Overview

NetStacks Controller supports active-active high availability, meaning every controller instance actively serves requests simultaneously. There is no primary/standby distinction among controller nodes — any instance can handle any API call, WebSocket connection, or SSH proxy session. The cluster is glued together by a shared PostgreSQL database, a shared VAULT_MASTER_KEY, and Valkey for ephemeral session coordination.

The HA architecture eliminates single points of failure across the stack:

  • Controller instances — Two or more instances run behind a load balancer. If one goes down, the remaining instances continue serving all traffic.
  • Valkey (session coordination) — A Redis-compatible in-memory store that tracks which controller instance owns each session, enabling cross-instance session routing. Valkey Sentinel provides automated failover for the Valkey tier itself.
  • PostgreSQL — The database stores all persistent state: users, devices, credentials, audit logs, certificates, and configuration (including the JWT signing secret). In HA mode, use an external PostgreSQL cluster with replication for database-level redundancy.
  • Load balancer — Nginx (or any WebSocket-aware load balancer) distributes traffic across controller instances using least_conn routing, with health-check-based removal of unhealthy nodes.

Capabilities enabled when Valkey is configured:

  • Instance registration and heartbeat — Each controller generates a random instance ID at startup and registers in Valkey with a periodic heartbeat. Stale instances are cleaned up automatically.
  • Session registry — Active sessions are tracked in Valkey with the owning instance's advertise address, enabling cross-instance session lookup and routing.
  • HA status — The Admin UI Dashboard includes an HA Status widget showing cluster mode, instance list, session distribution, and Valkey connectivity. It is backed by the GET /api/admin/ha-status endpoint.
Enterprise feature

High availability is an Enterprise Controller capability. Single-instance Controller deployments do not require Valkey or multiple controller instances — without Valkey configured, the controller simply runs in single-instance mode (mode: "single-instance").

Architecture

The following diagram shows a production HA deployment. The load balancer receives all external traffic and distributes it across the controller instances. Both controllers share the same PostgreSQL database and Valkey cluster for state coordination.

ha-architecture.txttext
                        +----------------------+
                        |   Terminal Clients   |
                        |  (Tauri desktop app) |
                        +----------+-----------+
                                   |
                              HTTPS / WSS
                                   |
                        +----------v-----------+
                        |  Nginx Load Balancer |
                        |  (least_conn, TLS    |
                        |   termination)       |
                        +----+------------+----+
                             |            |
                    +--------v--+    +----v--------+
                    |Controller |    | Controller  |
                    |Instance 1 |    | Instance 2  |
                    |  :3000    |    |   :3000     |
                    +--+----+---+    +--+----+-----+
                       |    |          |    |
              +--------+    +----+-----+    +--------+
              |                 |                    |
    +---------v--------+ +------v---------+ +--------v-------+
    |   PostgreSQL 16  | | Valkey Primary | | Network Devices|
    |   (pgvector ext) | |   + Replicas   | | (SSH/Telnet)   |
    |                  | |                | +----------------+
    |  Users, devices, | | Session state, |
    |  credentials,    | | instance reg,  |
    |  audit logs,     | | heartbeats     |
    |  jwt secret      | |                |
    +------------------+ +-------+--------+
                                 |
                        +--------v--------+
                        | Valkey Sentinel |
                        |  (3 instances)  |
                        | Monitors        |
                        | primary, auto-  |
                        | promotes replica|
                        +-----------------+

In this architecture:

  • Nginx terminates TLS and distributes traffic to controller instances using least_conn. It passes WebSocket upgrades through to the backend and proxies to the controllers over HTTPS (proxy_pass https://...).
  • Controller instances are stateless application servers. All persistent state lives in PostgreSQL; all ephemeral session state lives in Valkey. Scale horizontally by adding more instances.
  • Valkey Primary + Replicas handle session coordination. The primary handles writes; replicas provide read scaling and failover candidates.
  • Valkey Sentinel monitors the primary and automatically promotes a replica if the primary fails. Three sentinels are required for quorum.
  • PostgreSQL is the single source of truth for all persistent data. Use an external PostgreSQL cluster with streaming replication for database-level HA.
Reference compose ships with the Controller

The Enterprise Controller repository ships a reference HA compose file (docker-compose.production.yml) that builds local images named netstacks-controller and netstacks-nginx. The examples below mirror that shipped configuration. Substitute your own image names/registry as appropriate for your environment.

Prerequisites

Before deploying NetStacks Controller in any configuration, ensure the following requirements are met:

RequirementMinimumRecommended (HA)
Docker Engine24.0+25.0+
Docker Composev2.20+v2.24+
RAM per controller instance4 GB8 GB
CPU per controller instance2 vCPU4 vCPU
PostgreSQL16+ with pgvector16+ with pgvector, streaming replication
Valkey / RedisNot required (single instance)Valkey 8+ with Sentinel
Disk (database)20 GB100 GB+ SSD
Disk (Valkey)1 GB4 GB SSD (AOF persistence)
NetworkController reaches devices on SSH/Telnet portsLow-latency interconnect between controller instances
LicenseEnterprise Controller (Contact Sales)Enterprise Controller (Contact Sales)
Enterprise provisioning

The Enterprise Controller is provisioned through a sales engagement rather than a self-service license key. There is no per-instance seat key to copy between controllers — all instances in a cluster share one database and run as one logical deployment. Contact NetStacks to obtain the Enterprise Controller.

Cloud deployments

On AWS, use RDS for PostgreSQL with the pgvector extension (supported on RDS 16+) and ElastiCache for Valkey. On GCP, use Cloud SQL for PostgreSQL and Memorystore for Redis/Valkey. On Azure, use Azure Database for PostgreSQL Flexible Server and Azure Cache for Redis.

Single Instance Setup

Start with a single-instance deployment to verify your configuration before scaling to HA. This setup runs PostgreSQL and the controller API in a single Docker Compose stack. Valkey is optional — without it, the controller runs in single-instance mode.

Generate Secrets

Before creating the Docker Compose file, generate the required secrets. These values must be kept secure and consistent across all controller instances if you later scale to HA.

generate-secrets.shbash
# Generate the vault master key (64 hex characters = 32 bytes)
# This encrypts all credentials, SSH CA keys, and sensitive data in the database
export VAULT_MASTER_KEY=$(openssl rand -hex 32)
echo "VAULT_MASTER_KEY=$VAULT_MASTER_KEY"

# Generate the service-token secret (used for internal service-to-service auth)
export SERVICE_TOKEN_SECRET=$(openssl rand -hex 32)
echo "SERVICE_TOKEN_SECRET=$SERVICE_TOKEN_SECRET"

# Generate the database password
export DB_PASSWORD=$(openssl rand -base64 24)
echo "DB_PASSWORD=$DB_PASSWORD"

# Save these values securely. You will need VAULT_MASTER_KEY for every
# controller instance in HA mode.
Store the vault master key securely

The VAULT_MASTER_KEY is the root of trust for all encrypted data in NetStacks. If lost, encrypted credentials and SSH CA private keys cannot be recovered. Store it in a secrets manager (HashiCorp Vault, AWS Secrets Manager, etc.) or at minimum in a secure, backed-up location outside the Docker host.

The JWT signing secret lives in the database

Unlike the vault master key, the JWT signing secret is not a controller environment variable. It is stored as the auth.jwt_secret setting in PostgreSQL and configured from Admin → Settings. Because it lives in the shared database, every controller instance automatically uses the same signing key — tokens issued by one instance are accepted by all others. Change it from the default placeholder before going to production.

Docker Compose (Single Instance)

docker-compose.ymlyaml
services:
  postgres:
    image: pgvector/pgvector:pg16
    restart: unless-stopped
    environment:
      POSTGRES_USER: netstacks
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_DB: netstacks
    volumes:
      - postgres-data:/var/lib/postgresql/data
    ports:
      - "127.0.0.1:5432:5432"
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U netstacks"]
      interval: 5s
      timeout: 5s
      retries: 5

  controller:
    image: netstacks-controller:latest
    restart: unless-stopped
    depends_on:
      postgres:
        condition: service_healthy
    ports:
      - "3000:3000"
    environment:
      # Database
      DATABASE_URL: postgres://netstacks:${DB_PASSWORD}@postgres:5432/netstacks

      # Security — generate with: openssl rand -hex 32
      VAULT_MASTER_KEY: ${VAULT_MASTER_KEY}
      SERVICE_TOKEN_SECRET: ${SERVICE_TOKEN_SECRET}

      # Server bind (HOST defaults to 0.0.0.0, PORT defaults to 3000)
      HOST: 0.0.0.0
      PORT: 3000

      # Instance naming (optional, shown in the Admin UI HA status)
      NETSTACKS_INSTANCE_NAME: "controller-1"

      # TLS — the controller always serves HTTPS. With no cert/key paths set it
      # auto-generates a CA + server cert under TLS_DATA_DIR on first boot.
      TLS_DATA_DIR: /data/tls

      # Logging
      RUST_LOG: "info,netstacks_api=info"
    volumes:
      - controller-data:/data

volumes:
  postgres-data:
  controller-data:

Start the Stack

start-single-instance.shbash
# Create a .env file with your generated secrets
cat > .env << 'EOF'
VAULT_MASTER_KEY=<your-64-hex-char-key>
SERVICE_TOKEN_SECRET=<your-64-hex-char-secret>
DB_PASSWORD=<your-database-password>
EOF

# Start the stack
docker compose up -d

# Watch the logs for the initial admin password
docker compose logs -f controller

Verify the Deployment

The basic health endpoint is unauthenticated and returns the service status and version. It is available both at the root (legacy callers) and under /api (for callers using the/api base URL):

verify-deployment.shbash
# Basic health check (use -k for the self-signed cert). Both paths work:
curl -k https://localhost:3000/health
curl -k https://localhost:3000/api/health

# Expected response:
# { "status": "ok", "version": "0.0.14" }

# Readiness probe — checks the database (and Valkey if configured)
curl -k https://localhost:3000/ready
# { "status": "ready", "database": "connected", "valkey": "not configured" }

# Log in and get a token
curl -k -X POST https://localhost:3000/api/auth/login \
  -H "Content-Type: application/json" \
  -d '{"username": "admin", "password": "<password-from-logs>"}'
After first login

Log in to the Admin UI at https://your-host:3000, change the default admin password, and set a strong auth.jwt_secret under Admin → Settings before exposing the controller. See Settings and User Management.

High Availability Setup

The full HA deployment adds Valkey with Sentinel for session coordination, multiple controller instances, and an Nginx load balancer. All controller instances must share the same DATABASE_URL and VAULT_MASTER_KEY. The JWT signing secret is shared automatically because it lives in the database.

Shared vault key is critical

Every controller instance in the cluster must use an identical VAULT_MASTER_KEY. If keys differ between instances, credentials encrypted on one instance fail to decrypt on another. Generate the key once and distribute it to all instances via your secrets management system. The auth.jwt_secret setting does not need manual distribution — all instances read it from the shared database.

Full HA Docker Compose

This mirrors the shipped docker-compose.production.yml reference: PostgreSQL with pgvector, a Valkey primary plus two replicas plus three Sentinels, two controller instances sharing a YAML anchor, and an Nginx reverse proxy that terminates TLS.

docker-compose.production.ymlyaml
x-controller-common: &controller-common
  image: netstacks-controller:latest
  restart: unless-stopped
  depends_on:
    postgres:
      condition: service_healthy
    valkey-sentinel-1:
      condition: service_healthy
  environment: &controller-env
    DATABASE_URL: postgres://netstacks:${DB_PASSWORD}@postgres:5432/netstacks
    VAULT_MASTER_KEY: ${VAULT_MASTER_KEY}
    SERVICE_TOKEN_SECRET: ${SERVICE_TOKEN_SECRET}
    HOST: 0.0.0.0
    PORT: 3000
    RUST_LOG: ${RUST_LOG:-info,netstacks_api=info}
    SERVE_STATIC: "false"
    TLS_DATA_DIR: /data/tls

    # Valkey via Sentinel — full redis:// URLs, comma-separated.
    # The controller prefers VALKEY_SENTINEL_URLS over VALKEY_URL.
    VALKEY_SENTINEL_URLS: "redis://valkey-sentinel-1:26379,redis://valkey-sentinel-2:26379,redis://valkey-sentinel-3:26379"
    VALKEY_SENTINEL_MASTER: "mymaster"

services:
  # ---------------------------------------------------------------
  # PostgreSQL (use an external/managed DB for production)
  # ---------------------------------------------------------------
  postgres:
    image: pgvector/pgvector:pg16
    restart: unless-stopped
    environment:
      POSTGRES_USER: netstacks
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_DB: netstacks
    volumes:
      - postgres-data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U netstacks"]
      interval: 5s
      timeout: 5s
      retries: 5

  # ---------------------------------------------------------------
  # Valkey Primary + Replicas
  # ---------------------------------------------------------------
  valkey-primary:
    image: valkey/valkey:8-alpine
    restart: unless-stopped
    command: >
      valkey-server
      --appendonly yes
      --maxmemory 512mb
      --maxmemory-policy allkeys-lru
    volumes:
      - valkey-primary-data:/data
    healthcheck:
      test: ["CMD", "valkey-cli", "ping"]
      interval: 5s
      timeout: 3s
      retries: 5

  valkey-replica-1:
    image: valkey/valkey:8-alpine
    restart: unless-stopped
    command: >
      valkey-server --replicaof valkey-primary 6379 --appendonly yes
    volumes:
      - valkey-replica1-data:/data
    depends_on:
      valkey-primary:
        condition: service_healthy

  valkey-replica-2:
    image: valkey/valkey:8-alpine
    restart: unless-stopped
    command: >
      valkey-server --replicaof valkey-primary 6379 --appendonly yes
    volumes:
      - valkey-replica2-data:/data
    depends_on:
      valkey-primary:
        condition: service_healthy

  # ---------------------------------------------------------------
  # Valkey Sentinels (3 for quorum). Each writes its own config and
  # starts in sentinel mode — matches the shipped reference compose.
  # ---------------------------------------------------------------
  valkey-sentinel-1: &sentinel
    image: valkey/valkey:8-alpine
    restart: unless-stopped
    command: >
      sh -c '
        echo "sentinel monitor mymaster valkey-primary 6379 2" > /tmp/sentinel.conf &&
        echo "sentinel down-after-milliseconds mymaster 5000" >> /tmp/sentinel.conf &&
        echo "sentinel failover-timeout mymaster 10000" >> /tmp/sentinel.conf &&
        echo "sentinel parallel-syncs mymaster 1" >> /tmp/sentinel.conf &&
        valkey-server /tmp/sentinel.conf --sentinel
      '
    depends_on:
      valkey-primary:
        condition: service_healthy
    healthcheck:
      test: ["CMD", "valkey-cli", "-p", "26379", "ping"]
      interval: 5s
      timeout: 3s
      retries: 5

  valkey-sentinel-2:
    <<: *sentinel

  valkey-sentinel-3:
    <<: *sentinel

  # ---------------------------------------------------------------
  # Controller API instances (share the anchor above)
  # ---------------------------------------------------------------
  controller-1:
    <<: *controller-common
    environment:
      <<: *controller-env
      NETSTACKS_INSTANCE_NAME: "controller-1"
    volumes:
      - controller1-data:/data

  controller-2:
    <<: *controller-common
    environment:
      <<: *controller-env
      NETSTACKS_INSTANCE_NAME: "controller-2"
    volumes:
      - controller2-data:/data

  # ---------------------------------------------------------------
  # Nginx reverse proxy (TLS termination)
  # ---------------------------------------------------------------
  nginx:
    image: netstacks-nginx:latest
    restart: unless-stopped
    ports:
      - "443:443"
    volumes:
      - ${TLS_CERT_PATH:-./nginx/server.pem}:/data/tls/server.pem:ro
      - ${TLS_KEY_PATH:-./nginx/server-key.pem}:/data/tls/server-key.pem:ro
    depends_on:
      - controller-1
      - controller-2

volumes:
  postgres-data:
  valkey-primary-data:
  valkey-replica1-data:
  valkey-replica2-data:
  controller1-data:
  controller2-data:
Sentinel URL format

VALKEY_SENTINEL_URLS is a comma-separated list of full redis://host:port URLs (note the redis:// prefix on each entry), and VALKEY_SENTINEL_MASTER must match the master name in your sentinel config (mymaster above). When VALKEY_SENTINEL_URLS is set it takes priority over VALKEY_URL.

Nginx Configuration

The Nginx config handles TLS termination, least_conn load balancing, and WebSocket upgrade passthrough. WebSocket locations use a long proxy_read_timeout to support long-running terminal sessions. Nginx proxies to the controllers over HTTPS, so use proxy_ssl_verify off when the backend uses the controller's auto-generated self-signed cert.

nginx.confnginx
worker_processes auto;

events {
    worker_connections 2048;
}

http {
    # Upstream — controller instances
    upstream controller_api {
        least_conn;
        server controller-1:3000;
        server controller-2:3000;
    }

    # Connection upgrade map for WebSocket support
    map $http_upgrade $connection_upgrade {
        default upgrade;
        ''      close;
    }

    server {
        listen 443 ssl;
        server_name netstacks.example.net;

        # TLS certificates (mounted into the container)
        ssl_certificate     /data/tls/server.pem;
        ssl_certificate_key /data/tls/server-key.pem;
        ssl_protocols       TLSv1.2 TLSv1.3;
        ssl_ciphers         HIGH:!aNULL:!MD5;
        ssl_prefer_server_ciphers on;

        add_header Strict-Transport-Security "max-age=63072000; includeSubDomains" always;
        add_header X-Content-Type-Options "nosniff" always;
        add_header X-Frame-Options "DENY" always;

        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        # The controller serves HTTPS internally with a self-signed cert
        proxy_ssl_verify off;

        # REST + admin API
        location /api/ {
            proxy_pass https://controller_api;
            proxy_read_timeout 300s;
        }

        # WebSocket terminal sessions — long timeout, upgrade headers
        location /ws/ {
            proxy_pass https://controller_api;
            proxy_set_header Upgrade $http_upgrade;
            proxy_set_header Connection $connection_upgrade;
            proxy_read_timeout 3600s;
            proxy_send_timeout 3600s;
        }

        # Unauthenticated health/readiness for the load balancer
        location /health {
            proxy_pass https://controller_api;
            access_log off;
        }
        location /ready {
            proxy_pass https://controller_api;
            access_log off;
        }

        location / {
            proxy_pass https://controller_api;
        }
    }
}

Start the HA Stack

start-ha-stack.shbash
# Ensure your .env file has all required secrets
cat .env
# VAULT_MASTER_KEY=<64-hex-chars>
# SERVICE_TOKEN_SECRET=<64-hex-chars>
# DB_PASSWORD=<database-password>

# Start the full stack
docker compose -f docker-compose.production.yml up -d

# Verify all containers are running
docker compose -f docker-compose.production.yml ps

# Basic health through the load balancer (status + version)
curl -k https://localhost/health
# { "status": "ok", "version": "0.0.14" }

# Verify cluster membership via the admin HA-status endpoint (requires a token)
TOKEN=$(curl -sk -X POST https://localhost/api/auth/login \
  -H "Content-Type: application/json" \
  -d '{"username":"admin","password":"<password>"}' | jq -r '.token')

curl -sk https://localhost/api/admin/ha-status \
  -H "Authorization: Bearer $TOKEN" | jq .
# {
#   "mode": "active-active",
#   "this_instance": "<uuid>",
#   "valkey_status": "connected",
#   "instances": [ { "instance_id": "...", "advertise_addr": "https://...:3000",
#                    "active_sessions": 0, "is_self": true }, ... ],
#   "total_sessions": 0
# }
Scaling beyond two instances

To add a third (or more) controller instance, add another service that reuses the controller-common anchor with a new name, instance name, and volume, then add the new container to the upstream block in Nginx. No Valkey or database changes are needed — additional instances register automatically via Valkey on startup.

ADVERTISE_ADDR behind a load balancer

Each instance advertises an address that peers use to route cross-instance sessions. By default this is https://{HOST}:{PORT}, which is the in-container address. If instances cannot reach each other at that address, set ADVERTISE_ADDR explicitly to a hostname/IP that other instances can reach (for example https://controller-1:3000).

External Database

For production HA deployments, use an external PostgreSQL instance (or managed service) instead of the Docker Compose PostgreSQL container. This provides database-level redundancy, automated backups, point-in-time recovery, and independent scaling.

PostgreSQL Setup

setup-database.sqlsql
-- Connect to PostgreSQL as superuser
-- psql -h db.example.net -U postgres

-- Create the netstacks database and user
CREATE USER netstacks WITH PASSWORD 'your-secure-password';
CREATE DATABASE netstacks OWNER netstacks;

-- Connect to the netstacks database, then install pgvector
\c netstacks
CREATE EXTENSION IF NOT EXISTS vector;

-- Verify the extension is installed
SELECT extname, extversion FROM pg_extension WHERE extname = 'vector';
--  extname | extversion
-- ---------+------------
--  vector  | 0.7.4

-- Grant schema permissions
GRANT ALL PRIVILEGES ON DATABASE netstacks TO netstacks;
GRANT ALL PRIVILEGES ON SCHEMA public TO netstacks;
pgvector is required

The pgvector extension is required for the AI knowledge base and semantic search features. If using a managed PostgreSQL service, verify that pgvector is available. AWS RDS (PostgreSQL 16+), Google Cloud SQL, and Azure Database for PostgreSQL Flexible Server all support pgvector. See Knowledge Base.

Connection Pooling with PgBouncer

For deployments with many concurrent sessions, add PgBouncer between the controller instances and PostgreSQL to pool database connections. Each controller instance maintains its own connection pool, so without PgBouncer, total connection count grows with the number of instances.

pgbouncer.initext
# pgbouncer.ini
[databases]
netstacks = host=db.example.net port=5432 dbname=netstacks

[pgbouncer]
listen_addr = 0.0.0.0
listen_port = 6432
auth_type = md5
auth_file = /etc/pgbouncer/userlist.txt

# Transaction pooling is recommended for NetStacks
pool_mode = transaction
max_client_conn = 200
default_pool_size = 25
min_pool_size = 5
reserve_pool_size = 5

# Timeouts
server_idle_timeout = 600
client_idle_timeout = 0
server_connect_timeout = 15
pgbouncer-compose.ymlyaml
# Add PgBouncer to your compose file
  pgbouncer:
    image: edoburu/pgbouncer:1.23.1-p2
    restart: unless-stopped
    volumes:
      - ./pgbouncer.ini:/etc/pgbouncer/pgbouncer.ini:ro
      - ./userlist.txt:/etc/pgbouncer/userlist.txt:ro
    ports:
      - "127.0.0.1:6432:6432"
    depends_on:
      - postgres

# Then update DATABASE_URL on each controller instance:
# DATABASE_URL: postgres://netstacks:<password>@pgbouncer:6432/netstacks
Connection limits

PostgreSQL's default max_connections is 100. With two controller instances each using a pool of 25 connections, you need at least 50 connections plus overhead for maintenance and migrations. Raise max_connections in postgresql.conf, or use PgBouncer to multiplex many application connections over a smaller pool of database connections.

TLS Configuration

The NetStacks Controller always serves HTTPS — there is no plaintext HTTP mode and no on/off toggle. TLS behavior is driven by whether you provide a certificate and key. There are three common approaches.

Option 1: Auto-Generated Self-Signed CA (Default)

If you do not set both TLS_CERT_PATH and TLS_KEY_PATH, the controller generates a self-signed CA and server certificate on first boot, storing them under TLS_DATA_DIR (default /data/tls). The CA is reused across restarts so clients do not have to re-trust it. The certificate always includes localhost and 127.0.0.1; add more SANs (extra hostnames/IPs) with TLS_SANS. The Terminal app shows a certificate warning on first connection to an untrusted CA.

tls-auto-generate.ymlyaml
# Controller environment for auto-generated TLS
environment:
  TLS_DATA_DIR: /data/tls
  # Optional: extra Subject Alternative Names (comma-separated)
  TLS_SANS: "netstacks.example.net,10.0.1.20"

Option 2: Bring Your Own Certificate (Production)

For production, provide your own certificate and private key by setting both TLS_CERT_PATH and TLS_KEY_PATH and mounting the PEM files into the container. When both are set, the controller uses them instead of auto-generating.

generate-cert.shbash
# Generate a CSR and private key (if using an internal CA)
openssl req -new -newkey rsa:4096 -nodes \
  -keyout netstacks.key \
  -out netstacks.csr \
  -subj "/CN=netstacks.example.net/O=Example Corp"

# After your CA signs the CSR you will have:
#   netstacks.crt  (signed certificate)
#   ca-chain.crt   (CA certificate chain)
# Build a full chain file:
cat netstacks.crt ca-chain.crt > fullchain.pem

# Verify the certificate
openssl x509 -in fullchain.pem -text -noout | head -20
tls-byoc.ymlyaml
# Mount your certificates into the controller container
  controller-1:
    image: netstacks-controller:latest
    environment:
      TLS_CERT_PATH: /certs/fullchain.pem
      TLS_KEY_PATH: /certs/netstacks.key
    volumes:
      - ./certs/fullchain.pem:/certs/fullchain.pem:ro
      - ./certs/netstacks.key:/certs/netstacks.key:ro

Option 3: TLS Termination at Nginx (Recommended for HA)

In most HA deployments, TLS for external clients is terminated at the Nginx load balancer with a trusted (or Let's Encrypt) certificate, while the controllers keep their auto-generated self-signed certs internally. Nginx proxies to the backends with proxy_ssl_verify off.

letsencrypt-setup.shbash
# Obtain a Let's Encrypt certificate for the Nginx host
sudo certbot certonly --standalone \
  -d netstacks.example.net \
  --agree-tos --email [email protected] --non-interactive

# Certificate files:
#   /etc/letsencrypt/live/netstacks.example.net/fullchain.pem
#   /etc/letsencrypt/live/netstacks.example.net/privkey.pem

# Mount them where the nginx.conf expects server.pem / server-key.pem,
# or point ssl_certificate / ssl_certificate_key at the Let's Encrypt paths.
sudo systemctl enable --now certbot.timer   # auto-renew
sudo certbot renew --dry-run
Internal traffic stays encrypted

Even with Nginx terminating TLS for clients, traffic between Nginx and the controllers remains encrypted — the controllers only speak HTTPS. External clients see the trusted Nginx certificate; internal hops use the controller's self-signed certs.

Geo-HA Considerations

For organizations that need controller availability across geographic regions, NetStacks can be deployed in a multi-region configuration. There are important trade-offs compared to a single-region HA deployment.

Managed PostgreSQL Across Regions

Use a managed PostgreSQL service with cross-region replication for database-level geo-redundancy:

  • AWS RDS — Multi-AZ deployment with a cross-region read replica. In a failover scenario, promote the read replica to primary.
  • Google Cloud SQL — Cross-region replicas with automatic failover enabled.
  • Azure Database for PostgreSQL — Geo-redundant backup or read replicas in a secondary region.
Write latency

Cross-region PostgreSQL replication is asynchronous by default, meaning there is a replication lag window during which the secondary region may serve stale data. Synchronous replication across regions eliminates this but adds significant write latency (typically 50–200ms per write). Most deployments should use asynchronous replication and accept the small consistency window.

Regional Valkey Clusters

Valkey does not natively support cross-region replication. For geo-HA, deploy an independent Valkey cluster (primary + replicas + sentinels) in each region. Each regional controller cluster uses its own Valkey cluster.

  • Session state is regional — sessions created in Region A are tracked in Region A's Valkey cluster and served by Region A's controller instances.
  • If Region A fails, sessions on Region A controllers are lost. Users reconnect to Region B via DNS failover and establish new sessions. Persistent data (users, devices, credentials) is available in Region B via the database replica.

DNS-Based Routing

Use DNS routing to direct users to the nearest or healthiest region. Health checks should poll the unauthenticated /health endpoint:

dns-routing-example.txttext
# Example: AWS Route 53 health-checked failover
# Primary record (us-east-1)
netstacks.example.net  A  203.0.113.10  (failover: PRIMARY, health-check: /health)

# Secondary record (eu-west-1)
netstacks.example.net  A  198.51.100.20 (failover: SECONDARY, health-check: /health)

# Alternative: latency-based routing (active-active geo)
netstacks.example.net  A  203.0.113.10  (region: us-east-1, health-check: enabled)
netstacks.example.net  A  198.51.100.20 (region: eu-west-1, health-check: enabled)

Session Locality Limitations

  • Sessions are not portable across regions. A terminal session opened in Region A cannot be seamlessly migrated to Region B. If Region A fails, the user must establish a new session in Region B.
  • Session sharing is regional. Shared terminal sessions (a senior engineer watching a junior engineer's session) only work when both users connect to the same regional cluster. See Session Sharing.
  • Device proximity matters. Route users to the controller region closest to their target devices. An SSH session proxied through a distant region adds round-trip latency to every keystroke.
Recommended geo-HA pattern

For most organizations, an active-passive geo configuration is simpler and more reliable than active-active. Run the full HA stack in a primary region, with a warm standby in a secondary region (database replica + pre-deployed but stopped controller instances). Failover is a DNS change plus starting the standby controllers.

Environment Variable Reference

Reference of environment variables read by the NetStacks Controller. Variables marked Required must be set for the controller to start.

Core Configuration

VariableDescriptionDefaultRequired
DATABASE_URLPostgreSQL connection string, e.g. postgres://user:pass@host:5432/netstacks—Yes
VAULT_MASTER_KEY64 hex character (32 byte) key for encryption of credentials and secrets. Must be identical across all instances—Yes
SERVICE_TOKEN_SECRETSecret used to sign internal service tokens—Yes
HOSTBind address for the API server0.0.0.0No
PORTPort the controller API (HTTPS) listens on3000No
NETSTACKS_INSTANCE_NAMEHuman-readable name shown in the Admin UI HA statushostnameNo
ADVERTISE_ADDRAddress other instances use to reach this one. Override when behind a load balancerhttps://{HOST}:{PORT}No
RUST_LOGLog level filter, e.g. info,netstacks_api=infoinfoNo
SERVE_STATICWhether the controller serves the bundled UI assets. Set true to serve them from the controller; leave unset (or any non-true value) when a separate proxy serves themfalseNo
RECORDING_DIRDirectory for session recordings/var/lib/netstacks/recordingsNo
JWT signing secret is a database setting

The JWT signing secret is not a controller environment variable. It is the auth.jwt_secret setting in PostgreSQL (HS256), configured from Admin → Settings. Because it lives in the shared database, all instances use the same value automatically.

TLS Configuration

VariableDescriptionDefault
TLS_CERT_PATHPath to a PEM certificate. Set together with TLS_KEY_PATH to bring your own cert— (auto-generate)
TLS_KEY_PATHPath to the PEM private key matching the certificate— (auto-generate)
TLS_DATA_DIRDirectory where the auto-generated CA and server cert are stored and reused across restarts/data/tls
TLS_SANSComma-separated extra Subject Alternative Names added to the auto-generated cert (in addition to localhost and detected IPs)empty

Valkey / HA Configuration

VariableDescriptionDefault
VALKEY_SENTINEL_URLSComma-separated full Sentinel URLs for automated failover, e.g. redis://s1:26379,redis://s2:26379,redis://s3:26379. Takes priority over VALKEY_URLNone (HA disabled)
VALKEY_SENTINEL_MASTERSentinel master name for primary discovery. Must match your sentinel configmymaster
VALKEY_URLDirect Valkey connection URL, e.g. redis://valkey:6379. Used only when Sentinel URLs are not setNone (HA disabled)

If neither VALKEY_SENTINEL_URLS nor VALKEY_URL is set, the controller logs that it is running without HA support and operates in single-instance mode.

SSH Proxy Configuration

VariableDescriptionDefault
SSH_PROXY_ENABLEDEnable the bastion-style SSH proxy listener (true / 1 to enable)false
SSH_PROXY_ADDRBind address for the SSH proxy listener0.0.0.0:2222

Health Checks & Monitoring

The controller exposes health and status endpoints for load balancer checks, monitoring systems, and the Admin UI Dashboard.

Health Endpoints

EndpointAuthPurpose & response
GET /health
GET /api/health
NoLightweight liveness check. Returns { "status": "ok", "version": "..." }. Ideal for load balancer health checks.
GET /ready
GET /api/ready
NoReadiness probe. Tests the database (and Valkey if configured) and returns status, database, and valkey fields.
GET /api/admin/healthYes (admin)Detailed system health with per-component status and latencies (database, SSH proxy, etc.).
GET /api/admin/ha-statusYes (admin)Cluster mode, instance list, Valkey status, and session distribution. Backs the Dashboard HA Status widget.

Basic Health Response

health-check.shbash
curl -k https://netstacks.example.net/health
health-response.jsonjson
{
  "status": "ok",
  "version": "0.0.14"
}

Readiness Response

ready-response.jsonjson
{
  "status": "ready",
  "database": "connected",
  "valkey": "connected"
}

HA Status Response (admin)

ha-status-response.jsonjson
{
  "mode": "active-active",
  "this_instance": "3f9c2d8a-1b6e-4a77-9c0b-2e5f1d8a4b7c",
  "valkey_status": "connected",
  "instances": [
    {
      "instance_id": "3f9c2d8a-1b6e-4a77-9c0b-2e5f1d8a4b7c",
      "advertise_addr": "https://controller-1:3000",
      "active_sessions": 12,
      "is_self": true
    },
    {
      "instance_id": "a1b2c3d4-e5f6-7890-abcd-ef0123456789",
      "advertise_addr": "https://controller-2:3000",
      "active_sessions": 9,
      "is_self": false
    }
  ],
  "total_sessions": 21
}

mode is "active-active" when more than one instance is registered and "single-instance" otherwise. valkey_status is one of "connected", "disconnected", or "not configured".

Monitoring with Prometheus

The controller does not expose a native Prometheus metrics endpoint. Use a JSON exporter to scrape /api/admin/ha-status (with a bearer token) or the unauthenticated /ready endpoint and convert the fields to metrics.

json-exporter.ymlyaml
# json_exporter config (config.yml) — scrape /api/admin/ha-status
modules:
  netstacks_ha:
    headers:
      Authorization: "Bearer ${NETSTACKS_TOKEN}"
    metrics:
      - name: netstacks_cluster_instances
        path: '{ len(.instances) }'
        help: "Number of registered controller instances"
      - name: netstacks_sessions_total
        path: '{ .total_sessions }'
        help: "Total active sessions across the cluster"
      - name: netstacks_valkey_connected
        path: '{ .valkey_status }'
        help: "Valkey connection status"
        values:
          connected: 1
          disconnected: 0

Alerting Script

readiness-check.shbash
#!/bin/bash
# Poll the unauthenticated readiness endpoint from a monitoring host
RESPONSE=$(curl -sk https://netstacks.example.net/ready)
STATUS=$(echo "$RESPONSE" | jq -r '.status')
DB=$(echo "$RESPONSE" | jq -r '.database')

if [ "$STATUS" != "ready" ]; then
  echo "CRITICAL: controller not ready (db=$DB)"
  exit 2
fi

echo "OK: ready (db=$DB)"
exit 0
Load balancer health checks

Configure your load balancer to poll /health every 10 seconds with a 5-second timeout, and remove an instance from the pool after 3 consecutive failures. The endpoint is lightweight and unauthenticated, making it safe for frequent polling.

Backup & Restore

PostgreSQL is the single source of truth for all persistent data in NetStacks (including the JWT signing secret). Regular database backups are essential. Valkey data is ephemeral (session state, heartbeats) and does not need backing up — it rebuilds automatically when instances restart.

Manual Database Backup

manual-backup.shbash
# Backup the full database (compressed)
pg_dump -h db.example.net -U netstacks -d netstacks \
  --format=custom \
  --compress=9 \
  --file=netstacks-backup-$(date +%Y%m%d-%H%M%S).dump

# Backup only the schema (for documentation or migration reference)
pg_dump -h db.example.net -U netstacks -d netstacks \
  --schema-only \
  --file=netstacks-schema-$(date +%Y%m%d).sql

Automated Backups

Add an automated backup container to your compose stack that runs daily backups with configurable retention.

automated-backup.ymlyaml
  db-backup:
    image: prodrigestivill/postgres-backup-local:16
    restart: unless-stopped
    depends_on:
      postgres:
        condition: service_healthy
    environment:
      POSTGRES_HOST: postgres
      POSTGRES_DB: netstacks
      POSTGRES_USER: netstacks
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_EXTRA_OPTS: "--format=custom --compress=9"
      SCHEDULE: "0 2 * * *"        # daily at 02:00
      BACKUP_KEEP_DAYS: 7
      BACKUP_KEEP_WEEKS: 4
      BACKUP_KEEP_MONTHS: 6
    volumes:
      - db-backups:/backups

volumes:
  db-backups:

Restore from Backup

restore-backup.shbash
# Stop the controller instances first
docker compose -f docker-compose.production.yml stop controller-1 controller-2

# Restore from a custom-format backup
pg_restore -h db.example.net -U netstacks -d netstacks \
  --clean --if-exists --no-owner --no-privileges \
  netstacks-backup-20260410-020000.dump

# Verify the restore
psql -h db.example.net -U netstacks -d netstacks \
  -c "SELECT count(*) FROM users;"

# Restart the controller instances
docker compose -f docker-compose.production.yml start controller-1 controller-2
Vault master key recovery

Database backups contain encrypted credential data. To read those credentials after a restore, you must use the same VAULT_MASTER_KEY that was active when the data was encrypted. A backup restored with a different vault key will start successfully, but all encrypted credentials will be unreadable. Always back up the vault master key separately from the database.

Upgrading

NetStacks Controller supports rolling upgrades in HA mode. The controller applies database migrations on startup, so no manual migration step is needed.

Rolling Upgrade (HA Mode)

Upgrade one controller instance at a time to maintain availability. The load balancer health check removes the restarting instance and adds it back once /health returns 200.

rolling-upgrade.shbash
# Pull (or rebuild) the new controller image first
docker compose -f docker-compose.production.yml pull controller-1 controller-2

# Step 1: upgrade controller-1
docker compose -f docker-compose.production.yml up -d --no-deps controller-1

# Wait for controller-1 to report healthy through the load balancer
until curl -sk https://localhost/health | grep -q '"status":"ok"'; do
  echo "Waiting for controller-1..."
  sleep 5
done
echo "controller-1 is healthy"

# Step 2: upgrade controller-2
docker compose -f docker-compose.production.yml up -d --no-deps controller-2

until curl -sk https://localhost/health | grep -q '"status":"ok"'; do
  echo "Waiting for controller-2..."
  sleep 5
done
echo "Upgrade complete — verify cluster membership:"

# Confirm both instances are registered (admin token required)
curl -sk https://localhost/api/admin/ha-status \
  -H "Authorization: Bearer $TOKEN" | jq '.instances | length'
Migrations run on startup

Database migrations are applied automatically when a controller starts with a new version. Upgrade one instance at a time in a rolling fashion so the cluster always has a healthy member serving traffic.

Rollback

If an upgrade introduces issues, roll back by pinning the controller image to the previous tag and restarting.

rollback.shbash
# Pin the previous version (edit the compose file or use an env override)
#   image: netstacks-controller:0.0.13

docker compose -f docker-compose.production.yml up -d --no-deps controller-1
sleep 10
docker compose -f docker-compose.production.yml up -d --no-deps controller-2

# Verify
curl -sk https://localhost/health | jq .
Rollback and breaking migrations

If the new version applied migrations that alter existing tables (dropping columns, changing types), rolling back to the previous version may cause errors because the old code expects the old schema. Always read the release notes and take a database backup before major upgrades.

Blue-green deployment

For zero-risk upgrades, deploy the new version as a separate set of controller instances pointing at the same database, verify health on the green stack, then switch the load balancer to the green instances. If issues arise, switch back to blue instantly.

Questions & Answers

Does NetStacks Controller run active-active or active-passive?
Active-active. Every controller instance serves all traffic simultaneously; there is no primary controller. The HA status endpoint reports mode: "active-active" once more than one instance is registered in Valkey.
What is the controller health check URL?
GET /health (also GET /api/health), unauthenticated, returning { "status": "ok", "version": "..." }. For a readiness probe that tests the database and Valkey, use /ready. There is no /api/v1/* path.
Do I need Valkey for a single controller?
No. If neither VALKEY_SENTINEL_URLS nor VALKEY_URL is set, the controller runs in single-instance mode. Valkey is required only when running multiple instances that must coordinate sessions.
Which secrets must match across instances?
VAULT_MASTER_KEY must be identical on every instance. DATABASE_URL must point all instances at the same database. The JWT signing secret (auth.jwt_secret) is stored in that shared database, so it is automatically consistent.
How do I set the JWT signing secret?
It is not an environment variable. Set the auth.jwt_secret setting from Admin → Settings. Replace the default placeholder before production.
How is TLS enabled?
The controller always serves HTTPS. If you do not set both TLS_CERT_PATH and TLS_KEY_PATH, it auto-generates a CA and server cert under TLS_DATA_DIR (default /data/tls). Add extra SANs with TLS_SANS. There is no TLS_ENABLED toggle.
What is the correct Sentinel environment variable?
VALKEY_SENTINEL_URLS (plural) — a comma-separated list of full redis://host:port URLs — together with VALKEY_SENTINEL_MASTER (the master name, default mymaster).
Where do I see cluster health in the UI?
The Admin UI Dashboard includes an HA Status widget showing mode, instances, Valkey connectivity, and session distribution, backed by GET /api/admin/ha-status.