NetStacksNetStacks

Stack Instances

Enterprise

Deploy stack instances from the NetStacks Controller, track deployment status through the real state machine, and roll back via gNMI/NETCONF.

Overview

Controller-only feature

Stacks, stack instances, and deployments are part of the NetStacks Controller. The endpoints below live under the Controller config API at /api/config and require an operator session.

A stack instance is a saved, undeployed binding. It connects a stack to a set of target devices with concrete variable values and per-device overrides, and it carries a state field so you can stage a configuration before pushing it. Instances are managed with standard CRUD on /api/config/instances:

  • GET /api/config/instances — list instances (optionally filtered by ?stack_id=)
  • POST /api/config/instances — create an instance
  • GET /api/config/instances/{id} — fetch one instance
  • PUT /api/config/instances/{id} — update name, target devices, variable values, overrides, or state
  • DELETE /api/config/instances/{id} — delete the instance

When you are ready to push, you trigger a deployment from the instance with POST /api/config/instances/{id}/deploy. That creates a ConfigDeployment record and runs the deploy pipeline in the background. The deployment — not the instance — is what tracks live progress and per-device results.

Deployment State Machine

A deployment moves through a fixed set of statuses. There are six deployment statuses — there is no separate "validating" status; validation is an internal stage that runs while the deployment is in_progress.

StatusDescription
pendingThe deployment record was created but the background pipeline has not started yet. This is the status returned immediately by the deploy call.
in_progressThe pipeline is running: rendering templates, taking pre-flight backups, pushing configs, validating, and confirming.
completedAll target devices received their configuration successfully and the pipeline finished.
failedOne or more devices failed (and the stack is not atomic, or a pre-flight stage failed). Per-device error_message values explain why.
rolling_backA rollback is in progress — either an automatic atomic rollback after a partial failure, or an explicit rollback you triggered.
rolled_backRollback finished. Affected devices have been restored from their pre-flight backup configuration.
Typical transitions

Success path: pending → in_progress → completed. Atomic stack with a partial failure: in_progress → rolling_back → rolled_back. Non-atomic failure: in_progress → failed (succeeded devices keep their new config). You can then call the explicit rollback endpoint on a completed or failed deployment.

The Deploy Pipeline

Deploying an instance runs a six-stage pipeline against all target devices. These stages are internal phases of the in_progress status, not separate statuses. NetStacks reaches devices over their configured structured transport (gNMI or NETCONF), not raw CLI over SSH.

Stage 1 — Render

The stack is resolved into per-device render jobs. Each service template is rendered with the merged variable values (shared values plus per-device overrides). A ConfigDeviceDeployment record is created for each job, capturing the rendered_config, target_path, and operation.

Stage 2 — Pre-flight backup

Before any change, NetStacks pulls the current config from each device and stores it as a pre_deploy_backup version. This backup is always taken — there is no flag to skip it. The backup version id is linked to each device deployment as backup_config_id and is what rollback restores. A secondary CLI backup over SSH is also attempted on a best-effort, non-fatal basis. If the pre-flight backup pull fails for a device, the whole deployment is marked failed before any push happens.

Stage 3 — Push

Rendered configs are applied. Devices are pushed concurrently; services within a single device are applied in order. Each device deployment moves to in_progress, then completed or failed. For transports that support it, changes are applied as a confirmed commit so they can auto-revert if not confirmed.

Stage 4 — Validate

After the push, NetStacks pulls each target path back and compares hashes (for structured configs) or checks rendered lines in the running config (for CLI templates). A mismatch is logged as drift. Validation is best-effort: a failed read-back does not by itself fail the deployment, but detected drift on an atomic stack causes the confirm to be skipped so the change auto-reverts.

Stage 5 — Confirm

Confirmed commits are finalized for devices that passed validation. On atomic stacks, a device with detected drift is deliberately not confirmed, letting the device's confirmed-commit timer roll the change back.

Stage 6 — Sync

NetStacks pulls each device's full config one more time and stores it as a post_deploy_sync version so the config history reflects the deployed state.

Rollback

Rollback restores the pre-flight backup. It happens two ways:

  • Automatic (atomic stacks) — if any device fails during push on a stack marked atomic, the deployment enters rolling_back and the already-succeeded devices are reverted: gNMI devices have their pre-flight backup re-pushed, NETCONF devices auto-revert by skipping the confirm. The deployment ends in rolled_back.
  • Explicit — call POST /api/config/deployments/{id}/rollback on a completed or failed deployment. NetStacks re-pushes each device's backup_config_id as a full replace. Rollback fails if no backup configs are recorded.

Step-by-Step Guide

Step 1: Preview the render (recommended)

Before deploying, render the stack to see exactly what would be pushed without touching any device. Post the same per-service variable structure your instance uses to POST /api/config/stacks/{stack_id}/render. The response lists one render job per device/service with the full rendered_config, target_path, and operation.

Preview before production pushes

The render endpoint is the dry-run path: it connects to nothing and changes nothing. Review the rendered output for every device before triggering a live deploy.

Step 2: Trigger the deployment

Call POST /api/config/instances/{id}/deploy with an empty body. There are no deploy-time flags — backups are always taken, and rollback behavior is determined by whether the underlying stack is atomic. The call returns a new deployment in pending status; the pipeline then runs in the background.

Step 3: Monitor progress

Poll GET /api/config/deployments/{id} for the deployment plus its per-device records (total_devices, succeeded_count, failed_count, and each device's status). Stream stage-by-stage progress with GET /api/config/deployments/{id}/logs, which you can filter by ?device_id=.

Step 4: Inspect per-device results

Each ConfigDeviceDeployment exposes:

  • rendered_config — the config that was pushed
  • target_path and operation — where and how it was applied
  • backup_config_id — reference to the pre-flight backup version (fetch its text via the device config history)
  • error_message — the failure reason, if any

Step 5: Handle a failure or re-deploy

If a device failed, read its error_message and the deployment logs. There is no separate "redeploy" or "modify" action: edit the instance with PUT /api/config/instances/{id} if values changed, then call POST /api/config/instances/{id}/deploy again to create a fresh deployment.

Step 6: Roll back if needed

To revert a completed or failed deployment, call POST /api/config/deployments/{id}/rollback. NetStacks restores each device from its recorded pre-flight backup.

API Examples

Deploy a stack instance

Deploy takes no request body. The instance already holds the target devices, variable values, and overrides.

deploy-instance.shbash
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  "https://controller.example.net/api/config/instances/$INSTANCE_ID/deploy"

Deployment response

The deploy call returns the newly created deployment in pending status.total_devices is already set to the instance's target-device count, but succeeded_count and failed_count stay zero until the background pipeline runs.

deployment-response.jsonjson
{
  "id": "f8b3c1a2-1c2d-4e5f-8a9b-0c1d2e3f4a5b",
  "org_id": "0a1b2c3d-4e5f-6789-abcd-ef0123456789",
  "stack_id": "a7c4e2d1-9f8e-4321-bcde-1234567890ab",
  "name": "DC1-Core-Baseline",
  "status": "pending",
  "total_devices": 2,
  "succeeded_count": 0,
  "failed_count": 0,
  "created_by": "11112222-3333-4444-5555-666677778888",
  "started_at": null,
  "completed_at": null,
  "created_at": "2026-01-15T10:05:00Z",
  "updated_at": "2026-01-15T10:05:00Z"
}

Poll deployment detail with per-device results

get-deployment.shbash
curl -s \
  -H "Authorization: Bearer $TOKEN" \
  "https://controller.example.net/api/config/deployments/$DEPLOYMENT_ID"

The response flattens the deployment fields and adds a devices array of ConfigDeviceDeployment records:

deployment-detail.jsonjson
{
  "id": "f8b3c1a2-1c2d-4e5f-8a9b-0c1d2e3f4a5b",
  "name": "DC1-Core-Baseline",
  "status": "completed",
  "total_devices": 2,
  "succeeded_count": 2,
  "failed_count": 0,
  "started_at": "2026-01-15T10:05:01Z",
  "completed_at": "2026-01-15T10:05:18Z",
  "devices": [
    {
      "id": "dd000001-0000-0000-0000-000000000001",
      "deployment_id": "f8b3c1a2-1c2d-4e5f-8a9b-0c1d2e3f4a5b",
      "device_id": "de000001-0000-0000-0000-000000000001",
      "service_name": "NTP Configuration",
      "status": "completed",
      "rendered_config": "{\"openconfig-system:ntp\": {\"config\": {\"enabled\": true}}}",
      "target_path": "/system/ntp",
      "operation": "merge",
      "backup_config_id": "bc000001-0000-0000-0000-000000000001",
      "error_message": null,
      "started_at": "2026-01-15T10:05:02Z",
      "completed_at": "2026-01-15T10:05:08Z"
    }
  ]
}

Stream deployment logs

deployment-logs.shbash
curl -s \
  -H "Authorization: Bearer $TOKEN" \
  "https://controller.example.net/api/config/deployments/$DEPLOYMENT_ID/logs?limit=100"
logs-response.jsonjson
[
  { "level": "info", "device_id": null, "message": "Stage 1/6: Rendering stack templates" },
  { "level": "info", "device_id": null, "message": "Stage 2/6: Pre-flight backup" },
  { "level": "info", "device_id": "de000001-0000-0000-0000-000000000001", "message": "Structured backup saved (v7)" },
  { "level": "info", "device_id": null, "message": "Stage 3/6: Pushing configurations" },
  { "level": "info", "device_id": "de000001-0000-0000-0000-000000000001", "message": "Validation passed for path '/system/ntp'" }
]

Render preview (dry run)

Preview the rendered configuration for a stack without deploying. The body uses the per-service shape { "0": { "devices": [...], "shared_vars": {...}, "device_vars": {...} } }.

render-preview.shbash
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  "https://controller.example.net/api/config/stacks/$STACK_ID/render" \
  -d '{
    "variable_values": {
      "0": {
        "devices": ["de000001-0000-0000-0000-000000000001"],
        "shared_vars": { "ntp_server": "10.0.0.10" },
        "device_vars": {}
      }
    }
  }'
render-response.jsonjson
{
  "jobs": [
    {
      "device_id": "de000001-0000-0000-0000-000000000001",
      "device_name": "dc1-core-01",
      "template_id": "11110000-0000-0000-0000-000000000001",
      "template_name": "NTP Configuration",
      "service_order": 0,
      "rendered_config": "{\"openconfig-system:ntp\": {\"config\": {\"enabled\": true}}}",
      "target_path": "/system/ntp",
      "operation": "merge",
      "config_format": "json"
    }
  ]
}

Roll back a deployment

rollback.shbash
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  "https://controller.example.net/api/config/deployments/$DEPLOYMENT_ID/rollback"

This is only valid for deployments in completed or failed status. It re-pushes each device's pre-flight backup and returns the deployment with status rolled_back.

Questions & Answers

Q: What is the difference between an instance and a deployment?
A: An instance is a saved, undeployed binding of a stack to target devices, variable values, and overrides — managed with CRUD on /api/config/instances. A deployment is the live execution created when you call POST /api/config/instances/{id}/deploy; it tracks status, counts, per-device results, and logs.
Q: What are the deployment statuses?
A: Six: pending (created, pipeline not started), in_progress (pipeline running), completed (all devices succeeded), failed (one or more failed), rolling_back (reverting), and rolled_back (revert complete). There is no "validating" status — validation is an internal stage of in_progress.
Q: How do I deploy an instance — what goes in the request body?
A: Send an empty body to POST /api/config/instances/{id}/deploy. The instance already stores its devices and variables, so there are no dry_run, rollback_enabled, or validation_enabled flags.
Q: Are backups always taken?
A: Yes. Stage 2 of the pipeline unconditionally pulls and stores a pre_deploy_backup for every device before any change, plus a best-effort CLI backup over SSH. The backup version is referenced by each device deployment's backup_config_id.
Q: How does rollback work?
A: Two ways. On an atomic stack, a partial push failure automatically reverts the succeeded devices (gNMI re-push of the backup, NETCONF auto-revert by skipping confirm). For any completed or failed deployment, call POST /api/config/deployments/{id}/rollback to re-push each device's pre-flight backup. Rollback fails if no backups were recorded.
Q: Is there a dry-run mode?
A: Yes — POST /api/config/stacks/{stack_id}/render. It returns the rendered config, target path, and operation per device/service without connecting to or changing any device.
Q: How do I re-deploy or modify a deployment?
A: There is no in-place modify or redeploy action. Update the instance with PUT /api/config/instances/{id} if needed, then call the deploy endpoint again to create a new deployment.
Q: How does NetStacks reach the devices?
A: Over each device's configured structured transport — gNMI or NETCONF. Configuration is pushed as structured data at a target path, not as raw CLI over SSH. (REST transport is not yet implemented; SSH is used only for the secondary best-effort CLI backup.)

Troubleshooting

Device unreachable during deployment

Deployments connect over the device's structured transport (gNMI or NETCONF), not SSH port 22. Verify the device's transport type and port, that the gNMI/NETCONF service is listening, and that the Controller can reach it. Confirm a plain read first with POST /api/config/devices/{device_id}/pull, which uses the same transport path.

Authentication failure

Transport authentication errors usually mean the stored credential is wrong or expired. Verify the credential in the Credential Vault, update it, and re-run the deployment.

Configuration push rejected by device

The device may reject a set at the target path (schema mismatch, invalid value, conflicting config). Check the error_message on the failed device deployment and the deployment logs for the exact transport error, fix the Jinja2 template, and deploy again.

Rollback says no backups found

Explicit rollback requires recorded backup_config_id values. If the pre-flight backup stage never completed (for example the deployment failed during backup), there is nothing to restore. Re-run the deployment so a fresh pre-flight backup is captured.

Deployment stuck in in_progress

A deployment stays in_progress while the background pipeline runs. If it does not advance, one or more devices are likely slow to respond on the structured transport. Pull GET /api/config/deployments/{id}/logs to see which stage and device it is on; transport read/write timeouts surface there.

Drift detected but the change reverted

On an atomic stack, if Stage 4 validation detects drift between the rendered and read-back config, Stage 5 deliberately skips the confirm so the device's confirmed-commit timer auto-reverts. Inspect the logged hash mismatch, correct the template or target path, and redeploy.

Learn more about related NetStacks features: