NetStacksNetStacks

Execution Monitoring

Enterprise

Track MOP executions live in the MOP workspace and review scheduled/agent task execution history and status through the Controller API.

Overview

NetStacks Enterprise tracks two distinct kinds of automated work, each with its own monitoring surface and its own set of status values:

  • MOP executions — multi-step Method of Procedure runs against target devices. These are monitored live inside the MOP workspace, with per-device and per-step progress.
  • Scheduled & agent tasks — cron-driven backups and AI agent runs. Each run produces a task execution record that you review through the tasks history endpoints and cancel through the tasks API.
No single unified dashboard

There is no combined “Monitoring” page that merges MOP and scheduled-task activity into one view. MOP progress lives in the MOP workspace; scheduled and agent task runs are surfaced through the tasks panel and the Controller API. This page documents both surfaces and the real endpoints behind them.

Enterprise feature

Execution monitoring, MOP executions, and scheduled tasks are Controller (Enterprise) features. The standalone NetStacks terminal does not run a Controller scheduler or MOP execution engine.

Where Monitoring Lives

MOP executions — the MOP workspace

When you run a MOP, the execution opens in the MOP workspace. The live execution view shows the current phase, a per-device status board, and the output of each step as it runs. Lifecycle controls (pause, resume, abort) act on the execution directly from this view.

Scheduled & agent tasks — tasks history

Each time a scheduled task fires (on its cron schedule or via a manual run), the Controller creates a task execution record. Agent task runs are browsable through the agent task history endpoint, and any in-flight execution can be cancelled through the executions API.

History is per task type

Agent task runs are returned by GET /api/admin/agent-tasks/history. Backup runs write a config snapshot you review under device snapshots. There is no endpoint that lists every execution of every type in one call.

Execution Status Models

Scheduled-task executions and MOP executions use different status vocabularies. Do not assume one set applies to the other.

Scheduled & agent task execution status

StatusMeaning
pendingExecution record created, not yet started by the executor
runningThe executor has acquired a slot and is running the task
completedThe task finished and returned a result
failedThe task returned an error
cancelledAn operator cancelled the execution before it finished
timeoutSurfaced by the agent task history view for runs that exceeded their limit

MOP execution status

StatusSet byMeaning
runningstart / resumeThe execution is active (or has been resumed)
pausedpauseThe execution is held, typically at a phase checkpoint
abortedabortThe execution was stopped before completing
completedcompleteThe execution finished all phases

MOP device & step status

Within a MOP execution, each device and each step carries its own status. Devices move through pending, running, skipped, and rolling_back. Steps move through pending, running, passed, failed, approved, and skipped. The device status board in the MOP workspace renders these directly.

Monitoring MOP Executions

A MOP execution is created with an execution strategy and a control mode. The defaults are sequential strategy, manual control mode, and an on_failure behaviour of pause. Phase checkpoints (after pre-checks, after changes, after post-checks) default to pausing so an operator can review before continuing.

Live execution view

The MOP workspace live view surfaces, in real time:

  • Overall execution status (running, paused, aborted, completed)
  • The current phase (pre-check, change, post-check)
  • A per-device status board with each device's current state
  • Per-step status, captured output, and duration in milliseconds

Acting on a running execution

From the live view you can:

  • Pause — hold the execution (status becomes paused)
  • Resume — continue a paused execution (status returns to running)
  • Abort — stop the execution (status becomes aborted)
  • Skip or retry a single device, or roll back a device
  • Approve, skip, or re-run an individual step
Abort does not auto-roll-back

Aborting a MOP execution stops further steps, but commands already sent to devices are not reversed automatically. If the MOP defines rollback steps, roll back affected devices explicitly, or run the rollback phase, to return devices to a known-good state.

Monitoring Scheduled & Agent Tasks

The Controller's task executor runs each task in a background slot with a concurrency limit. Only two task types currently execute: backup and agent_task. Other task types (deployment, health_check, custom_script) are defined in the schema but return a “not yet implemented” error if scheduled.

Backup tasks

A backup task collects device configs and stores them as a config snapshot. The execution result is a small JSON summary — not a per-device log transcript — with the snapshot id and success/failure counts. Review the collected configs and the snapshot under device snapshots.

Agent tasks

An agent task runs an AI prompt on its schedule. Each run is recorded as a task execution with the prompt attached, and you browse past runs through the agent task history endpoint. A run can be triggered manually with the schedule's /run endpoint, which returns the new execution record immediately and runs it in the background.

Cancelling a task execution

Any in-flight task execution is cancelled with POST /api/admin/tasks/executions/{id}/cancel. This aborts the run if it is executing on the current node and marks the execution record cancelled. The underlying scheduled task stays enabled and will fire again at its next cron time.

Retries & Timeouts

Scheduled tasks store three tuning fields. Their schema defaults are max_retries = 3, retry_delay_seconds = 60, and timeout_seconds = 3600 for agent tasks (the general create path also defaults the timeout to 3600). These values are persisted with the task and returned by the get/create/update endpoints.

Retry config is stored, not yet a retry loop

The current executor runs each task once per trigger; it does not wrap the run in an automatic max_retries loop. If a run fails, the execution is marked failed and the task simply runs again at its next scheduled time. Treat max_retries and retry_delay_seconds as forward-looking configuration rather than behaviour you can rely on today.

SSH command timeout

Inside a MOP execution, each step runs a single SSH command with a fixed per-command timeout (60 seconds). The backup task uses the platform's config-collection settings (including its own concurrency and retry handling for config collection), independent of a task's retry_delay_seconds field.

API Examples

List agent task execution history

Browse past agent task runs (admin-gated):

list-agent-history.httphttp
GET /api/admin/agent-tasks/history?limit=20&offset=0
Authorization: Bearer <token>

Cancel a running task execution

Cancel an in-flight scheduled or agent task execution by its execution id. A successful cancel returns 204 No Content:

cancel-task-execution.httphttp
POST /api/admin/tasks/executions/{execution_id}/cancel
Authorization: Bearer <token>

# Response: 204 No Content

Run an agent task now

Trigger a scheduled agent task immediately. The Controller returns the new execution record (HTTP 202) and runs it in the background:

run-agent-task.httphttp
POST /api/admin/tasks/agent-schedules/{task_id}/run
Authorization: Bearer <token>

# Response: 202 Accepted  (body: the created TaskExecution)

Backup task execution result

A completed backup task returns a JSON summary referencing the snapshot it wrote — not a device-by-device console log:

backup-result.jsonjson
{
  "status": "complete",
  "snapshot_id": "a1f3c8e2-7b9d-4e21-9c6a-0d2f4b8e1a55",
  "total_devices": 4,
  "successful": 4,
  "failed": 0
}

// If some devices fail, status is "partial"; if all fail, "failed".
// Per-device collection errors are recorded on the snapshot,
// not returned inline in this summary.

Abort a MOP execution

Stop a running MOP execution. The execution status transitions to aborted and a completion timestamp is set:

abort-mop-execution.httphttp
POST /api/mop-executions/{execution_id}/abort
Authorization: Bearer <token>

# Related MOP lifecycle endpoints:
#   POST /api/mop-executions/{id}/start     -> running
#   POST /api/mop-executions/{id}/pause     -> paused
#   POST /api/mop-executions/{id}/resume    -> running
#   POST /api/mop-executions/{id}/complete  -> completed

Run a phase on a device (result shape)

Executing a phase against a device returns the steps that ran with pass/fail counts:

execute-phase-result.jsonjson
POST /api/mop-executions/{exec_id}/devices/{device_id}/execute-phase
Content-Type: application/json
Authorization: Bearer <token>

{ "step_type": "pre_check" }

// Response (PhaseExecutionResult):
{
  "device_id": "9c6a0d2f-...",
  "step_type": "pre_check",
  "steps_executed": 2,
  "steps_passed": 2,
  "steps_failed": 0,
  "snapshot_id": null,
  "combined_output": "...",
  "results": [
    { "step_id": "...", "status": "passed", "output": "...", "duration_ms": 1180 },
    { "step_id": "...", "status": "passed", "output": "...", "duration_ms": 1840 }
  ]
}

Health check output (illustrative / planned)

Health-check task execution is not yet implemented — scheduling a health_check task returns a “not yet implemented” error. The block below is a sketch of what a per-device health summary could look like once available; it does not reflect current output:

health-check-illustrative.txttext
# ILLUSTRATIVE ONLY — health_check does not execute today
Task: Edge Switch Health Monitor
Type: health_check
Status: failed

Failing Device:
  edge-sw-03.branch2 (10.5.3.1)
    Check: ssh
    Error: "Connection timed out - no route to host"

Passing Devices:
  edge-sw-01.branch1 (10.5.1.1) - all checks passed
  edge-sw-02.branch1 (10.5.1.2) - all checks passed

Questions & Answers

Q: Is there a single Monitoring dashboard for all automation?
A: No. NetStacks does not ship a unified monitoring page. MOP executions are monitored live in the MOP workspace; scheduled and agent task runs are reviewed through the tasks history endpoints (for example GET /api/admin/agent-tasks/history) and the executions API.
Q: What statuses can a scheduled or agent task execution have?
A: pending, running, completed, failed, and cancelled. The agent task history view also surfaces timeout for runs that exceeded their limit.
Q: What statuses can a MOP execution have?
A: running, paused, aborted, and completed. Start and resume set running; pause sets paused; abort sets aborted; complete sets completed. These are distinct from the scheduled-task statuses.
Q: How do I cancel a running task?
A: Call POST /api/admin/tasks/executions/{id}/cancel. It returns 204 No Content, aborts the run if it is executing on the current node, and marks the execution cancelled. The scheduled task itself stays enabled for its next cron run.
Q: How do I stop a MOP execution?
A: Use POST /api/mop-executions/{id}/abort (status becomes aborted), or /pause to hold it and /resume to continue. Abort does not roll back changes already applied — roll back affected devices explicitly if needed.
Q: Does NetStacks automatically retry a failed task?
A: Not today. The executor runs a task once per trigger. The max_retries and retry_delay_seconds fields are stored on the task (defaults 3 and 60s) but are not yet enforced as an automatic retry loop. A failed run simply re-runs at its next scheduled time.
Q: Which scheduled task types actually run?
A: backup and agent_task. The deployment, health_check, and custom_script types exist in the schema but currently return a “not yet implemented” error when executed.
Q: Where do backup results show up?
A: A backup task writes a config snapshot and returns a JSON summary with snapshot_id, total_devices, successful, and failed counts. Review the snapshot and the collected configs under device snapshots.

Troubleshooting

A task execution is stuck in “running”

If the Controller restarted mid-run, an execution record can remain in running because the in-memory cancellation token for that run no longer exists. Cancelling it via POST /api/admin/tasks/executions/{id}/cancel still marks the database record cancelled even when the run is no longer active on this node.

A scheduled task always fails immediately

Confirm the task type is one that executes. Scheduling a deployment, health_check, or custom_script task fails with a “not yet implemented” error by design. Use backup or agent_task instead.

Backup task reports failures for some devices

The backup result reports successful and failed counts; per-device errors (no credential configured, credential decrypt error, collection failed, device not found) are recorded on the snapshot rather than in the summary. Open the snapshot to see which devices failed and why, then fix the device's default credential or connectivity.

A MOP step shows no output

MOP steps run a single SSH command with a 60-second timeout. If a step failed with an SSH error, the device may be unreachable or the credential may be wrong. Verify the device's credential and connectivity, then retry the device or re-run the step.

Use mock steps to dry-run

MOP steps support a mock mode: when a step has mock output enabled, executing it returns the mock text and marks the step passed without touching the device. This is useful for validating a procedure's flow before running it for real.

Monitoring works alongside these NetStacks features: