Execution Monitoring
EnterpriseTrack MOP executions live in the MOP workspace and review scheduled/agent task execution history and status through the Controller API.
Overview
NetStacks Enterprise tracks two distinct kinds of automated work, each with its own monitoring surface and its own set of status values:
- MOP executions — multi-step Method of Procedure runs against target devices. These are monitored live inside the MOP workspace, with per-device and per-step progress.
- Scheduled & agent tasks — cron-driven backups and AI agent runs. Each run produces a task execution record that you review through the tasks history endpoints and cancel through the tasks API.
There is no combined “Monitoring” page that merges MOP and scheduled-task activity into one view. MOP progress lives in the MOP workspace; scheduled and agent task runs are surfaced through the tasks panel and the Controller API. This page documents both surfaces and the real endpoints behind them.
Execution monitoring, MOP executions, and scheduled tasks are Controller (Enterprise) features. The standalone NetStacks terminal does not run a Controller scheduler or MOP execution engine.
Where Monitoring Lives
MOP executions — the MOP workspace
When you run a MOP, the execution opens in the MOP workspace. The live execution view shows the current phase, a per-device status board, and the output of each step as it runs. Lifecycle controls (pause, resume, abort) act on the execution directly from this view.
Scheduled & agent tasks — tasks history
Each time a scheduled task fires (on its cron schedule or via a manual run), the Controller creates a task execution record. Agent task runs are browsable through the agent task history endpoint, and any in-flight execution can be cancelled through the executions API.
Agent task runs are returned by GET /api/admin/agent-tasks/history. Backup runs write a config snapshot you review under device snapshots. There is no endpoint that lists every execution of every type in one call.
Execution Status Models
Scheduled-task executions and MOP executions use different status vocabularies. Do not assume one set applies to the other.
Scheduled & agent task execution status
| Status | Meaning |
|---|---|
| pending | Execution record created, not yet started by the executor |
| running | The executor has acquired a slot and is running the task |
| completed | The task finished and returned a result |
| failed | The task returned an error |
| cancelled | An operator cancelled the execution before it finished |
| timeout | Surfaced by the agent task history view for runs that exceeded their limit |
MOP execution status
| Status | Set by | Meaning |
|---|---|---|
| running | start / resume | The execution is active (or has been resumed) |
| paused | pause | The execution is held, typically at a phase checkpoint |
| aborted | abort | The execution was stopped before completing |
| completed | complete | The execution finished all phases |
MOP device & step status
Within a MOP execution, each device and each step carries its own status. Devices move through pending, running, skipped, and rolling_back. Steps move through pending, running, passed, failed, approved, and skipped. The device status board in the MOP workspace renders these directly.
Monitoring MOP Executions
A MOP execution is created with an execution strategy and a control mode. The defaults are sequential strategy, manual control mode, and an on_failure behaviour of pause. Phase checkpoints (after pre-checks, after changes, after post-checks) default to pausing so an operator can review before continuing.
Live execution view
The MOP workspace live view surfaces, in real time:
- Overall execution status (running, paused, aborted, completed)
- The current phase (pre-check, change, post-check)
- A per-device status board with each device's current state
- Per-step status, captured output, and duration in milliseconds
Acting on a running execution
From the live view you can:
- Pause — hold the execution (status becomes paused)
- Resume — continue a paused execution (status returns to running)
- Abort — stop the execution (status becomes aborted)
- Skip or retry a single device, or roll back a device
- Approve, skip, or re-run an individual step
Aborting a MOP execution stops further steps, but commands already sent to devices are not reversed automatically. If the MOP defines rollback steps, roll back affected devices explicitly, or run the rollback phase, to return devices to a known-good state.
Monitoring Scheduled & Agent Tasks
The Controller's task executor runs each task in a background slot with a concurrency limit. Only two task types currently execute: backup and agent_task. Other task types (deployment, health_check, custom_script) are defined in the schema but return a “not yet implemented” error if scheduled.
Backup tasks
A backup task collects device configs and stores them as a config snapshot. The execution result is a small JSON summary — not a per-device log transcript — with the snapshot id and success/failure counts. Review the collected configs and the snapshot under device snapshots.
Agent tasks
An agent task runs an AI prompt on its schedule. Each run is recorded as a task execution with the prompt attached, and you browse past runs through the agent task history endpoint. A run can be triggered manually with the schedule's /run endpoint, which returns the new execution record immediately and runs it in the background.
Any in-flight task execution is cancelled with POST /api/admin/tasks/executions/{id}/cancel. This aborts the run if it is executing on the current node and marks the execution record cancelled. The underlying scheduled task stays enabled and will fire again at its next cron time.
Retries & Timeouts
Scheduled tasks store three tuning fields. Their schema defaults are max_retries = 3, retry_delay_seconds = 60, and timeout_seconds = 3600 for agent tasks (the general create path also defaults the timeout to 3600). These values are persisted with the task and returned by the get/create/update endpoints.
The current executor runs each task once per trigger; it does not wrap the run in an automatic max_retries loop. If a run fails, the execution is marked failed and the task simply runs again at its next scheduled time. Treat max_retries and retry_delay_seconds as forward-looking configuration rather than behaviour you can rely on today.
SSH command timeout
Inside a MOP execution, each step runs a single SSH command with a fixed per-command timeout (60 seconds). The backup task uses the platform's config-collection settings (including its own concurrency and retry handling for config collection), independent of a task's retry_delay_seconds field.
API Examples
List agent task execution history
Browse past agent task runs (admin-gated):
GET /api/admin/agent-tasks/history?limit=20&offset=0
Authorization: Bearer <token>Cancel a running task execution
Cancel an in-flight scheduled or agent task execution by its execution id. A successful cancel returns 204 No Content:
POST /api/admin/tasks/executions/{execution_id}/cancel
Authorization: Bearer <token>
# Response: 204 No ContentRun an agent task now
Trigger a scheduled agent task immediately. The Controller returns the new execution record (HTTP 202) and runs it in the background:
POST /api/admin/tasks/agent-schedules/{task_id}/run
Authorization: Bearer <token>
# Response: 202 Accepted (body: the created TaskExecution)Backup task execution result
A completed backup task returns a JSON summary referencing the snapshot it wrote — not a device-by-device console log:
{
"status": "complete",
"snapshot_id": "a1f3c8e2-7b9d-4e21-9c6a-0d2f4b8e1a55",
"total_devices": 4,
"successful": 4,
"failed": 0
}
// If some devices fail, status is "partial"; if all fail, "failed".
// Per-device collection errors are recorded on the snapshot,
// not returned inline in this summary.Abort a MOP execution
Stop a running MOP execution. The execution status transitions to aborted and a completion timestamp is set:
POST /api/mop-executions/{execution_id}/abort
Authorization: Bearer <token>
# Related MOP lifecycle endpoints:
# POST /api/mop-executions/{id}/start -> running
# POST /api/mop-executions/{id}/pause -> paused
# POST /api/mop-executions/{id}/resume -> running
# POST /api/mop-executions/{id}/complete -> completedRun a phase on a device (result shape)
Executing a phase against a device returns the steps that ran with pass/fail counts:
POST /api/mop-executions/{exec_id}/devices/{device_id}/execute-phase
Content-Type: application/json
Authorization: Bearer <token>
{ "step_type": "pre_check" }
// Response (PhaseExecutionResult):
{
"device_id": "9c6a0d2f-...",
"step_type": "pre_check",
"steps_executed": 2,
"steps_passed": 2,
"steps_failed": 0,
"snapshot_id": null,
"combined_output": "...",
"results": [
{ "step_id": "...", "status": "passed", "output": "...", "duration_ms": 1180 },
{ "step_id": "...", "status": "passed", "output": "...", "duration_ms": 1840 }
]
}Health check output (illustrative / planned)
Health-check task execution is not yet implemented — scheduling a health_check task returns a “not yet implemented” error. The block below is a sketch of what a per-device health summary could look like once available; it does not reflect current output:
# ILLUSTRATIVE ONLY — health_check does not execute today
Task: Edge Switch Health Monitor
Type: health_check
Status: failed
Failing Device:
edge-sw-03.branch2 (10.5.3.1)
Check: ssh
Error: "Connection timed out - no route to host"
Passing Devices:
edge-sw-01.branch1 (10.5.1.1) - all checks passed
edge-sw-02.branch1 (10.5.1.2) - all checks passedQuestions & Answers
- Q: Is there a single Monitoring dashboard for all automation?
- A: No. NetStacks does not ship a unified monitoring page. MOP executions are monitored live in the MOP workspace; scheduled and agent task runs are reviewed through the tasks history endpoints (for example
GET /api/admin/agent-tasks/history) and the executions API. - Q: What statuses can a scheduled or agent task execution have?
- A:
pending,running,completed,failed, andcancelled. The agent task history view also surfacestimeoutfor runs that exceeded their limit. - Q: What statuses can a MOP execution have?
- A:
running,paused,aborted, andcompleted. Start and resume setrunning; pause setspaused; abort setsaborted; complete setscompleted. These are distinct from the scheduled-task statuses. - Q: How do I cancel a running task?
- A: Call
POST /api/admin/tasks/executions/{id}/cancel. It returns204 No Content, aborts the run if it is executing on the current node, and marks the execution cancelled. The scheduled task itself stays enabled for its next cron run. - Q: How do I stop a MOP execution?
- A: Use
POST /api/mop-executions/{id}/abort(status becomesaborted), or/pauseto hold it and/resumeto continue. Abort does not roll back changes already applied — roll back affected devices explicitly if needed. - Q: Does NetStacks automatically retry a failed task?
- A: Not today. The executor runs a task once per trigger. The
max_retriesandretry_delay_secondsfields are stored on the task (defaults 3 and 60s) but are not yet enforced as an automatic retry loop. A failed run simply re-runs at its next scheduled time. - Q: Which scheduled task types actually run?
- A:
backupandagent_task. Thedeployment,health_check, andcustom_scripttypes exist in the schema but currently return a “not yet implemented” error when executed. - Q: Where do backup results show up?
- A: A backup task writes a config snapshot and returns a JSON summary with
snapshot_id,total_devices,successful, andfailedcounts. Review the snapshot and the collected configs under device snapshots.
Troubleshooting
A task execution is stuck in “running”
If the Controller restarted mid-run, an execution record can remain in running because the in-memory cancellation token for that run no longer exists. Cancelling it via POST /api/admin/tasks/executions/{id}/cancel still marks the database record cancelled even when the run is no longer active on this node.
A scheduled task always fails immediately
Confirm the task type is one that executes. Scheduling a deployment, health_check, or custom_script task fails with a “not yet implemented” error by design. Use backup or agent_task instead.
Backup task reports failures for some devices
The backup result reports successful and failed counts; per-device errors (no credential configured, credential decrypt error, collection failed, device not found) are recorded on the snapshot rather than in the summary. Open the snapshot to see which devices failed and why, then fix the device's default credential or connectivity.
A MOP step shows no output
MOP steps run a single SSH command with a 60-second timeout. If a step failed with an SSH error, the device may be unreachable or the credential may be wrong. Verify the device's credential and connectivity, then retry the device or re-run the step.
MOP steps support a mock mode: when a step has mock output enabled, executing it returns the mock text and marks the step passed without touching the device. This is useful for validating a procedure's flow before running it for real.
Related Features
Monitoring works alongside these NetStacks features:
- Method of Procedures (MOPs) — build the multi-step procedures whose executions you monitor live
- Scheduled Tasks — create backup and agent tasks that produce execution records
- Cron Expressions — scheduling syntax that drives scheduled task runs
- Approvals — gate MOP changes behind reviewer sign-off before execution
- Config Snapshots — where backup task results are stored and reviewed
- Activity Monitor — live operational activity across the Controller
- Audit Logs — durable record of operations, including MOP step dispatch
- NOC Agents — the AI agents that back agent task executions