NetStacksNetStacks

Alert Pipeline

Enterprise

Ingest alerts via webhook and Kafka, normalize and deduplicate them, then route through priority rules to AI triage agents, fast-path actions, or suppression.

Overview

The Alerts plugin (display name Alert Ingestion & Pipeline) ingests alerts from external monitoring systems, normalizes their severity into a common scale, deduplicates repeat occurrences into a single alert, and then runs each new alert through a priority-ordered set of routing rules. A rule can hand the alert to an AI triage agent, apply a fast-path action without AI, or suppress it.

Enterprise plugin

Alert Pipeline is a NetStacks Enterprise plugin. It runs as a container (netstacks-plugin-alerts) on internal port 8080 and is reached through the controller proxy at /api/plugins/alerts/*. There is no separate :8443 listener.

What it does

  • Multi-source ingestion — HTTP webhooks and an optional Kafka consumer feed alerts into the pipeline.
  • Severity normalization — source-specific severity strings are mapped to critical, warning, or info.
  • Deduplication — repeat alerts that share a fingerprint within a time window increment an occurrence count instead of creating new alerts.
  • Routing & AI triage — priority rules dispatch alerts to AI agents, fast-path actions, or suppression.
  • Notifications & escalation — email, webhook, Slack, and PagerDuty channels plus time-based escalation policies.
Warning

All plugin endpoints are reached through the authenticated controller proxy and require an org_id query parameter. Do not expose the plugin container port (8080) directly; route through /api/plugins/alerts.

How It Works

Alert processing pipeline

  1. Ingestion — an alert arrives at /ingest/webhook or is consumed from a Kafka topic.
  2. Severity mapping — the source-specific severity is looked up in the org's severity mappings and normalized to critical, warning, or info. With no explicit mapping, common keywords are inferred; if no severity is supplied it defaults to info.
  3. Device resolution — if a device_hostname is provided without a device_id, the plugin resolves it against the core inventory by hostname or IP.
  4. Fingerprint & deduplication — a SHA-256 fingerprint is computed (default: org + source + summary). If an open or acknowledged alert with the same fingerprint exists inside the dedup window, its occurrence count and last-occurrence timestamp are updated and routing is skipped.
  5. Routing — for a genuinely new alert, the routing engine evaluates rules by ascending priority. The first rule whose conditions all match wins; if none match, the alert goes to the default triage agent.

Alert states and triage states

Every alert carries two independent state fields:

state
One of open, acknowledged, resolved, or suppressed. This is the human/operational lifecycle, updated by operators or by fast-path and suppress rules.
triage_state
The AI triage lifecycle: triaging (an agent is running), skipped (fast-path or suppress, no AI), or failed (routing or agent dispatch errored). Alerts also expose triage outputs such as root_cause, impact_summary, resolution, and resolved_by_agent.
Triage events

Each step the engine takes is recorded as a triage event (ingested, routing_matched, routing_skipped, agent_started, agent_failed). Retrieve them per alert at /admin/triage-events/{alert_id}.

Ingestion Sources

Webhook ingestion

Post alerts to /api/plugins/alerts/ingest/webhook?org_id=<org>. The payload model accepts source (required), severity (optional, source-specific — it will be mapped), summary (required), details (a free-form object), and either device_id or device_hostname to associate the alert with a device.

ingest-webhook.shbash
# Ingest an alert via the controller proxy
curl -X POST \
  "https://netstacks.example.com/api/plugins/alerts/ingest/webhook?org_id=ORG_ID" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -d '{
    "source": "prometheus",
    "severity": "page",
    "summary": "BGP peer 10.0.0.2 down on core-rtr-01",
    "details": {
      "protocol": "bgp",
      "neighbor": "10.0.0.2",
      "asn": "65001"
    },
    "device_hostname": "core-rtr-01.dc1"
  }'

A new alert returns HTTP 201 with the full alert record. A deduplicated alert returns the existing record with an incremented occurrence_count and does not re-trigger routing. To test your configuration, send a real low-severity alert to this endpoint and inspect the result and its triage events — there is no separate test endpoint.

Kafka consumer

The plugin can consume alerts from Kafka topics. Enable it through plugin settings; messages are parsed and fed into the same severity, dedup, and routing pipeline as webhook alerts.

kafka-settings.jsonjson
{
  "kafka.enabled": true,
  "kafka.bootstrap_servers": "kafka:9092",
  "kafka.topics": "network-alerts",
  "kafka.consumer_group": "netstacks-alerts",
  "kafka.security_protocol": "PLAINTEXT",
  "kafka.sasl_mechanism": "",
  "kafka.sasl_username": "",
  "kafka.sasl_password": ""
}

Defaults: topics network-alerts, consumer group netstacks-alerts, security protocol PLAINTEXT. For authenticated brokers set kafka.security_protocol to SASL_SSL and supply the SASL mechanism, username, and password.

SNMP traps

A trap sidecar can receive SNMP traps (default listener port 162, configurable via snmp.trap_port) and forward them as webhook alerts to the plugin's internal /ingest/webhook URL. Relevant settings include snmp.webhook_url, snmp.default_org_id, snmp.community_filter, and snmp.custom_oid_mappings.

Syslog and the direct SNMP route

The direct /ingest/syslog and /ingest/snmp HTTP routes are present but currently return 501 Not Implemented. For SNMP, use the trap sidecar (which forwards to the webhook endpoint). For syslog, forward through any tool that can emit an HTTP webhook into /ingest/webhook.

Severity Mapping & Deduplication

Severity mappings

Normalized severity is always one of critical, warning, or info. Map each source's own severity strings (for example Prometheus page, Datadog error) to one of these levels. Mappings are per-source and managed at /api/plugins/alerts/admin/severity-mappings.

severity-mapping.shbash
# Map Prometheus "page" severity to critical
curl -X POST \
  "https://netstacks.example.com/api/plugins/alerts/admin/severity-mappings?org_id=ORG_ID" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -d '{
    "source": "prometheus",
    "source_severity": "page",
    "mapped_severity": "critical"
  }'

When no explicit mapping exists, the plugin infers a level from common keywords (for example critical, emergency, fatal → critical; warning, error, major → warning), falling back to info.

Deduplication configs

Deduplication is configured separately from routing, at /api/plugins/alerts/admin/dedup-configs. A config sets the window_seconds (default 300, minimum 60) and the fingerprint_fields that define what counts as a duplicate (default ["source", "summary"]). You can scope a config to matching sources or summaries with source_pattern and summary_pattern.

dedup-config.shbash
# Create a dedup config: 10-minute window keyed on source + neighbor
curl -X POST \
  "https://netstacks.example.com/api/plugins/alerts/admin/dedup-configs?org_id=ORG_ID" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -d '{
    "name": "BGP peer flaps",
    "source_pattern": "prometheus",
    "fingerprint_fields": ["source", "summary"],
    "window_seconds": 600
  }'

The default plugin-wide deduplication window is deduplication_window_seconds (default 300) and the default severity is default_severity (default warning), both set in plugin settings.

Routing Rules

Routing rules decide what happens to each new alert. Rules are evaluated in ascending priority (lower number runs first), and the first rule whose conditions all match wins. If no rule matches, the alert is routed to the default triage agent. Manage rules at /api/plugins/alerts/admin/routing-rules.

Rule shape

A routing rule has these fields:

  • name, description — identity.
  • priority — integer (lower = evaluated first).
  • enabled — whether the rule participates.
  • conditions — a flat list, combined with AND logic. Each condition has field, operator, and value.
  • action — one of route_to_agent, fast_path, or suppress.
  • agent_id — the triage agent for route_to_agent (omit to use the default agent).
  • fast_path_config — settings for the fast_path action.

Conditions

A condition's field reads a top-level alert field (such as severity, source, summary) or a nested value via dot-path into details (for example details.protocol). Supported operators: eq, neq, contains, not_contains, regex, gt, lt, in, and not_in.

Route critical BGP alerts to a triage agent

rule-route-to-agent.jsonjson
{
  "name": "Critical BGP to network triage agent",
  "description": "Send critical BGP alerts to the AI triage agent",
  "priority": 10,
  "enabled": true,
  "conditions": [
    { "field": "severity", "operator": "eq", "value": "critical" },
    { "field": "details.protocol", "operator": "eq", "value": "bgp" }
  ],
  "action": "route_to_agent",
  "agent_id": "AGENT_UUID"
}

Fast-path: auto-acknowledge low-value info alerts

The fast_path action applies deterministic updates and skips AI triage. fast_path_config keys include auto_acknowledge, set_severity, suppress_minutes, and add_tags.

rule-fast-path.jsonjson
{
  "name": "Auto-ack informational link-state notices",
  "priority": 20,
  "enabled": true,
  "conditions": [
    { "field": "severity", "operator": "eq", "value": "info" },
    { "field": "summary", "operator": "contains", "value": "link up" }
  ],
  "action": "fast_path",
  "fast_path_config": {
    "auto_acknowledge": true,
    "add_tags": { "category": "link-state" }
  }
}

Suppress known noise

rule-suppress.jsonjson
{
  "name": "Suppress maintenance-window chatter",
  "priority": 5,
  "enabled": true,
  "conditions": [
    { "field": "source", "operator": "eq", "value": "lab-monitor" }
  ],
  "action": "suppress"
}

A suppress rule sets the alert state to suppressed and triage_state to skipped. Put low-priority (low-number) suppress rules first so noise is dropped before more general routing rules run.

Tip

After creating rules, send a representative alert to /ingest/webhook and check the returned record plus its triage events to confirm which rule matched and what action was taken. Each rule also tracks match_count and last_matched_at.

Notification Channels & Escalation

Notification channels

Channels deliver alert notifications. The supported channel_type values are email, webhook, slack, and pagerduty. Each channel stores a type-specific config object. Manage channels at /api/plugins/alerts/admin/pipeline/channels.

channel-slack.shbash
# Slack channel (incoming webhook)
curl -X POST \
  "https://netstacks.example.com/api/plugins/alerts/admin/pipeline/channels?org_id=ORG_ID" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -d '{
    "name": "noc-slack",
    "channel_type": "slack",
    "config": { "webhook_url": "https://hooks.slack.com/services/XXX/YYY/ZZZ" }
  }'

Config keys by channel type:

email
smtp_host, smtp_port (default 587), from_address, to_addresses (list), tls (default true), optional username / password.
webhook
url and optional headers.
slack
webhook_url (Slack incoming webhook).
pagerduty
PagerDuty Events v2 routing key; alert severity is mapped to a PagerDuty severity automatically.

Escalation policies

Escalation policies notify a channel when an alert of a given severity stays unresolved beyond a threshold. A policy has a name, severity, escalate_after_mins, and the target channel_id. Manage them at /api/plugins/alerts/admin/escalation-policies; a background worker evaluates them on a schedule.

escalation-policy.shbash
# Escalate unresolved critical alerts to PagerDuty after 15 minutes
curl -X POST \
  "https://netstacks.example.com/api/plugins/alerts/admin/escalation-policies?org_id=ORG_ID" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -d '{
    "name": "Critical -> PagerDuty",
    "severity": "critical",
    "escalate_after_mins": 15,
    "channel_id": "CHANNEL_UUID",
    "enabled": true
  }'

Terminal Panels

The Alerts plugin contributes two panels to the NetStacks Terminal activity bar, defined in its manifest:

  • Alerts (panel id alert-list) — backed by /admin/alerts, refreshing every 30 seconds. Columns: Severity, Source, Summary, State, Device, Created.
  • Pipeline Rules (panel id pipeline-rules) — backed by /admin/pipeline/rules, refreshing every 60 seconds. Columns: Name, Priority, Enabled, Matches.

The Alerts list endpoint supports server-side filtering by state, severity, source, device_id, and triage_state, so you can narrow to, for example, open critical alerts still triaging.

alerts-admin.shbash
# List open critical alerts
curl "https://netstacks.example.com/api/plugins/alerts/admin/alerts?org_id=ORG_ID&state=open&severity=critical" \
  -H "Authorization: Bearer YOUR_API_TOKEN"

# Update an alert's state (acknowledge)
curl -X PUT \
  "https://netstacks.example.com/api/plugins/alerts/admin/alerts/ALERT_ID?org_id=ORG_ID" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -d '{ "state": "acknowledged", "acknowledged_by": "jdoe" }'
Resolve and suppress

The same PUT endpoint accepts state values of resolved and suppressed. Supplying suppressed_until sets a time-bound suppression.

Q&A

Q: What alert sources are supported?
A: HTTP webhooks (the primary, fully-implemented path) and an optional Kafka consumer feed alerts into the pipeline. SNMP traps are received by a sidecar that forwards them to the webhook endpoint. The direct /ingest/syslog and /ingest/snmp HTTP routes currently return 501 Not Implemented.
Q: What is the real webhook endpoint and payload?
A: POST to /api/plugins/alerts/ingest/webhook?org_id=<org> with JSON containing source and summary (required), plus optional severity, details, and either device_id or device_hostname. There is no message, labels, or timestamp field, and no :8443 port.
Q: How does deduplication work?
A: Each alert gets a SHA-256 fingerprint (default org + source + summary). If an open or acknowledged alert with the same fingerprint exists within the dedup window (default 300 seconds), its occurrence_count and last-occurrence timestamp are updated and routing is skipped. Windows and fingerprint fields are configurable via dedup-configs.
Q: How do routing rules decide what happens to an alert?
A: Rules are evaluated in ascending priority (lower first); the first rule whose conditions all match (AND logic) wins. Its action is one of route_to_agent (AI triage), fast_path (deterministic updates, no AI), or suppress. If nothing matches, the alert goes to the default triage agent.
Q: What states can an alert be in?
A: The operational state is open, acknowledged, resolved, or suppressed. Separately, triage_state tracks AI triage: triaging, skipped, or failed.
Q: Which notification channels exist?
A: email, webhook, slack, and pagerduty. Channels are managed at /admin/pipeline/channels, and escalation policies at /admin/escalation-policies notify a channel when an alert of a given severity stays unresolved past a threshold.
Q: How do I test the pipeline?
A: There is no dedicated test endpoint or UI. Send a real low-severity alert to /ingest/webhook and inspect the returned alert record and its triage events (via /admin/triage-events/{alert_id}) to confirm severity mapping, dedup behavior, and which routing rule matched.

Troubleshooting

Alerts not arriving

  • Confirm the Alerts plugin is enabled (it is reached through the controller proxy at /api/plugins/alerts).
  • Verify your request includes the org_id query parameter and a valid bearer token.
  • Check the payload validates: source and summary are required.
  • If you are posting to /ingest/syslog or /ingest/snmp and getting 501, switch to the webhook endpoint (or the SNMP trap sidecar).

Severity looks wrong

  • Add an explicit severity mapping for the source/severity pair at /admin/severity-mappings rather than relying on keyword inference.
  • Remember unmapped alerts with no severity default to info.

Deduplication too aggressive or too loose

  • Adjust window_seconds on the dedup config — shorter windows create more distinct alerts, longer windows merge more.
  • Tune fingerprint_fields to control what counts as a duplicate.
  • Note: only open and acknowledged alerts are dedup targets — resolved or suppressed alerts will not absorb new occurrences.

Routing rules not matching

  • Check rule priority — a lower-number rule may be matching (and stopping evaluation) first.
  • Verify field paths: nested values need the details. prefix (for example details.protocol).
  • Confirm the operator is one of the supported set and the value matches as a string for eq/contains operators.
  • Inspect the alert's triage events to see which rule matched or that the default agent was used.

Notifications not sending

  • Verify the channel config has the required keys for its type (for example webhook url, Slack webhook_url, email smtp_host / from_address / to_addresses).
  • For escalation, confirm the policy is enabled and references a valid channel_id with the correct severity.
  • Plugin System — how Enterprise plugins run as containers and are proxied at /api/plugins/{name}.
  • Incidents & ITSM — alerts can link to incidents during triage.
  • AI Agents — the triage agents that route_to_agent rules dispatch alerts to.
  • Task Monitoring — observe automation triggered from alert triage.