Alert Pipeline
EnterpriseIngest alerts via webhook and Kafka, normalize and deduplicate them, then route through priority rules to AI triage agents, fast-path actions, or suppression.
Overview
The Alerts plugin (display name Alert Ingestion & Pipeline) ingests alerts from external monitoring systems, normalizes their severity into a common scale, deduplicates repeat occurrences into a single alert, and then runs each new alert through a priority-ordered set of routing rules. A rule can hand the alert to an AI triage agent, apply a fast-path action without AI, or suppress it.
Alert Pipeline is a NetStacks Enterprise plugin. It runs as a container (netstacks-plugin-alerts) on internal port 8080 and is reached through the controller proxy at /api/plugins/alerts/*. There is no separate :8443 listener.
What it does
- Multi-source ingestion — HTTP webhooks and an optional Kafka consumer feed alerts into the pipeline.
- Severity normalization — source-specific severity strings are mapped to
critical,warning, orinfo. - Deduplication — repeat alerts that share a fingerprint within a time window increment an occurrence count instead of creating new alerts.
- Routing & AI triage — priority rules dispatch alerts to AI agents, fast-path actions, or suppression.
- Notifications & escalation — email, webhook, Slack, and PagerDuty channels plus time-based escalation policies.
All plugin endpoints are reached through the authenticated controller proxy and require an org_id query parameter. Do not expose the plugin container port (8080) directly; route through /api/plugins/alerts.
How It Works
Alert processing pipeline
- Ingestion — an alert arrives at
/ingest/webhookor is consumed from a Kafka topic. - Severity mapping — the source-specific severity is looked up in the org's severity mappings and normalized to
critical,warning, orinfo. With no explicit mapping, common keywords are inferred; if no severity is supplied it defaults toinfo. - Device resolution — if a
device_hostnameis provided without adevice_id, the plugin resolves it against the core inventory by hostname or IP. - Fingerprint & deduplication — a SHA-256 fingerprint is computed (default: org + source + summary). If an
openoracknowledgedalert with the same fingerprint exists inside the dedup window, its occurrence count and last-occurrence timestamp are updated and routing is skipped. - Routing — for a genuinely new alert, the routing engine evaluates rules by ascending priority. The first rule whose conditions all match wins; if none match, the alert goes to the default triage agent.
Alert states and triage states
Every alert carries two independent state fields:
state- One of
open,acknowledged,resolved, orsuppressed. This is the human/operational lifecycle, updated by operators or by fast-path and suppress rules. triage_state- The AI triage lifecycle:
triaging(an agent is running),skipped(fast-path or suppress, no AI), orfailed(routing or agent dispatch errored). Alerts also expose triage outputs such asroot_cause,impact_summary,resolution, andresolved_by_agent.
Each step the engine takes is recorded as a triage event (ingested, routing_matched, routing_skipped, agent_started, agent_failed). Retrieve them per alert at /admin/triage-events/{alert_id}.
Ingestion Sources
Webhook ingestion
Post alerts to /api/plugins/alerts/ingest/webhook?org_id=<org>. The payload model accepts source (required), severity (optional, source-specific — it will be mapped), summary (required), details (a free-form object), and either device_id or device_hostname to associate the alert with a device.
# Ingest an alert via the controller proxy
curl -X POST \
"https://netstacks.example.com/api/plugins/alerts/ingest/webhook?org_id=ORG_ID" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-d '{
"source": "prometheus",
"severity": "page",
"summary": "BGP peer 10.0.0.2 down on core-rtr-01",
"details": {
"protocol": "bgp",
"neighbor": "10.0.0.2",
"asn": "65001"
},
"device_hostname": "core-rtr-01.dc1"
}'A new alert returns HTTP 201 with the full alert record. A deduplicated alert returns the existing record with an incremented occurrence_count and does not re-trigger routing. To test your configuration, send a real low-severity alert to this endpoint and inspect the result and its triage events — there is no separate test endpoint.
Kafka consumer
The plugin can consume alerts from Kafka topics. Enable it through plugin settings; messages are parsed and fed into the same severity, dedup, and routing pipeline as webhook alerts.
{
"kafka.enabled": true,
"kafka.bootstrap_servers": "kafka:9092",
"kafka.topics": "network-alerts",
"kafka.consumer_group": "netstacks-alerts",
"kafka.security_protocol": "PLAINTEXT",
"kafka.sasl_mechanism": "",
"kafka.sasl_username": "",
"kafka.sasl_password": ""
}Defaults: topics network-alerts, consumer group netstacks-alerts, security protocol PLAINTEXT. For authenticated brokers set kafka.security_protocol to SASL_SSL and supply the SASL mechanism, username, and password.
SNMP traps
A trap sidecar can receive SNMP traps (default listener port 162, configurable via snmp.trap_port) and forward them as webhook alerts to the plugin's internal /ingest/webhook URL. Relevant settings include snmp.webhook_url, snmp.default_org_id, snmp.community_filter, and snmp.custom_oid_mappings.
The direct /ingest/syslog and /ingest/snmp HTTP routes are present but currently return 501 Not Implemented. For SNMP, use the trap sidecar (which forwards to the webhook endpoint). For syslog, forward through any tool that can emit an HTTP webhook into /ingest/webhook.
Severity Mapping & Deduplication
Severity mappings
Normalized severity is always one of critical, warning, or info. Map each source's own severity strings (for example Prometheus page, Datadog error) to one of these levels. Mappings are per-source and managed at /api/plugins/alerts/admin/severity-mappings.
# Map Prometheus "page" severity to critical
curl -X POST \
"https://netstacks.example.com/api/plugins/alerts/admin/severity-mappings?org_id=ORG_ID" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-d '{
"source": "prometheus",
"source_severity": "page",
"mapped_severity": "critical"
}'When no explicit mapping exists, the plugin infers a level from common keywords (for example critical, emergency, fatal → critical; warning, error, major → warning), falling back to info.
Deduplication configs
Deduplication is configured separately from routing, at /api/plugins/alerts/admin/dedup-configs. A config sets the window_seconds (default 300, minimum 60) and the fingerprint_fields that define what counts as a duplicate (default ["source", "summary"]). You can scope a config to matching sources or summaries with source_pattern and summary_pattern.
# Create a dedup config: 10-minute window keyed on source + neighbor
curl -X POST \
"https://netstacks.example.com/api/plugins/alerts/admin/dedup-configs?org_id=ORG_ID" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-d '{
"name": "BGP peer flaps",
"source_pattern": "prometheus",
"fingerprint_fields": ["source", "summary"],
"window_seconds": 600
}'The default plugin-wide deduplication window is deduplication_window_seconds (default 300) and the default severity is default_severity (default warning), both set in plugin settings.
Routing Rules
Routing rules decide what happens to each new alert. Rules are evaluated in ascending priority (lower number runs first), and the first rule whose conditions all match wins. If no rule matches, the alert is routed to the default triage agent. Manage rules at /api/plugins/alerts/admin/routing-rules.
Rule shape
A routing rule has these fields:
name,description— identity.priority— integer (lower = evaluated first).enabled— whether the rule participates.conditions— a flat list, combined with AND logic. Each condition hasfield,operator, andvalue.action— one ofroute_to_agent,fast_path, orsuppress.agent_id— the triage agent forroute_to_agent(omit to use the default agent).fast_path_config— settings for thefast_pathaction.
Conditions
A condition's field reads a top-level alert field (such as severity, source, summary) or a nested value via dot-path into details (for example details.protocol). Supported operators: eq, neq, contains, not_contains, regex, gt, lt, in, and not_in.
Route critical BGP alerts to a triage agent
{
"name": "Critical BGP to network triage agent",
"description": "Send critical BGP alerts to the AI triage agent",
"priority": 10,
"enabled": true,
"conditions": [
{ "field": "severity", "operator": "eq", "value": "critical" },
{ "field": "details.protocol", "operator": "eq", "value": "bgp" }
],
"action": "route_to_agent",
"agent_id": "AGENT_UUID"
}Fast-path: auto-acknowledge low-value info alerts
The fast_path action applies deterministic updates and skips AI triage. fast_path_config keys include auto_acknowledge, set_severity, suppress_minutes, and add_tags.
{
"name": "Auto-ack informational link-state notices",
"priority": 20,
"enabled": true,
"conditions": [
{ "field": "severity", "operator": "eq", "value": "info" },
{ "field": "summary", "operator": "contains", "value": "link up" }
],
"action": "fast_path",
"fast_path_config": {
"auto_acknowledge": true,
"add_tags": { "category": "link-state" }
}
}Suppress known noise
{
"name": "Suppress maintenance-window chatter",
"priority": 5,
"enabled": true,
"conditions": [
{ "field": "source", "operator": "eq", "value": "lab-monitor" }
],
"action": "suppress"
}A suppress rule sets the alert state to suppressed and triage_state to skipped. Put low-priority (low-number) suppress rules first so noise is dropped before more general routing rules run.
After creating rules, send a representative alert to /ingest/webhook and check the returned record plus its triage events to confirm which rule matched and what action was taken. Each rule also tracks match_count and last_matched_at.
Notification Channels & Escalation
Notification channels
Channels deliver alert notifications. The supported channel_type values are email, webhook, slack, and pagerduty. Each channel stores a type-specific config object. Manage channels at /api/plugins/alerts/admin/pipeline/channels.
# Slack channel (incoming webhook)
curl -X POST \
"https://netstacks.example.com/api/plugins/alerts/admin/pipeline/channels?org_id=ORG_ID" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-d '{
"name": "noc-slack",
"channel_type": "slack",
"config": { "webhook_url": "https://hooks.slack.com/services/XXX/YYY/ZZZ" }
}'Config keys by channel type:
emailsmtp_host,smtp_port(default 587),from_address,to_addresses(list),tls(default true), optionalusername/password.webhookurland optionalheaders.slackwebhook_url(Slack incoming webhook).pagerduty- PagerDuty Events v2 routing key; alert severity is mapped to a PagerDuty severity automatically.
Escalation policies
Escalation policies notify a channel when an alert of a given severity stays unresolved beyond a threshold. A policy has a name, severity, escalate_after_mins, and the target channel_id. Manage them at /api/plugins/alerts/admin/escalation-policies; a background worker evaluates them on a schedule.
# Escalate unresolved critical alerts to PagerDuty after 15 minutes
curl -X POST \
"https://netstacks.example.com/api/plugins/alerts/admin/escalation-policies?org_id=ORG_ID" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-d '{
"name": "Critical -> PagerDuty",
"severity": "critical",
"escalate_after_mins": 15,
"channel_id": "CHANNEL_UUID",
"enabled": true
}'Terminal Panels
The Alerts plugin contributes two panels to the NetStacks Terminal activity bar, defined in its manifest:
- Alerts (panel id
alert-list) — backed by/admin/alerts, refreshing every 30 seconds. Columns: Severity, Source, Summary, State, Device, Created. - Pipeline Rules (panel id
pipeline-rules) — backed by/admin/pipeline/rules, refreshing every 60 seconds. Columns: Name, Priority, Enabled, Matches.
The Alerts list endpoint supports server-side filtering by state, severity, source, device_id, and triage_state, so you can narrow to, for example, open critical alerts still triaging.
# List open critical alerts
curl "https://netstacks.example.com/api/plugins/alerts/admin/alerts?org_id=ORG_ID&state=open&severity=critical" \
-H "Authorization: Bearer YOUR_API_TOKEN"
# Update an alert's state (acknowledge)
curl -X PUT \
"https://netstacks.example.com/api/plugins/alerts/admin/alerts/ALERT_ID?org_id=ORG_ID" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-d '{ "state": "acknowledged", "acknowledged_by": "jdoe" }'The same PUT endpoint accepts state values of resolved and suppressed. Supplying suppressed_until sets a time-bound suppression.
Q&A
- Q: What alert sources are supported?
- A: HTTP webhooks (the primary, fully-implemented path) and an optional Kafka consumer feed alerts into the pipeline. SNMP traps are received by a sidecar that forwards them to the webhook endpoint. The direct
/ingest/syslogand/ingest/snmpHTTP routes currently return501 Not Implemented.
- Q: What is the real webhook endpoint and payload?
- A: POST to
/api/plugins/alerts/ingest/webhook?org_id=<org>with JSON containingsourceandsummary(required), plus optionalseverity,details, and eitherdevice_idordevice_hostname. There is nomessage,labels, ortimestampfield, and no:8443port.
- Q: How does deduplication work?
- A: Each alert gets a SHA-256 fingerprint (default org + source + summary). If an
openoracknowledgedalert with the same fingerprint exists within the dedup window (default 300 seconds), itsoccurrence_countand last-occurrence timestamp are updated and routing is skipped. Windows and fingerprint fields are configurable via dedup-configs.
- Q: How do routing rules decide what happens to an alert?
- A: Rules are evaluated in ascending priority (lower first); the first rule whose conditions all match (AND logic) wins. Its
actionis one ofroute_to_agent(AI triage),fast_path(deterministic updates, no AI), orsuppress. If nothing matches, the alert goes to the default triage agent.
- Q: What states can an alert be in?
- A: The operational
stateisopen,acknowledged,resolved, orsuppressed. Separately,triage_statetracks AI triage:triaging,skipped, orfailed.
- Q: Which notification channels exist?
- A:
email,webhook,slack, andpagerduty. Channels are managed at/admin/pipeline/channels, and escalation policies at/admin/escalation-policiesnotify a channel when an alert of a given severity stays unresolved past a threshold.
- Q: How do I test the pipeline?
- A: There is no dedicated test endpoint or UI. Send a real low-severity alert to
/ingest/webhookand inspect the returned alert record and its triage events (via/admin/triage-events/{alert_id}) to confirm severity mapping, dedup behavior, and which routing rule matched.
Troubleshooting
Alerts not arriving
- Confirm the Alerts plugin is enabled (it is reached through the controller proxy at
/api/plugins/alerts). - Verify your request includes the
org_idquery parameter and a valid bearer token. - Check the payload validates:
sourceandsummaryare required. - If you are posting to
/ingest/syslogor/ingest/snmpand getting501, switch to the webhook endpoint (or the SNMP trap sidecar).
Severity looks wrong
- Add an explicit severity mapping for the source/severity pair at
/admin/severity-mappingsrather than relying on keyword inference. - Remember unmapped alerts with no severity default to
info.
Deduplication too aggressive or too loose
- Adjust
window_secondson the dedup config — shorter windows create more distinct alerts, longer windows merge more. - Tune
fingerprint_fieldsto control what counts as a duplicate. - Note: only
openandacknowledgedalerts are dedup targets — resolved or suppressed alerts will not absorb new occurrences.
Routing rules not matching
- Check rule
priority— a lower-number rule may be matching (and stopping evaluation) first. - Verify
fieldpaths: nested values need thedetails.prefix (for exampledetails.protocol). - Confirm the
operatoris one of the supported set and thevaluematches as a string for eq/contains operators. - Inspect the alert's triage events to see which rule matched or that the default agent was used.
Notifications not sending
- Verify the channel
confighas the required keys for its type (for example webhookurl, Slackwebhook_url, emailsmtp_host/from_address/to_addresses). - For escalation, confirm the policy is enabled and references a valid
channel_idwith the correct severity.
Related Features
- Plugin System — how Enterprise plugins run as containers and are proxied at
/api/plugins/{name}. - Incidents & ITSM — alerts can link to incidents during triage.
- AI Agents — the triage agents that
route_to_agentrules dispatch alerts to. - Task Monitoring — observe automation triggered from alert triage.