NetStacksNetStacks

Web Sources

Enterprise

Crawl documentation sites with a same-origin BFS crawler and ingest the extracted content into the NetStacks Knowledge Base for vector search.

Overview

Web Sources crawl web-based documentation and ingest the extracted content into the NetStacks Knowledge Base. Instead of uploading files or adding URLs one at a time, you point a web source at a documentation site and NetStacks runs a breadth-first crawl, extracts the main content of each page, converts it to markdown-style text, and feeds it through the standard Knowledge Base ingestion pipeline (chunking, embedding, pgvector storage).

This is ideal for ingesting vendor documentation (Cisco, Juniper, Arista), internal wikis, and any web-accessible reference material your team consults during troubleshooting.

What the crawler does

CapabilityBehavior
Same-origin BFS crawlStarts from the configured URL and follows in-scope links breadth-first up to max_depth and max_pages.
Scope enforcementA link is in scope only if its scheme, host, and port match the source URL. An optional path_prefix further restricts crawling to a subtree (e.g. /eos/latest).
URL normalizationFragments (#...) are stripped, empty query strings are removed, and trailing slashes are dropped (except the root path) so the same page is not indexed twice.
Main-content extractionPicks the first matching content container and strips script, style, nav, header, footer, aside, and noscript.
Change detectionEach page's extracted text is hashed with SHA-256. Unchanged pages are skipped on re-crawl; changed pages are re-embedded.
Stale-page cleanupAfter a crawl that completes fully, documents whose URLs were no longer discovered are deleted from the Knowledge Base.
Category & tagsAn optional category and tag list are applied to every document created by the source, for filtered search.
Enterprise (Controller) feature

Web Sources run inside the NetStacks Controller (Enterprise) and extend the Knowledge Base with automated crawling alongside file uploads and individual URL sources. Endpoints live under /api/knowledge/web-sources and require an authenticated operator. NetStacks Free does not include the Knowledge Base.

How It Works

Crawl pipeline

When you trigger a crawl, the Controller spawns a background task that runs this loop for the web source:

  1. Seed the queue — The source URL is normalized and pushed onto a breadth-first queue at depth 0. The crawler uses the User-Agent NetStacks-Crawler/1.0, a 30-second per-request timeout, and follows up to 5 redirects.
  2. Fetch and validate — For each URL the crawler issues an HTTP GET. It skips the page (counting it as failed) if the response redirects out of scope, the Content-Type is not text/html, or the body exceeds the 5 MB cap. A fetch failure on the very first (root) URL aborts the whole crawl with an error.
  3. Extract content — The HTML is parsed and the first matching container is selected from main, article, [role='main'], #content, .content, then body. The DOM is walked and converted to markdown-style text (headings, lists, code fences, simple tables), skipping script/style/nav/header/footer/aside/noscript. Pages with no extractable content are counted as failed and skipped.
  4. Hash and compare — The extracted text is hashed with SHA-256. If a document for that URL already exists with the same hash it is skipped; if the hash differs, the document and its chunks are updated and re-embedded; if no document exists, a new one is created and embedded.
  5. Discover links — If the page is below max_depth, in-scope links are extracted (honoring <base href>, skipping javascript:/mailto:/tel:/data:), normalized, deduplicated, and queued at the next depth.
  6. Rate limit — The crawler sleeps rate_limit_ms milliseconds between requests (default 1000 ms).
  7. Finalize — When the queue empties (a full crawl) the crawler deletes any previously-ingested pages that were not seen this run. Crawl progress (pages_crawled) is persisted throughout so the list view stays current.

Crawl states

A web source's state field is one of three values:

  • idle — not currently crawling (the default after creation and after a crawl finishes).
  • crawling — a background crawl is in progress. Triggering a crawl while in this state returns 409 Conflict.
  • failed — the last crawl failed; the reason is stored in last_error.

Progress is tracked via pages_crawled, and last_crawled_at records the most recent run. On Controller startup, sources stuck in crawling (e.g. from a crash) are reset.

Use path_prefix to stay focused

Set path_prefix to scope the crawl to a documentation subtree. For example, a source URL of https://docs.example.com/eos/latest/ with path_prefix /eos/latest keeps the crawler out of blogs, release notes, and unrelated product sections on the same host. Combine this with a modest max_depth (2–3) for most sites.

Configuration Reference

These fields are accepted when creating or updating a web source. The Controller clamps numeric values to safe ranges, so out-of-range inputs are silently corrected rather than rejected.

FieldTypeDefaultNotes
namestringrequiredDisplay name for the source.
urlstringrequiredStarting URL. Must parse as a valid URL or the request returns 400. Defines the crawl origin (scheme + host + port).
path_prefixstringnoneOptional. Restricts crawling to URLs whose path starts with this prefix.
categorystringnoneOptional category applied to ingested documents.
tagsstring[]noneOptional tags applied to ingested documents.
max_pagesinteger100Clamped to 1–10000.
max_depthinteger3Clamped to 0–10. Depth 0 crawls only the start URL.
rate_limit_msinteger1000Delay between requests, clamped to 100–60000 ms.
allow_insecure_sslbooleanfalseWhen true, accepts invalid/self-signed TLS certificates on the target site.
Large sites

The crawler follows every in-scope link within the configured depth, which can expand quickly on large sites. Keep max_pages conservative (the default is 100; the hard cap is 10000) and use path_prefix to limit scope before raising max_depth.

Step-by-Step Guide

Workflow 1: Add and crawl a web source

  1. Open the Knowledge Base — In the Controller admin UI, go to the Knowledge Base area and select Web Sources.
  2. Add a web source — Provide a name and the starting URL (for example https://docs.arista.com/eos/latest/).
  3. Scope and limit the crawl — Optionally set a path prefix, then set max_depth (e.g. 3) and max_pages (e.g. 500), and adjust the per-request rate limit if needed.
  4. Add a category and tags — e.g. arista, eos, switching for filtered search.
  5. Trigger the crawl — Start the crawl. It runs in the background; the source moves to crawling and the pages-crawled counter updates as it progresses.

Workflow 2: Monitor progress

  1. Watch the state and counter — The source shows idle, crawling, or failed, plus the pages_crawled count.
  2. Inspect ingested documents — List the documents that belong to the source to confirm content extracted correctly.
  3. Check errors on failure — If the state is failed, read last_error for the cause.

Workflow 3: Re-crawl to pick up changes

  1. Trigger the crawl again — Re-run the crawl on the existing source.
  2. Only changes are re-processed — Pages with an unchanged SHA-256 content hash are skipped; changed pages are re-embedded and new pages are added.
  3. Stale pages are removed — After a full crawl, documents whose URLs were no longer discovered are deleted from the Knowledge Base.

API Examples

Base path

All endpoints are nested under /api/knowledge/web-sources on the Controller and require an authenticated operator. IDs are UUIDs.

Create a web source

name and url are required; everything else falls back to the defaults in the configuration reference. A successful create returns 201 Created with the full source object.

create-web-source.shbash
curl -X POST https://controller.example.com/api/knowledge/web-sources \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Arista EOS Docs",
    "url": "https://docs.arista.com/eos/latest/",
    "path_prefix": "/eos/latest",
    "category": "vendor-docs",
    "tags": ["arista", "eos", "switching"],
    "max_depth": 3,
    "max_pages": 500,
    "rate_limit_ms": 1000,
    "allow_insecure_ssl": false
  }'
create-response.jsonjson
{
  "id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f",
  "org_id": "1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d",
  "name": "Arista EOS Docs",
  "url": "https://docs.arista.com/eos/latest/",
  "path_prefix": "/eos/latest",
  "category": "vendor-docs",
  "tags": ["arista", "eos", "switching"],
  "max_pages": 500,
  "max_depth": 3,
  "rate_limit_ms": 1000,
  "allow_insecure_ssl": false,
  "state": "idle",
  "last_crawled_at": null,
  "last_error": null,
  "pages_crawled": 0,
  "created_by": "0f1e2d3c-4b5a-6978-8796-a5b4c3d2e1f0",
  "created_at": "2026-03-25T10:30:00Z",
  "updated_at": "2026-03-25T10:30:00Z"
}

Creating a second source with a URL that already exists returns 409 Conflict; an unparseable url returns 400 Bad Request.

Trigger a crawl

Starting a crawl is asynchronous. A successful trigger returns 202 Accepted and the crawl runs in a background task. If the source is already crawling, you get 409 Conflict.

trigger-crawl.shbash
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f

curl -X POST \
  https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID/crawl \
  -H "Authorization: Bearer $TOKEN"
trigger-response.jsonjson
{
  "status": "crawl_started",
  "source_id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f"
}

Check crawl status

Fetch the source to read its current state and progress:

get-web-source.shbash
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f

curl -s \
  https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID \
  -H "Authorization: Bearer $TOKEN" | jq '.'
get-response.jsonjson
{
  "id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f",
  "name": "Arista EOS Docs",
  "url": "https://docs.arista.com/eos/latest/",
  "path_prefix": "/eos/latest",
  "category": "vendor-docs",
  "tags": ["arista", "eos", "switching"],
  "max_pages": 500,
  "max_depth": 3,
  "rate_limit_ms": 1000,
  "allow_insecure_ssl": false,
  "state": "crawling",
  "last_crawled_at": "2026-03-25T10:42:18Z",
  "last_error": null,
  "pages_crawled": 127,
  "created_at": "2026-03-25T10:30:00Z",
  "updated_at": "2026-03-25T10:42:18Z"
}

List all web sources

list-web-sources.shbash
curl -s https://controller.example.com/api/knowledge/web-sources \
  -H "Authorization: Bearer $TOKEN" | jq '.[] | {id, name, state, pages_crawled, max_pages}'
list-response.jsonjson
[
  {
    "id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f",
    "name": "Arista EOS Docs",
    "state": "idle",
    "pages_crawled": 487,
    "max_pages": 500
  },
  {
    "id": "b3c4d5e6-7f80-4192-a3b4-c5d6e7f80910",
    "name": "Juniper Junos Docs",
    "state": "crawling",
    "pages_crawled": 203,
    "max_pages": 1000
  }
]

List the documents a source produced

Each web source owns the documents it ingested. List them to verify extraction:

list-source-documents.shbash
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f

curl -s \
  https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID/documents \
  -H "Authorization: Bearer $TOKEN" | jq '.'

Update crawl settings

Send only the fields you want to change. Numeric values are clamped to the same ranges as on create.

update-web-source.shbash
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f

curl -X PUT \
  https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "max_pages": 1000,
    "rate_limit_ms": 2000
  }'

Delete a web source

Returns 204 No Content on success and removes the source.

delete-web-source.shbash
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f

curl -X DELETE \
  https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID \
  -H "Authorization: Bearer $TOKEN" -i

Questions & Answers

Which pages does the crawler follow?
Only same-origin pages: a link is in scope when its scheme, host, and port match the source URL. If you set a path_prefix, the link's path must also start with that prefix. Cross-host links, and links with javascript:, mailto:, tel:, or data: schemes, are ignored.
How deep should I crawl?
A max_depth of 2–3 is usually enough for well-structured documentation sites. The cap is 10, and depth 0 crawls only the start URL. Combine depth with path_prefix to avoid wandering into unrelated areas.
Is there rate limiting?
Yes. The crawler sleeps rate_limit_ms milliseconds between requests, defaulting to 1000 ms and clamped to 100–60000 ms. There is no per-host concurrency — pages are fetched sequentially.
Does the crawler obey robots.txt?
No. The current crawler does not fetch or honor robots.txt. It identifies itself with the User-Agent NetStacks-Crawler/1.0 and relies on scope rules, max_pages, max_depth, and the rate limit to stay polite. Only crawl sites you are authorized to crawl.
What about JavaScript-rendered sites?
The crawler fetches raw HTML and does not execute JavaScript. Single-page apps and client-rendered sites may yield little or no extractable content. For those, export to a file and use the file upload source, or add specific pages individually.
What page types are skipped?
Responses that are not text/html, bodies larger than 5 MB, redirects that leave the configured scope, and pages with no extractable main content are all skipped (and counted as failed). The crawl continues past these — only a failure on the very first (root) URL aborts the whole run.
Can I crawl HTTPS sites with self-signed certificates?
Yes — set allow_insecure_ssl to true to accept invalid or self-signed TLS certificates. Leave it false (the default) for public sites.
How are duplicates and changes handled?
URLs are normalized (fragments removed, empty queries dropped, trailing slashes trimmed) and deduplicated during the crawl. Each page's extracted text is hashed with SHA-256; unchanged pages are skipped on re-crawl, and changed pages are re-embedded.
What happens to pages that disappear?
When a crawl completes fully (the queue empties without hitting the page limit), documents whose URLs were not seen during that run are deleted from the Knowledge Base, keeping the index in sync with the live site.

Troubleshooting

SymptomLikely causeResolution
Crawl immediately failsThe root URL could not be fetched, or the embedding service is unavailable.A fetch error on the start URL aborts the whole crawl — read last_error and confirm the Controller can curl the URL (DNS, firewall, TLS). A 503 on trigger means the embedding service did not initialize.
Trigger returns 409The source is already in the crawling state.Wait for the current crawl to finish. Sources stuck in crawling after a crash are reset on Controller startup.
Few or no documents ingestedJavaScript-rendered content, or the main content lives outside the recognized containers.The extractor only reads main, article, [role='main'], #content, .content, or body. For SPA-heavy sites, export to a file and use file upload, or add pages as individual URL sources.
Crawl stops short of the sitemax_pages or max_depth reached, or a path_prefix excluded the rest.Raise max_pages (cap 10000) or max_depth (cap 10), or widen path_prefix. Remember scope is limited to the same scheme, host, and port.
Pages skipped unexpectedlyNon-HTML Content-Type, body over 5 MB, or an out-of-scope redirect.These are skipped by design. Check the target page's response headers; redirects to another host or path leave scope and are dropped.
TLS errors on an internal siteSelf-signed or otherwise invalid certificate.Set allow_insecure_ssl to true on the source, or install a trusted certificate on the target.
Re-crawl re-processes everythingDynamic content in the main body changes the SHA-256 hash each run.Boilerplate (nav, header, footer, aside) is stripped, but timestamps or tokens inside the main content will still alter the hash. There is no per-page change suppression beyond the hash comparison.
  • Knowledge Base — The parent feature web sources feed into. Manage all document sources, chunking, and vector search.
  • LLM Configuration — Configure the embedding provider used to vectorize crawled content and the LLMs that power RAG answers.
  • AI Chat — Uses Knowledge Base content (including web sources) through RAG to answer grounded in your documentation.
  • Knowledge Packs — Curated, distributable knowledge bundles that complement crawled web content.