Web Sources
EnterpriseCrawl documentation sites with a same-origin BFS crawler and ingest the extracted content into the NetStacks Knowledge Base for vector search.
Overview
Web Sources crawl web-based documentation and ingest the extracted content into the NetStacks Knowledge Base. Instead of uploading files or adding URLs one at a time, you point a web source at a documentation site and NetStacks runs a breadth-first crawl, extracts the main content of each page, converts it to markdown-style text, and feeds it through the standard Knowledge Base ingestion pipeline (chunking, embedding, pgvector storage).
This is ideal for ingesting vendor documentation (Cisco, Juniper, Arista), internal wikis, and any web-accessible reference material your team consults during troubleshooting.
What the crawler does
| Capability | Behavior |
|---|---|
| Same-origin BFS crawl | Starts from the configured URL and follows in-scope links breadth-first up to max_depth and max_pages. |
| Scope enforcement | A link is in scope only if its scheme, host, and port match the source URL. An optional path_prefix further restricts crawling to a subtree (e.g. /eos/latest). |
| URL normalization | Fragments (#...) are stripped, empty query strings are removed, and trailing slashes are dropped (except the root path) so the same page is not indexed twice. |
| Main-content extraction | Picks the first matching content container and strips script, style, nav, header, footer, aside, and noscript. |
| Change detection | Each page's extracted text is hashed with SHA-256. Unchanged pages are skipped on re-crawl; changed pages are re-embedded. |
| Stale-page cleanup | After a crawl that completes fully, documents whose URLs were no longer discovered are deleted from the Knowledge Base. |
| Category & tags | An optional category and tag list are applied to every document created by the source, for filtered search. |
Web Sources run inside the NetStacks Controller (Enterprise) and extend the Knowledge Base with automated crawling alongside file uploads and individual URL sources. Endpoints live under /api/knowledge/web-sources and require an authenticated operator. NetStacks Free does not include the Knowledge Base.
How It Works
Crawl pipeline
When you trigger a crawl, the Controller spawns a background task that runs this loop for the web source:
- Seed the queue — The source URL is normalized and pushed onto a breadth-first queue at depth 0. The crawler uses the User-Agent
NetStacks-Crawler/1.0, a 30-second per-request timeout, and follows up to 5 redirects. - Fetch and validate — For each URL the crawler issues an HTTP GET. It skips the page (counting it as failed) if the response redirects out of scope, the
Content-Typeis nottext/html, or the body exceeds the 5 MB cap. A fetch failure on the very first (root) URL aborts the whole crawl with an error. - Extract content — The HTML is parsed and the first matching container is selected from
main,article,[role='main'],#content,.content, thenbody. The DOM is walked and converted to markdown-style text (headings, lists, code fences, simple tables), skippingscript/style/nav/header/footer/aside/noscript. Pages with no extractable content are counted as failed and skipped. - Hash and compare — The extracted text is hashed with SHA-256. If a document for that URL already exists with the same hash it is skipped; if the hash differs, the document and its chunks are updated and re-embedded; if no document exists, a new one is created and embedded.
- Discover links — If the page is below
max_depth, in-scope links are extracted (honoring<base href>, skippingjavascript:/mailto:/tel:/data:), normalized, deduplicated, and queued at the next depth. - Rate limit — The crawler sleeps
rate_limit_msmilliseconds between requests (default 1000 ms). - Finalize — When the queue empties (a full crawl) the crawler deletes any previously-ingested pages that were not seen this run. Crawl progress (
pages_crawled) is persisted throughout so the list view stays current.
Crawl states
A web source's state field is one of three values:
idle— not currently crawling (the default after creation and after a crawl finishes).crawling— a background crawl is in progress. Triggering a crawl while in this state returns409 Conflict.failed— the last crawl failed; the reason is stored inlast_error.
Progress is tracked via pages_crawled, and last_crawled_at records the most recent run. On Controller startup, sources stuck in crawling (e.g. from a crash) are reset.
Set path_prefix to scope the crawl to a documentation subtree. For example, a source URL of https://docs.example.com/eos/latest/ with path_prefix /eos/latest keeps the crawler out of blogs, release notes, and unrelated product sections on the same host. Combine this with a modest max_depth (2–3) for most sites.
Configuration Reference
These fields are accepted when creating or updating a web source. The Controller clamps numeric values to safe ranges, so out-of-range inputs are silently corrected rather than rejected.
| Field | Type | Default | Notes |
|---|---|---|---|
name | string | required | Display name for the source. |
url | string | required | Starting URL. Must parse as a valid URL or the request returns 400. Defines the crawl origin (scheme + host + port). |
path_prefix | string | none | Optional. Restricts crawling to URLs whose path starts with this prefix. |
category | string | none | Optional category applied to ingested documents. |
tags | string[] | none | Optional tags applied to ingested documents. |
max_pages | integer | 100 | Clamped to 1–10000. |
max_depth | integer | 3 | Clamped to 0–10. Depth 0 crawls only the start URL. |
rate_limit_ms | integer | 1000 | Delay between requests, clamped to 100–60000 ms. |
allow_insecure_ssl | boolean | false | When true, accepts invalid/self-signed TLS certificates on the target site. |
The crawler follows every in-scope link within the configured depth, which can expand quickly on large sites. Keep max_pages conservative (the default is 100; the hard cap is 10000) and use path_prefix to limit scope before raising max_depth.
Step-by-Step Guide
Workflow 1: Add and crawl a web source
- Open the Knowledge Base — In the Controller admin UI, go to the Knowledge Base area and select Web Sources.
- Add a web source — Provide a name and the starting URL (for example
https://docs.arista.com/eos/latest/). - Scope and limit the crawl — Optionally set a path prefix, then set
max_depth(e.g. 3) andmax_pages(e.g. 500), and adjust the per-request rate limit if needed. - Add a category and tags — e.g.
arista,eos,switchingfor filtered search. - Trigger the crawl — Start the crawl. It runs in the background; the source moves to
crawlingand the pages-crawled counter updates as it progresses.
Workflow 2: Monitor progress
- Watch the state and counter — The source shows
idle,crawling, orfailed, plus thepages_crawledcount. - Inspect ingested documents — List the documents that belong to the source to confirm content extracted correctly.
- Check errors on failure — If the state is
failed, readlast_errorfor the cause.
Workflow 3: Re-crawl to pick up changes
- Trigger the crawl again — Re-run the crawl on the existing source.
- Only changes are re-processed — Pages with an unchanged SHA-256 content hash are skipped; changed pages are re-embedded and new pages are added.
- Stale pages are removed — After a full crawl, documents whose URLs were no longer discovered are deleted from the Knowledge Base.
API Examples
All endpoints are nested under /api/knowledge/web-sources on the Controller and require an authenticated operator. IDs are UUIDs.
Create a web source
name and url are required; everything else falls back to the defaults in the configuration reference. A successful create returns 201 Created with the full source object.
curl -X POST https://controller.example.com/api/knowledge/web-sources \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Arista EOS Docs",
"url": "https://docs.arista.com/eos/latest/",
"path_prefix": "/eos/latest",
"category": "vendor-docs",
"tags": ["arista", "eos", "switching"],
"max_depth": 3,
"max_pages": 500,
"rate_limit_ms": 1000,
"allow_insecure_ssl": false
}'{
"id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f",
"org_id": "1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d",
"name": "Arista EOS Docs",
"url": "https://docs.arista.com/eos/latest/",
"path_prefix": "/eos/latest",
"category": "vendor-docs",
"tags": ["arista", "eos", "switching"],
"max_pages": 500,
"max_depth": 3,
"rate_limit_ms": 1000,
"allow_insecure_ssl": false,
"state": "idle",
"last_crawled_at": null,
"last_error": null,
"pages_crawled": 0,
"created_by": "0f1e2d3c-4b5a-6978-8796-a5b4c3d2e1f0",
"created_at": "2026-03-25T10:30:00Z",
"updated_at": "2026-03-25T10:30:00Z"
}Creating a second source with a URL that already exists returns 409 Conflict; an unparseable url returns 400 Bad Request.
Trigger a crawl
Starting a crawl is asynchronous. A successful trigger returns 202 Accepted and the crawl runs in a background task. If the source is already crawling, you get 409 Conflict.
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f
curl -X POST \
https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID/crawl \
-H "Authorization: Bearer $TOKEN"{
"status": "crawl_started",
"source_id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f"
}Check crawl status
Fetch the source to read its current state and progress:
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f
curl -s \
https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID \
-H "Authorization: Bearer $TOKEN" | jq '.'{
"id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f",
"name": "Arista EOS Docs",
"url": "https://docs.arista.com/eos/latest/",
"path_prefix": "/eos/latest",
"category": "vendor-docs",
"tags": ["arista", "eos", "switching"],
"max_pages": 500,
"max_depth": 3,
"rate_limit_ms": 1000,
"allow_insecure_ssl": false,
"state": "crawling",
"last_crawled_at": "2026-03-25T10:42:18Z",
"last_error": null,
"pages_crawled": 127,
"created_at": "2026-03-25T10:30:00Z",
"updated_at": "2026-03-25T10:42:18Z"
}List all web sources
curl -s https://controller.example.com/api/knowledge/web-sources \
-H "Authorization: Bearer $TOKEN" | jq '.[] | {id, name, state, pages_crawled, max_pages}'[
{
"id": "8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f",
"name": "Arista EOS Docs",
"state": "idle",
"pages_crawled": 487,
"max_pages": 500
},
{
"id": "b3c4d5e6-7f80-4192-a3b4-c5d6e7f80910",
"name": "Juniper Junos Docs",
"state": "crawling",
"pages_crawled": 203,
"max_pages": 1000
}
]List the documents a source produced
Each web source owns the documents it ingested. List them to verify extraction:
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f
curl -s \
https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID/documents \
-H "Authorization: Bearer $TOKEN" | jq '.'Update crawl settings
Send only the fields you want to change. Numeric values are clamped to the same ranges as on create.
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f
curl -X PUT \
https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"max_pages": 1000,
"rate_limit_ms": 2000
}'Delete a web source
Returns 204 No Content on success and removes the source.
SOURCE_ID=8f2c1d6e-3b4a-4c5d-9e1f-7a8b9c0d1e2f
curl -X DELETE \
https://controller.example.com/api/knowledge/web-sources/$SOURCE_ID \
-H "Authorization: Bearer $TOKEN" -iQuestions & Answers
- Which pages does the crawler follow?
- Only same-origin pages: a link is in scope when its scheme, host, and port match the source URL. If you set a
path_prefix, the link's path must also start with that prefix. Cross-host links, and links withjavascript:,mailto:,tel:, ordata:schemes, are ignored. - How deep should I crawl?
- A
max_depthof 2–3 is usually enough for well-structured documentation sites. The cap is 10, and depth 0 crawls only the start URL. Combine depth withpath_prefixto avoid wandering into unrelated areas. - Is there rate limiting?
- Yes. The crawler sleeps
rate_limit_msmilliseconds between requests, defaulting to 1000 ms and clamped to 100–60000 ms. There is no per-host concurrency — pages are fetched sequentially. - Does the crawler obey robots.txt?
- No. The current crawler does not fetch or honor
robots.txt. It identifies itself with the User-AgentNetStacks-Crawler/1.0and relies on scope rules,max_pages,max_depth, and the rate limit to stay polite. Only crawl sites you are authorized to crawl. - What about JavaScript-rendered sites?
- The crawler fetches raw HTML and does not execute JavaScript. Single-page apps and client-rendered sites may yield little or no extractable content. For those, export to a file and use the file upload source, or add specific pages individually.
- What page types are skipped?
- Responses that are not
text/html, bodies larger than 5 MB, redirects that leave the configured scope, and pages with no extractable main content are all skipped (and counted as failed). The crawl continues past these — only a failure on the very first (root) URL aborts the whole run. - Can I crawl HTTPS sites with self-signed certificates?
- Yes — set
allow_insecure_ssltotrueto accept invalid or self-signed TLS certificates. Leave itfalse(the default) for public sites. - How are duplicates and changes handled?
- URLs are normalized (fragments removed, empty queries dropped, trailing slashes trimmed) and deduplicated during the crawl. Each page's extracted text is hashed with SHA-256; unchanged pages are skipped on re-crawl, and changed pages are re-embedded.
- What happens to pages that disappear?
- When a crawl completes fully (the queue empties without hitting the page limit), documents whose URLs were not seen during that run are deleted from the Knowledge Base, keeping the index in sync with the live site.
Troubleshooting
| Symptom | Likely cause | Resolution |
|---|---|---|
| Crawl immediately fails | The root URL could not be fetched, or the embedding service is unavailable. | A fetch error on the start URL aborts the whole crawl — read last_error and confirm the Controller can curl the URL (DNS, firewall, TLS). A 503 on trigger means the embedding service did not initialize. |
| Trigger returns 409 | The source is already in the crawling state. | Wait for the current crawl to finish. Sources stuck in crawling after a crash are reset on Controller startup. |
| Few or no documents ingested | JavaScript-rendered content, or the main content lives outside the recognized containers. | The extractor only reads main, article, [role='main'], #content, .content, or body. For SPA-heavy sites, export to a file and use file upload, or add pages as individual URL sources. |
| Crawl stops short of the site | max_pages or max_depth reached, or a path_prefix excluded the rest. | Raise max_pages (cap 10000) or max_depth (cap 10), or widen path_prefix. Remember scope is limited to the same scheme, host, and port. |
| Pages skipped unexpectedly | Non-HTML Content-Type, body over 5 MB, or an out-of-scope redirect. | These are skipped by design. Check the target page's response headers; redirects to another host or path leave scope and are dropped. |
| TLS errors on an internal site | Self-signed or otherwise invalid certificate. | Set allow_insecure_ssl to true on the source, or install a trusted certificate on the target. |
| Re-crawl re-processes everything | Dynamic content in the main body changes the SHA-256 hash each run. | Boilerplate (nav, header, footer, aside) is stripped, but timestamps or tokens inside the main content will still alter the hash. There is no per-page change suppression beyond the hash comparison. |
Related Features
- Knowledge Base — The parent feature web sources feed into. Manage all document sources, chunking, and vector search.
- LLM Configuration — Configure the embedding provider used to vectorize crawled content and the LLMs that power RAG answers.
- AI Chat — Uses Knowledge Base content (including web sources) through RAG to answer grounded in your documentation.
- Knowledge Packs — Curated, distributable knowledge bundles that complement crawled web content.