NetStacksNetStacks

LLM Configuration

Connect NetStacks to OpenAI, Anthropic, Ollama, OpenRouter, LiteLLM, or any OpenAI-compatible LLM. BYOK keys, models, token budgets, and self-hosting.

Overview

LLM Configuration controls how NetStacks connects to AI language model providers. NetStacks uses a BYOK (Bring Your Own Key) model: you supply your own provider API keys, so you keep full control over cost, rate limits, and data-handling agreements. NetStacks bundles no AI credits and proxies nothing through its own infrastructure.

NetStacks ships in two deployment shapes, and LLM configuration differs between them:

  • Terminal (standalone) — the free, open-source desktop app. Each operator configures their own provider(s) locally. Provider settings (including API keys) are stored as JSON in the app’s local settings, on that machine. There is no organization credential vault, no shared token budget, and no central failover.
  • Controller (Enterprise) — the self-hosted/server deployment. An administrator configures providers centrally. API keys are stored encrypted in the Controller credential vault (referenced by a credential ID, never returned in API responses), and the Controller adds priority-based failover, per-organization token budgets, and usage analytics.
Where keys live

In the standalone Terminal, your API key is saved on your own machine in the local settings file. In the Controller, the key is stored encrypted in the credential vault and AI requests are made server-side, so Terminal clients connected to a Controller never receive the raw key.

Supported Providers

The standalone Terminal supports six provider configurations. The Controller supports five provider types in its llm_provider_type enum — openai, anthropic, ollama, openrouter, and custom. LiteLLM is a first-class option in the Terminal; in the Controller you reach a LiteLLM gateway through the Custom (OpenAI-compatible) type by pointing the base URL at your gateway.

ProviderKey requiredTerminalControllerNotes
OpenAIYesYesYes (openai)Default model gpt-4o; supports a custom base URL (Azure).
AnthropicYesYesYes (anthropic)Claude models; strong reasoning, recommended for agents.
OllamaNoYesYes (ollama)Self-hosted; data never leaves your network.
OpenRouterYesYesYes (openrouter)OpenAI-compatible gateway to many models.
LiteLLMOptionalYesVia CustomTerminal-native. In the Controller, configure as Custom and point the base URL at your LiteLLM proxy.
CustomVariesYesYes (custom)Any OpenAI-compatible endpoint (vLLM, text-generation-inference, Azure OpenAI, a LiteLLM gateway). Terminal Custom also supports OAuth2 client-credentials and a Gemini/Vertex API format.

Default models & endpoints (Terminal)

If you do not override them, the standalone Terminal uses these defaults per provider:

ProviderDefault modelDefault base URL
Anthropicclaude-3-5-sonnet-20241022Provider default
OpenAIgpt-4oProvider default
Ollamallama3.2http://localhost:11434
OpenRouteranthropic/claude-3.5-sonnetProvider default
LiteLLMgpt-4ohttp://localhost:4000
Verify exact model IDs with your provider

Provider model catalogs change frequently. Always set default_model to a model ID your provider currently serves, and use the Test action after saving.

Terminal (Standalone) Setup

In the standalone Terminal, AI providers are configured per-machine. There is no shared vault or org budget — each install holds its own settings.

  1. Open Settings, then the AI tab. The AI tab covers providers, models, MCP servers, sanitization, suggestions, and automation mode.
  2. Choose a provider. Select Anthropic, OpenAI, Ollama, OpenRouter, LiteLLM, or Custom.
  3. Enter credentials. Paste an API key for cloud providers. For Ollama, no key is required — set the base URL (defaults to http://localhost:11434). For LiteLLM, set the gateway URL (defaults to http://localhost:4000); the key is optional.
  4. Pick a model. Set the model ID, e.g. claude-3-5-sonnet-20241022, gpt-4o, or llama3.2.
  5. Save. Settings (including the key) are written to the app’s local settings on this machine. AI features become active immediately.

The AI tab also lets you enable a configuration preview for command suggestions, manage MCP servers, and configure the credential sanitizer that scrubs secrets before any prompt is sent.

One default provider powers every AI feature

The provider/model you set as the default at the top of the AI tab is used by every AI surface — the chat side panel, the floating “Ask AI” pop-overs and hovers, Tab-to-fill suggestions, and background agents. You don’t configure providers per-feature. The only optional override is a specific modelfor the agent toolset (it still uses the default provider). If a pop-over ever reports “not configured” while the side panel works, your default provider simply doesn’t have a key — set the default to the provider you actually configured.

First-run Setup Wizard

On first launch, a Setup Wizard walks you through choosing a provider, getting an API key, and testing the connection. Reopen it anytime from the command palette → “Setup Wizard”.

Controller (Enterprise) Setup Enterprise

Controller (Enterprise)

Everything in this section — central provider configuration, the credential vault, priority failover, token budgets, and usage analytics — applies to the self-hosted Controller. The standalone Terminal does not have these features.

In the Controller admin UI, LLM providers are managed under Settings, AI section, with a dedicated LLM Providers page at /llm-providers.

  1. Open the LLM Providers page (or the AI section of Settings) and choose Add LLM Provider.
  2. Select the provider type — one of openai, anthropic, ollama, openrouter, or custom.
  3. Attach an API key credential. The provider stores an api_key_credential_id that references a credential in the encrypted vault — not the raw key. Ollama without auth needs no credential.
  4. Set the base URL if needed (Ollama URL, Azure OpenAI endpoint, or a LiteLLM/vLLM gateway). Optionally enable accept_invalid_certs for self-signed internal endpoints.
  5. Set the default model and, optionally, max_input_tokens, max_output_tokens, temperature (0.0–2.0), and top_p.
  6. Set priority for failover (lower number = tried first; defaults to 100). Configure multiple enabled providers so requests fail over automatically.
  7. Test the provider to confirm the key, endpoint, and model are reachable, then save.
Key never leaves the server

API key material stays in the Controller vault. List and detail responses expose only has_api_key and the credential display name — never the secret. AI calls are made server-side, so Terminal clients pointed at the Controller never hold provider keys.

Self-Hosting with Ollama

Ollama runs LLMs entirely on your own hardware, so prompts and terminal context never leave your network. It works in both the standalone Terminal and the Controller.

  1. Install Ollama from the Ollama project.
  2. Pull a model, for example ollama pull llama3.2.
  3. Add Ollama as a provider with the base URL of your instance (http://localhost:11434 by default, or an internal host such as http://ollama.internal:11434).
  4. Set the model to a name shown by ollama list (e.g. llama3.2, mistral, qwen2.5).
ollama-quickstart.shbash
# Install (see ollama.ai for your platform), then:
ollama pull llama3.2

# Confirm it serves and lists the model
ollama list

# Test the OpenAI-compatible endpoint Ollama exposes
curl -s http://localhost:11434/api/tags | head
First request is slow

On the first request after start, Ollama loads the model into memory, which can take tens of seconds for large models. Subsequent requests are fast. Ensure the host has enough RAM/VRAM for the model size.

Token Budgets & Analytics Enterprise

Controller (Enterprise) only

Token budgets and usage analytics are Controller features. The standalone Terminal does not enforce organization-wide token budgets.

The Controller token budget has three knobs:

  • Monthly soft limit (monthly_soft_limit_tokens) — a warning threshold; AI keeps working past it.
  • Monthly hard limit (monthly_hard_limit_tokens) — a strict cap.
  • Warning threshold percent (warning_threshold_percent) — the percentage at which a notification fires.

Budget status is read at GET /api/admin/analytics/tokens/budget and updated with PUT to the same path. The status payload reports current_usage_tokens, the configured limits, usage_percent, a status string, and a projected_month_end_tokens estimate.

Usage analytics live under /api/admin/analytics/tokens/ and can be broken down several ways:

  • GET /api/admin/analytics/tokens/timeseries — usage over time.
  • GET /api/admin/analytics/tokens/by-feature — chat vs. agents vs. suggestions vs. embeddings.
  • GET /api/admin/analytics/tokens/by-user — per user.
  • GET /api/admin/analytics/tokens/by-provider — per provider.

The same admin UI exposes a Token Analytics page in the AI navigation group. The Settings, AI section also includes an AI Data Security control for redacting credentials and sensitive data before prompts are sent to an LLM.

API & Config Examples

The JSON examples below match the Controller llm_provider model. All Controller admin endpoints are mounted under /api/admin (there is no /v1 segment).

OpenAI provider (Controller)

openai-provider.jsonjson
{
  "provider_type": "openai",
  "name": "OpenAI Production",
  "enabled": true,
  "priority": 1,
  "api_key_credential_id": "a1b2c3d4-0000-0000-0000-000000000000",
  "default_model": "gpt-4o",
  "max_input_tokens": 128000,
  "max_output_tokens": 4096,
  "temperature": 0.7,
  "top_p": 0.9
}

Anthropic provider (Controller)

anthropic-provider.jsonjson
{
  "provider_type": "anthropic",
  "name": "Anthropic - Agent Reasoning",
  "enabled": true,
  "priority": 2,
  "api_key_credential_id": "e5f6a7b8-0000-0000-0000-000000000000",
  "default_model": "claude-sonnet-4-20250514",
  "max_input_tokens": 200000,
  "max_output_tokens": 8192,
  "temperature": 0.3
}

Ollama provider (Controller)

ollama-provider.jsonjson
{
  "provider_type": "ollama",
  "name": "Local Ollama - Privacy Mode",
  "enabled": true,
  "priority": 3,
  "base_url": "http://ollama.internal:11434",
  "default_model": "llama3.2",
  "max_output_tokens": 4096,
  "temperature": 0.7
}

LiteLLM gateway via the Custom type (Controller)

litellm-as-custom.jsonjson
{
  "provider_type": "custom",
  "name": "LiteLLM Gateway",
  "enabled": true,
  "priority": 4,
  "api_key_credential_id": "c0ffee00-0000-0000-0000-000000000000",
  "base_url": "http://litellm.internal:4000",
  "default_model": "gpt-4o"
}

Terminal provider config (standalone)

The standalone Terminal stores its provider config locally as JSON, tagged by provider. The key lives in the local settings, not a vault.

terminal-ai-provider.jsonjson
{
  "provider": "anthropic",
  "api_key": "sk-ant-...",
  "model": "claude-3-5-sonnet-20241022"
}

// Ollama (no key)
{
  "provider": "ollama",
  "model": "llama3.2",
  "base_url": "http://localhost:11434"
}

// LiteLLM proxy (key optional)
{
  "provider": "litellm",
  "model": "gpt-4o",
  "base_url": "http://localhost:4000"
}

Create a provider (Controller admin API)

create-provider.shbash
# Admin LLM provider routes are nested at /api/admin/llm-providers
curl -X POST https://controller.example.com/api/admin/llm-providers \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "provider_type": "anthropic",
    "name": "Anthropic Primary",
    "priority": 1,
    "api_key_credential_id": "e5f6a7b8-0000-0000-0000-000000000000",
    "default_model": "claude-sonnet-4-20250514"
  }'

# Test a configured provider (replace :id with the returned provider id)
curl -X POST https://controller.example.com/api/admin/llm-providers/$ID/test \
  -H "Authorization: Bearer $TOKEN"

Multi-provider failover (Controller)

failover-setup.shbash
# Primary (priority 1)
curl -X POST https://controller.example.com/api/admin/llm-providers \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"provider_type":"anthropic","name":"Anthropic Primary","priority":1,
       "api_key_credential_id":"...","default_model":"claude-sonnet-4-20250514"}'

# Fallback (priority 2) - tried if the primary is unreachable
curl -X POST https://controller.example.com/api/admin/llm-providers \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"provider_type":"openai","name":"OpenAI Fallback","priority":2,
       "api_key_credential_id":"...","default_model":"gpt-4o"}'

Token budget (Controller)

token-budget.shbash
# Update the monthly budget
curl -X PUT https://controller.example.com/api/admin/analytics/tokens/budget \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{
    "monthly_soft_limit_tokens": 5000000,
    "monthly_hard_limit_tokens": 10000000,
    "warning_threshold_percent": 80
  }'

# Read current budget status
curl -s https://controller.example.com/api/admin/analytics/tokens/budget \
  -H "Authorization: Bearer $TOKEN" | jq '.'

# Response is wrapped under "budget"
{
  "budget": {
    "current_usage_tokens": 3250000,
    "soft_limit_tokens": 5000000,
    "hard_limit_tokens": 10000000,
    "usage_percent": 32.5,
    "status": "ok",
    "projected_month_end_tokens": 7800000
  }
}

Usage by feature (Controller)

usage-by-feature.shbash
# Token usage broken down by feature; start/end are ISO8601 query params
curl -s https://controller.example.com/api/admin/analytics/tokens/by-feature \
  -H "Authorization: Bearer $TOKEN" \
  -G --data-urlencode "start=2026-03-01T00:00:00Z" \
     --data-urlencode "end=2026-04-01T00:00:00Z" | jq '.'

# Response is wrapped under "data"
{
  "data": [
    { "feature": "chat",        "total_tokens": 1850000, "total_cost_cents": 925, "request_count": 342 },
    { "feature": "agents",      "total_tokens":  980000, "total_cost_cents": 490, "request_count": 45 },
    { "feature": "suggestions", "total_tokens":  320000, "total_cost_cents": 32,  "request_count": 1205 },
    { "feature": "embeddings",  "total_tokens":  100000, "total_cost_cents": 1,   "request_count": 89 }
  ]
}

Questions & Answers

Which LLM providers does NetStacks support?
The standalone Terminal supports six configurations: Anthropic, OpenAI, Ollama, OpenRouter, LiteLLM, and Custom (any OpenAI-compatible endpoint). The Controller supports five provider types — openai, anthropic, ollama, openrouter, and custom. A LiteLLM gateway is used in the Controller via the Custom type by pointing the base URL at the gateway.
What is BYOK?
BYOK (Bring Your Own Key) means you supply your own provider API keys. NetStacks includes no bundled AI credits and proxies nothing through its own infrastructure. You create keys directly with your provider and enter them in NetStacks, keeping full control over cost, rate limits, and data-handling terms.
Where are my API keys stored?
In the standalone Terminal, the key is stored locally as part of the AI provider settings JSON on that machine. In the Controller, the key is stored encrypted in the credential vault and referenced by api_key_credential_id; the raw key is never returned by the API, and AI calls run server-side.
Can I run a fully self-hosted LLM?
Yes. Configure Ollama and point it at your local instance (default http://localhost:11434); inference stays on your hardware so data never leaves your network. You can also use the Custom type to reach any OpenAI-compatible endpoint such as vLLM, text-generation-inference, or Azure OpenAI in your own subscription.
How do token budgets work?
Token budgets are a Controller feature. You set a monthly_soft_limit_tokens (warning only), monthly_hard_limit_tokens (strict cap), and a warning_threshold_percent. Status comes from GET /api/admin/analytics/tokens/budget and includes a projected month-end estimate. The standalone Terminal does not enforce org budgets.
What are the correct admin API paths?
Controller admin routes are mounted under /api/admin (no /v1). LLM providers are at /api/admin/llm-providers and token analytics at /api/admin/analytics/tokens/…. The budget response is wrapped under budget and the by-feature response under data.
Is my data sent to external AI providers?
When you use a cloud provider, prompts and terminal context go to that provider’s API — but only after the credential sanitizer scrubs secrets such as Cisco enable secrets, type-7/0 passwords, SNMP community strings, SNMPv3 auth, and Juniper $9$ secrets. For zero external egress, use Ollama with a local model. With a Controller, AI calls are made server-side so Terminal clients never talk to providers directly.

Troubleshooting

IssueLikely causeFix
Provider connection failingInvalid key, wrong endpoint, or blocked egressUse the Test action to diagnose. Confirm the key is current and the base URL is reachable from the machine making the call (Controller for Enterprise, the desktop for standalone). Check outbound HTTPS to the provider.
Self-signed cert rejected (Controller)Internal endpoint uses a self-signed certificateEnable accept_invalid_certs on the provider, or install the endpoint’s CA on the Controller host.
Token budget exceeded (Controller)Monthly usage hit the hard limitCheck GET /api/admin/analytics/tokens/budget for usage and projection. Raise the hard limit, wait for reset, or move high-volume features (suggestions) to a cheaper model. Use the by-feature/by-user breakdowns to find top consumers.
Ollama not respondingService down or model not pulledRun ollama list to confirm the service and model. If missing, run ollama pull <model>. Make sure the base URL matches the Ollama listen address. The first request loads the model (tens of seconds).
Failover not triggering (Controller)Backup disabled or priority misconfiguredEnsure backup providers are enabled with valid credentials and a higher priority number than the primary. Test each provider individually. Failover responds to connection errors and timeouts.
  • AI Chat — conversational network-operations assistance over the configured LLM.
  • AI Agents — autonomous agents that use the configured LLM for multi-step investigation.
  • Command Suggestions — context-aware autocomplete and the credential sanitizer that scrubs secrets before prompts are sent.
  • Knowledge Base — uses provider embedding models for vector search.
  • MCP Servers — extend AI features with Model Context Protocol tools.
  • System Settings — Controller configuration including the credential vault that holds encrypted provider keys.