LLM Configuration
Connect NetStacks to OpenAI, Anthropic, Ollama, OpenRouter, LiteLLM, or any OpenAI-compatible LLM. BYOK keys, models, token budgets, and self-hosting.
Overview
LLM Configuration controls how NetStacks connects to AI language model providers. NetStacks uses a BYOK (Bring Your Own Key) model: you supply your own provider API keys, so you keep full control over cost, rate limits, and data-handling agreements. NetStacks bundles no AI credits and proxies nothing through its own infrastructure.
NetStacks ships in two deployment shapes, and LLM configuration differs between them:
- Terminal (standalone) — the free, open-source desktop app. Each operator configures their own provider(s) locally. Provider settings (including API keys) are stored as JSON in the app’s local settings, on that machine. There is no organization credential vault, no shared token budget, and no central failover.
- Controller (Enterprise) — the self-hosted/server deployment. An administrator configures providers centrally. API keys are stored encrypted in the Controller credential vault (referenced by a credential ID, never returned in API responses), and the Controller adds priority-based failover, per-organization token budgets, and usage analytics.
In the standalone Terminal, your API key is saved on your own machine in the local settings file. In the Controller, the key is stored encrypted in the credential vault and AI requests are made server-side, so Terminal clients connected to a Controller never receive the raw key.
Supported Providers
The standalone Terminal supports six provider configurations. The Controller supports five provider types in its llm_provider_type enum — openai, anthropic, ollama, openrouter, and custom. LiteLLM is a first-class option in the Terminal; in the Controller you reach a LiteLLM gateway through the Custom (OpenAI-compatible) type by pointing the base URL at your gateway.
| Provider | Key required | Terminal | Controller | Notes |
|---|---|---|---|---|
| OpenAI | Yes | Yes | Yes (openai) | Default model gpt-4o; supports a custom base URL (Azure). |
| Anthropic | Yes | Yes | Yes (anthropic) | Claude models; strong reasoning, recommended for agents. |
| Ollama | No | Yes | Yes (ollama) | Self-hosted; data never leaves your network. |
| OpenRouter | Yes | Yes | Yes (openrouter) | OpenAI-compatible gateway to many models. |
| LiteLLM | Optional | Yes | Via Custom | Terminal-native. In the Controller, configure as Custom and point the base URL at your LiteLLM proxy. |
| Custom | Varies | Yes | Yes (custom) | Any OpenAI-compatible endpoint (vLLM, text-generation-inference, Azure OpenAI, a LiteLLM gateway). Terminal Custom also supports OAuth2 client-credentials and a Gemini/Vertex API format. |
Default models & endpoints (Terminal)
If you do not override them, the standalone Terminal uses these defaults per provider:
| Provider | Default model | Default base URL |
|---|---|---|
| Anthropic | claude-3-5-sonnet-20241022 | Provider default |
| OpenAI | gpt-4o | Provider default |
| Ollama | llama3.2 | http://localhost:11434 |
| OpenRouter | anthropic/claude-3.5-sonnet | Provider default |
| LiteLLM | gpt-4o | http://localhost:4000 |
Provider model catalogs change frequently. Always set default_model to a model ID your provider currently serves, and use the Test action after saving.
Terminal (Standalone) Setup
In the standalone Terminal, AI providers are configured per-machine. There is no shared vault or org budget — each install holds its own settings.
- Open Settings, then the AI tab. The AI tab covers providers, models, MCP servers, sanitization, suggestions, and automation mode.
- Choose a provider. Select Anthropic, OpenAI, Ollama, OpenRouter, LiteLLM, or Custom.
- Enter credentials. Paste an API key for cloud providers. For Ollama, no key is required — set the base URL (defaults to
http://localhost:11434). For LiteLLM, set the gateway URL (defaults tohttp://localhost:4000); the key is optional. - Pick a model. Set the model ID, e.g.
claude-3-5-sonnet-20241022,gpt-4o, orllama3.2. - Save. Settings (including the key) are written to the app’s local settings on this machine. AI features become active immediately.
The AI tab also lets you enable a configuration preview for command suggestions, manage MCP servers, and configure the credential sanitizer that scrubs secrets before any prompt is sent.
The provider/model you set as the default at the top of the AI tab is used by every AI surface — the chat side panel, the floating “Ask AI” pop-overs and hovers, Tab-to-fill suggestions, and background agents. You don’t configure providers per-feature. The only optional override is a specific modelfor the agent toolset (it still uses the default provider). If a pop-over ever reports “not configured” while the side panel works, your default provider simply doesn’t have a key — set the default to the provider you actually configured.
On first launch, a Setup Wizard walks you through choosing a provider, getting an API key, and testing the connection. Reopen it anytime from the command palette → “Setup Wizard”.
Controller (Enterprise) Setup Enterprise
Everything in this section — central provider configuration, the credential vault, priority failover, token budgets, and usage analytics — applies to the self-hosted Controller. The standalone Terminal does not have these features.
In the Controller admin UI, LLM providers are managed under Settings, AI section, with a dedicated LLM Providers page at /llm-providers.
- Open the LLM Providers page (or the AI section of Settings) and choose Add LLM Provider.
- Select the provider type — one of
openai,anthropic,ollama,openrouter, orcustom. - Attach an API key credential. The provider stores an
api_key_credential_idthat references a credential in the encrypted vault — not the raw key. Ollama without auth needs no credential. - Set the base URL if needed (Ollama URL, Azure OpenAI endpoint, or a LiteLLM/vLLM gateway). Optionally enable
accept_invalid_certsfor self-signed internal endpoints. - Set the default model and, optionally,
max_input_tokens,max_output_tokens,temperature(0.0–2.0), andtop_p. - Set priority for failover (lower number = tried first; defaults to 100). Configure multiple enabled providers so requests fail over automatically.
- Test the provider to confirm the key, endpoint, and model are reachable, then save.
API key material stays in the Controller vault. List and detail responses expose only has_api_key and the credential display name — never the secret. AI calls are made server-side, so Terminal clients pointed at the Controller never hold provider keys.
Self-Hosting with Ollama
Ollama runs LLMs entirely on your own hardware, so prompts and terminal context never leave your network. It works in both the standalone Terminal and the Controller.
- Install Ollama from the Ollama project.
- Pull a model, for example
ollama pull llama3.2. - Add Ollama as a provider with the base URL of your instance (
http://localhost:11434by default, or an internal host such ashttp://ollama.internal:11434). - Set the model to a name shown by
ollama list(e.g.llama3.2,mistral,qwen2.5).
# Install (see ollama.ai for your platform), then:
ollama pull llama3.2
# Confirm it serves and lists the model
ollama list
# Test the OpenAI-compatible endpoint Ollama exposes
curl -s http://localhost:11434/api/tags | headOn the first request after start, Ollama loads the model into memory, which can take tens of seconds for large models. Subsequent requests are fast. Ensure the host has enough RAM/VRAM for the model size.
Token Budgets & Analytics Enterprise
Token budgets and usage analytics are Controller features. The standalone Terminal does not enforce organization-wide token budgets.
The Controller token budget has three knobs:
- Monthly soft limit (
monthly_soft_limit_tokens) — a warning threshold; AI keeps working past it. - Monthly hard limit (
monthly_hard_limit_tokens) — a strict cap. - Warning threshold percent (
warning_threshold_percent) — the percentage at which a notification fires.
Budget status is read at GET /api/admin/analytics/tokens/budget and updated with PUT to the same path. The status payload reports current_usage_tokens, the configured limits, usage_percent, a status string, and a projected_month_end_tokens estimate.
Usage analytics live under /api/admin/analytics/tokens/ and can be broken down several ways:
GET /api/admin/analytics/tokens/timeseries— usage over time.GET /api/admin/analytics/tokens/by-feature— chat vs. agents vs. suggestions vs. embeddings.GET /api/admin/analytics/tokens/by-user— per user.GET /api/admin/analytics/tokens/by-provider— per provider.
The same admin UI exposes a Token Analytics page in the AI navigation group. The Settings, AI section also includes an AI Data Security control for redacting credentials and sensitive data before prompts are sent to an LLM.
API & Config Examples
The JSON examples below match the Controller llm_provider model. All Controller admin endpoints are mounted under /api/admin (there is no /v1 segment).
OpenAI provider (Controller)
{
"provider_type": "openai",
"name": "OpenAI Production",
"enabled": true,
"priority": 1,
"api_key_credential_id": "a1b2c3d4-0000-0000-0000-000000000000",
"default_model": "gpt-4o",
"max_input_tokens": 128000,
"max_output_tokens": 4096,
"temperature": 0.7,
"top_p": 0.9
}Anthropic provider (Controller)
{
"provider_type": "anthropic",
"name": "Anthropic - Agent Reasoning",
"enabled": true,
"priority": 2,
"api_key_credential_id": "e5f6a7b8-0000-0000-0000-000000000000",
"default_model": "claude-sonnet-4-20250514",
"max_input_tokens": 200000,
"max_output_tokens": 8192,
"temperature": 0.3
}Ollama provider (Controller)
{
"provider_type": "ollama",
"name": "Local Ollama - Privacy Mode",
"enabled": true,
"priority": 3,
"base_url": "http://ollama.internal:11434",
"default_model": "llama3.2",
"max_output_tokens": 4096,
"temperature": 0.7
}LiteLLM gateway via the Custom type (Controller)
{
"provider_type": "custom",
"name": "LiteLLM Gateway",
"enabled": true,
"priority": 4,
"api_key_credential_id": "c0ffee00-0000-0000-0000-000000000000",
"base_url": "http://litellm.internal:4000",
"default_model": "gpt-4o"
}Terminal provider config (standalone)
The standalone Terminal stores its provider config locally as JSON, tagged by provider. The key lives in the local settings, not a vault.
{
"provider": "anthropic",
"api_key": "sk-ant-...",
"model": "claude-3-5-sonnet-20241022"
}
// Ollama (no key)
{
"provider": "ollama",
"model": "llama3.2",
"base_url": "http://localhost:11434"
}
// LiteLLM proxy (key optional)
{
"provider": "litellm",
"model": "gpt-4o",
"base_url": "http://localhost:4000"
}Create a provider (Controller admin API)
# Admin LLM provider routes are nested at /api/admin/llm-providers
curl -X POST https://controller.example.com/api/admin/llm-providers \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"provider_type": "anthropic",
"name": "Anthropic Primary",
"priority": 1,
"api_key_credential_id": "e5f6a7b8-0000-0000-0000-000000000000",
"default_model": "claude-sonnet-4-20250514"
}'
# Test a configured provider (replace :id with the returned provider id)
curl -X POST https://controller.example.com/api/admin/llm-providers/$ID/test \
-H "Authorization: Bearer $TOKEN"Multi-provider failover (Controller)
# Primary (priority 1)
curl -X POST https://controller.example.com/api/admin/llm-providers \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"provider_type":"anthropic","name":"Anthropic Primary","priority":1,
"api_key_credential_id":"...","default_model":"claude-sonnet-4-20250514"}'
# Fallback (priority 2) - tried if the primary is unreachable
curl -X POST https://controller.example.com/api/admin/llm-providers \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"provider_type":"openai","name":"OpenAI Fallback","priority":2,
"api_key_credential_id":"...","default_model":"gpt-4o"}'Token budget (Controller)
# Update the monthly budget
curl -X PUT https://controller.example.com/api/admin/analytics/tokens/budget \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{
"monthly_soft_limit_tokens": 5000000,
"monthly_hard_limit_tokens": 10000000,
"warning_threshold_percent": 80
}'
# Read current budget status
curl -s https://controller.example.com/api/admin/analytics/tokens/budget \
-H "Authorization: Bearer $TOKEN" | jq '.'
# Response is wrapped under "budget"
{
"budget": {
"current_usage_tokens": 3250000,
"soft_limit_tokens": 5000000,
"hard_limit_tokens": 10000000,
"usage_percent": 32.5,
"status": "ok",
"projected_month_end_tokens": 7800000
}
}Usage by feature (Controller)
# Token usage broken down by feature; start/end are ISO8601 query params
curl -s https://controller.example.com/api/admin/analytics/tokens/by-feature \
-H "Authorization: Bearer $TOKEN" \
-G --data-urlencode "start=2026-03-01T00:00:00Z" \
--data-urlencode "end=2026-04-01T00:00:00Z" | jq '.'
# Response is wrapped under "data"
{
"data": [
{ "feature": "chat", "total_tokens": 1850000, "total_cost_cents": 925, "request_count": 342 },
{ "feature": "agents", "total_tokens": 980000, "total_cost_cents": 490, "request_count": 45 },
{ "feature": "suggestions", "total_tokens": 320000, "total_cost_cents": 32, "request_count": 1205 },
{ "feature": "embeddings", "total_tokens": 100000, "total_cost_cents": 1, "request_count": 89 }
]
}Questions & Answers
- Which LLM providers does NetStacks support?
- The standalone Terminal supports six configurations: Anthropic, OpenAI, Ollama, OpenRouter, LiteLLM, and Custom (any OpenAI-compatible endpoint). The Controller supports five provider types —
openai,anthropic,ollama,openrouter, andcustom. A LiteLLM gateway is used in the Controller via the Custom type by pointing the base URL at the gateway. - What is BYOK?
- BYOK (Bring Your Own Key) means you supply your own provider API keys. NetStacks includes no bundled AI credits and proxies nothing through its own infrastructure. You create keys directly with your provider and enter them in NetStacks, keeping full control over cost, rate limits, and data-handling terms.
- Where are my API keys stored?
- In the standalone Terminal, the key is stored locally as part of the AI provider settings JSON on that machine. In the Controller, the key is stored encrypted in the credential vault and referenced by
api_key_credential_id; the raw key is never returned by the API, and AI calls run server-side. - Can I run a fully self-hosted LLM?
- Yes. Configure Ollama and point it at your local instance (default
http://localhost:11434); inference stays on your hardware so data never leaves your network. You can also use the Custom type to reach any OpenAI-compatible endpoint such as vLLM, text-generation-inference, or Azure OpenAI in your own subscription. - How do token budgets work?
- Token budgets are a Controller feature. You set a
monthly_soft_limit_tokens(warning only),monthly_hard_limit_tokens(strict cap), and awarning_threshold_percent. Status comes fromGET /api/admin/analytics/tokens/budgetand includes a projected month-end estimate. The standalone Terminal does not enforce org budgets. - What are the correct admin API paths?
- Controller admin routes are mounted under
/api/admin(no/v1). LLM providers are at/api/admin/llm-providersand token analytics at/api/admin/analytics/tokens/…. The budget response is wrapped underbudgetand the by-feature response underdata. - Is my data sent to external AI providers?
- When you use a cloud provider, prompts and terminal context go to that provider’s API — but only after the credential sanitizer scrubs secrets such as Cisco enable secrets, type-7/0 passwords, SNMP community strings, SNMPv3 auth, and Juniper
$9$secrets. For zero external egress, use Ollama with a local model. With a Controller, AI calls are made server-side so Terminal clients never talk to providers directly.
Troubleshooting
| Issue | Likely cause | Fix |
|---|---|---|
| Provider connection failing | Invalid key, wrong endpoint, or blocked egress | Use the Test action to diagnose. Confirm the key is current and the base URL is reachable from the machine making the call (Controller for Enterprise, the desktop for standalone). Check outbound HTTPS to the provider. |
| Self-signed cert rejected (Controller) | Internal endpoint uses a self-signed certificate | Enable accept_invalid_certs on the provider, or install the endpoint’s CA on the Controller host. |
| Token budget exceeded (Controller) | Monthly usage hit the hard limit | Check GET /api/admin/analytics/tokens/budget for usage and projection. Raise the hard limit, wait for reset, or move high-volume features (suggestions) to a cheaper model. Use the by-feature/by-user breakdowns to find top consumers. |
| Ollama not responding | Service down or model not pulled | Run ollama list to confirm the service and model. If missing, run ollama pull <model>. Make sure the base URL matches the Ollama listen address. The first request loads the model (tens of seconds). |
| Failover not triggering (Controller) | Backup disabled or priority misconfigured | Ensure backup providers are enabled with valid credentials and a higher priority number than the primary. Test each provider individually. Failover responds to connection errors and timeouts. |
Related Features
- AI Chat — conversational network-operations assistance over the configured LLM.
- AI Agents — autonomous agents that use the configured LLM for multi-step investigation.
- Command Suggestions — context-aware autocomplete and the credential sanitizer that scrubs secrets before prompts are sent.
- Knowledge Base — uses provider embedding models for vector search.
- MCP Servers — extend AI features with Model Context Protocol tools.
- System Settings — Controller configuration including the credential vault that holds encrypted provider keys.