00Deployment / Sovereign AI + local models

Sovereign AI analytics, inside your network.

Run the complete GenBI stack on infrastructure you govern. Sovereignty is built into the architecture, and the AI is a capability you own, not a service you rent.

  • Standalone or Kubernetes
  • Fully air-gapped option
  • No data migration
Private deployment / online
Your network
01 / Request
Ask a business question
UI · API · MCP
02 / Ground
Wren context layer
Schema · metrics · permissions
03 / Infer
Local model endpoint
Ollama · vLLM · compatible API
04 / Query
Your data sources
Queried in place — no migration
Egress
0
external AI calls
Data path
IN-PLACE
schema and rows stay yours
Air-gap ready
No metered APIs
Unlimited users

Reference architecture / self-hosted GenBI

Runs in your network
100%
Per-token API fees
$0
Data sources, one layer
20+
Users, session-based license
Unlimited
01Data sovereignty

Everything runs inside your boundary.

The context layer, the GenBI agent, and the model all deploy on your infrastructure. Residency and jurisdiction become properties of the architecture, and your security review has nothing external to chase.

Private deployment boundary
Your network
Standalone or KubernetesAir-gap ready
Source

Your databases

Postgres, MySQL, BigQuery, Snowflake, and 20+ sources queried in place.

No ETL, no migration
GenBI

Wren AI

Grounds every question in your schema, metrics, and permissions.

Standalone or Kubernetes
Inference

Local LLM

Validated open-weight models served by Ollama, vLLM, or any OpenAI-compatible endpoint.

Your GPUs
External AI APIs
Egress blocked by design
Blocked
0outbound calls
0prompts shared
0tokens metered
01

Sovereignty by architecture

Data can't cross a border it never approaches. Residency and jurisdiction are guaranteed by design, not by a vendor contract.

02

Air-gapped ready

Runs on networks with no internet route at all. Included with Enterprise Plus.

03

Governed by default

Permission-aware SQL, audit logs, and a constrained query surface keep every answer inside policy.

02AI sovereignty

Own the models, not just the data.

AI sovereignty means the model itself is yours to run. Serve a validated open-weight model with Ollama or vLLM on your own GPUs, point Wren AI at your base URL, and no outside vendor can reprice, deprecate, or switch off your analytics.

  • 01

    A tested, validated model list

    We benchmark open-weight models on real GenBI workloads and recommend only the ones that pass.

  • 02

    The context layer does the heavy lifting

    Schema, joins, and business definitions travel with every request, so local models write governed SQL instead of guessing.

  • 03

    Swap models without re-platforming

    A newer validated model is a config change: modeling, permissions, and saved questions carry over, so you're never tied to one vendor's roadmap.

config.yaml
# Point Wren AI at the endpoint you already run.
type: llm
provider: litellm_llm
models:
  - alias: default
    model: openai/validated-model   # from the tested list
    api_base: "http://localhost:11434/v1"
    timeout: 600
Serves viaOllamavLLMOpenAI-compatible
Ask for the validated model list
03Economics

Pay for machines, not tokens.

Every question, report, and agent run is an LLM call. On metered APIs that bill compounds with adoption; on your own hardware it's a capacity plan.

GenBI on metered APIs

Cost driver
Per token — every question, retry, and agent step
As adoption grows
Spend climbs with success
Budgeting
Variable, hard to forecast
Data path
Prompts and schema leave your network

GenBI on local models

Fixed cost
Cost driver
Machines you size once, run at capacity
As adoption grows
Cost per question falls
Budgeting
A fixed line item
Data path
Nothing leaves your network
Licensing

Concurrent sessions, not seats.

Self-hosted licensing counts simultaneous active sessions — never named users.

Users
Unlimited
One session
≈ 5 active users
API & agents
Count same as UI

The bigger the rollout, the stronger the case — every dashboard, report, and agent runs on inference you already own.

04 / Scale

Built for large-scale GenBI rollouts.

Run standalone or in a Kubernetes cluster, on infrastructure you size yourself — already proving out in production.

80%
Query hit rate at Phison, fully on-prem
20+
Databases connected without ETL
100+
Hours saved monthly by one data team
Wren AI enables natural language data interaction across 20+ databases without ETL or migration. Integrated with Phison's aiDAPTIV+ architecture, it delivers up to an 80% query hit rate and secure, on-prem AI operations—accelerating adoption with lower costs and faster deployment.
Wei
CTO, Phison (A Public Company)
05Getting there

Two ways to run it in your network.

Start free with the open-source context layer, or deploy the full self-hosted platform on Enterprise Plus — both run entirely inside your network with local models.

Path 01 / Open Source

Wren AI Open Source

The open context layer for AI agents — a CLI and Rust context engine that give your own agents governed access to data. Built for developers; no UI.

  • wren CLI: query, plan, validate & profile across 20+ data sources
  • MDL context layer: models, relationships & calculated fields
  • Rust context engine powered by Apache DataFusion
  • Bring Your Own LLM
Recommended
Path 02 / Enterprise Plus

Wren AI Enterprise Plus

The full self-hosted platform, for regulated industries and organization-wide rollouts that need a formal security review.

  • Full platform: UI, dashboards, governance
  • Standalone or Kubernetes, including air-gapped environments
  • SSO and SCIM 2.0
  • Unlimited users, starts at 10 concurrent sessions
  • Procurement and security review
Reference architecture · Free PDF

Take the sovereign GenBI blueprint into your next architecture review.

A concise, four-page field guide to sovereign GenBI: the network boundary, governed context layer, local inference, security controls, and fixed-cost economics behind production deployments.

  • 01Sovereignty by architecture, not contract
  • 02Local models and governed context
  • 03Row and column security
  • 04Fixed-cost economics and licensing

Free PDF · 4 pages · Built for platform, data, and security teams

Cover of The On-Premise GenBI Reference Architecture
On-Premise GenBI whitepaper architecture page
On-Premise GenBI whitepaper security and model page
FAQ / On-premise

On-premise questions, answered.

Wren AI runs against a curated set of open-weight models we test for SQL accuracy on GenBI workloads, served through any OpenAI-compatible endpoint such as Ollama or vLLM. We share the current validated list during your evaluation, and moving to a newer validated model is a configuration change, not a migration.

Both are met by architecture rather than contract. Data sovereignty: every prompt, schema detail, and answer is processed inside your network and jurisdiction; nothing reaches an external AI vendor. AI sovereignty: inference runs on validated open-weight models on hardware you control, so the capability can't be repriced, deprecated, or revoked from outside. This holds in your data center, a sovereign-cloud tenancy, or a fully air-gapped site.

Yes. The context layer, the GenBI agent, and the model all run inside your network, so the deployment works with no internet route at all. On-prem and air-gapped deployment support is included in the Enterprise Plus plan.

The quality of GenBI answers depends heavily on the model's own intelligence, so we maintain a list of tested, highly recommended local models and strongly suggest running one of them. Every request is also grounded in your context layer (schema, joins, metrics, and permissions), which narrows the model's job. Reach out to request the current list.

Two fixed parts: inference runs on hardware you provision, so the cost per question falls as usage grows, and the Wren AI license is priced by concurrent sessions rather than seats, with unlimited users. Enterprise Plus starts at 10 concurrent sessions; estimate your sessions or compare plans on the pricing page.

A concurrent session is a single LLM-driven request to Wren AI (a question, chart, report, or GenBI App), counted from request to output, whether it comes from the UI, API, Teams, Slack, or MCP. Licensing counts simultaneous active sessions, not named users; as a guide, one concurrent session supports around five active users.

Schedule a demo. We scope your deployment, share the validated local-model list, and set up a pilot in your environment on the Enterprise Plus plan, so what you evaluate is exactly what you roll out.

Next step / architecture review

Sovereignty is an architecture decision.

We'll walk your data and security teams through a self-hosted rollout with local models — or start with the reference architecture your infrastructure team can review today.