Your databases
Postgres, MySQL, BigQuery, Snowflake, and 20+ sources queried in place.
Run the complete GenBI stack on infrastructure you govern. Sovereignty is built into the architecture, and the AI is a capability you own, not a service you rent.
Reference architecture / self-hosted GenBI
The context layer, the GenBI agent, and the model all deploy on your infrastructure. Residency and jurisdiction become properties of the architecture, and your security review has nothing external to chase.
Postgres, MySQL, BigQuery, Snowflake, and 20+ sources queried in place.
Grounds every question in your schema, metrics, and permissions.
Validated open-weight models served by Ollama, vLLM, or any OpenAI-compatible endpoint.
Data can't cross a border it never approaches. Residency and jurisdiction are guaranteed by design, not by a vendor contract.
Runs on networks with no internet route at all. Included with Enterprise Plus.
Permission-aware SQL, audit logs, and a constrained query surface keep every answer inside policy.
AI sovereignty means the model itself is yours to run. Serve a validated open-weight model with Ollama or vLLM on your own GPUs, point Wren AI at your base URL, and no outside vendor can reprice, deprecate, or switch off your analytics.
We benchmark open-weight models on real GenBI workloads and recommend only the ones that pass.
Schema, joins, and business definitions travel with every request, so local models write governed SQL instead of guessing.
A newer validated model is a config change: modeling, permissions, and saved questions carry over, so you're never tied to one vendor's roadmap.
# Point Wren AI at the endpoint you already run.
type: llm
provider: litellm_llm
models:
- alias: default
model: openai/validated-model # from the tested list
api_base: "http://localhost:11434/v1"
timeout: 600Every question, report, and agent run is an LLM call. On metered APIs that bill compounds with adoption; on your own hardware it's a capacity plan.
Self-hosted licensing counts simultaneous active sessions — never named users.
The bigger the rollout, the stronger the case — every dashboard, report, and agent runs on inference you already own.
Run standalone or in a Kubernetes cluster, on infrastructure you size yourself — already proving out in production.
Wren AI enables natural language data interaction across 20+ databases without ETL or migration. Integrated with Phison's aiDAPTIV+ architecture, it delivers up to an 80% query hit rate and secure, on-prem AI operations—accelerating adoption with lower costs and faster deployment.
Start free with the open-source context layer, or deploy the full self-hosted platform on Enterprise Plus — both run entirely inside your network with local models.
The open context layer for AI agents — a CLI and Rust context engine that give your own agents governed access to data. Built for developers; no UI.
The full self-hosted platform, for regulated industries and organization-wide rollouts that need a formal security review.
A concise, four-page field guide to sovereign GenBI: the network boundary, governed context layer, local inference, security controls, and fixed-cost economics behind production deployments.
Free PDF · 4 pages · Built for platform, data, and security teams



Wren AI runs against a curated set of open-weight models we test for SQL accuracy on GenBI workloads, served through any OpenAI-compatible endpoint such as Ollama or vLLM. We share the current validated list during your evaluation, and moving to a newer validated model is a configuration change, not a migration.
Both are met by architecture rather than contract. Data sovereignty: every prompt, schema detail, and answer is processed inside your network and jurisdiction; nothing reaches an external AI vendor. AI sovereignty: inference runs on validated open-weight models on hardware you control, so the capability can't be repriced, deprecated, or revoked from outside. This holds in your data center, a sovereign-cloud tenancy, or a fully air-gapped site.
Yes. The context layer, the GenBI agent, and the model all run inside your network, so the deployment works with no internet route at all. On-prem and air-gapped deployment support is included in the Enterprise Plus plan.
The quality of GenBI answers depends heavily on the model's own intelligence, so we maintain a list of tested, highly recommended local models and strongly suggest running one of them. Every request is also grounded in your context layer (schema, joins, metrics, and permissions), which narrows the model's job. Reach out to request the current list.
Two fixed parts: inference runs on hardware you provision, so the cost per question falls as usage grows, and the Wren AI license is priced by concurrent sessions rather than seats, with unlimited users. Enterprise Plus starts at 10 concurrent sessions; estimate your sessions or compare plans on the pricing page.
A concurrent session is a single LLM-driven request to Wren AI (a question, chart, report, or GenBI App), counted from request to output, whether it comes from the UI, API, Teams, Slack, or MCP. Licensing counts simultaneous active sessions, not named users; as a guide, one concurrent session supports around five active users.
Schedule a demo. We scope your deployment, share the validated local-model list, and set up a pilot in your environment on the Enterprise Plus plan, so what you evaluate is exactly what you roll out.
We'll walk your data and security teams through a self-hosted rollout with local models — or start with the reference architecture your infrastructure team can review today.