The On-Premise GenBI Reference Architecture.
The complete blueprint for sovereign, governed, natural-language analytics run entirely inside your network: a four-page brief your infrastructure and security teams can review as-is.
On-premise by architecture
Keep every question inside your boundary.
100%
runs in your network
from context grounding and inference through query execution
$0
per-token API fees
when inference runs on locally served models you control
80%*
query hit rate
in Phison’s fully on-premise production deployment
* Up to 80%, reported by Phison’s CTO. Deployment results vary.

- 01
Why GenBI is going sovereign
The three forces (data sovereignty, AI sovereignty, and agent governance) pushing enterprise GenBI inside the firewall.
- 02
The reference architecture
One network boundary, four layers: clients, the governed context layer, local inference and data, and hardware you own.
- 03
Layer by layer
The MDL context model, row/column-level security with signed JWTs, validated local models on Ollama or vLLM, and data queried in place.
- 04
Economics & plans
Inference as a capacity plan on hardware you size, and fixed-cost licensing by concurrent sessions, never seats.
Get the reference architecture
Fill out the form and we'll take you straight to the PDF.
Free PDF · 4 pages

