Reference architecture
The On-Premise GenBI Reference Architecture.
The complete blueprint for sovereign, governed, natural-language analytics run entirely inside your network: a four-page brief your infrastructure and security teams can review as-is.
On-premise by architecture
Keep every question inside your boundary.
100%
runs in your network
from context grounding and inference through query execution
$0
per-token API fees
when inference runs on locally served models you control
80%*
cache hit rate
in Phison’s fully on-premise production deployment
* Up to 80% cache hit rate, reported by Phison’s CTO for its on-prem AI infrastructure. This is not an answer-accuracy metric. Deployment results vary.

Why GenBI is going sovereign
The three forces (data sovereignty, AI sovereignty, and agent governance) pushing enterprise GenBI inside the firewall.
The reference architecture
One network boundary, four layers: clients, the governed context layer, local inference and data, and hardware you own.
Layer by layer
The MDL context model, row/column-level security with signed JWTs, validated local models on Ollama or vLLM, and data queried in place.
Economics & plans
Inference as a capacity plan on hardware you size, and fixed-cost licensing by concurrent sessions, never seats.
Get the reference architecture
Fill out the form and we'll take you straight to the PDF.
Free PDF, 4 pages

