What each model actually gives you
Dedicated single-tenant. The hardware is exclusively yours, operated by the provider in their facility. You get isolation without operational burden. You do not get physical custody, and the provider's jurisdiction still applies.
Customer on-premises. The platform is installed in your data centre. You get physical custody and jurisdictional control. You take on power, cooling, hardware lifecycle, spares and staff.
Air-gapped. On-premises with no external connectivity. You get the strongest possible boundary. Every update becomes a logistics exercise, and every troubleshooting session happens without remote assistance.
The progression is not simply 'more secure'. Each step trades operational simplicity for a specific control, and the question is always whether you need that particular control.
The costs that are not on the quote
Hardware and licensing are the visible costs. They are rarely the largest.
- Staff. On-premises infrastructure needs people who can maintain it. Not a fraction of somebody's week — a genuine operational responsibility with cover for holidays and departures.
- Facilities. Modern accelerators are dense, hot and power-hungry. Many existing data halls cannot support a GPU rack without electrical and cooling work, and that work is often the long pole in the schedule.
- Lifecycle. Hardware you own depreciates whether it is busy or idle, and it does not upgrade itself when a substantially better accelerator ships in eighteen months.
- Idle capacity. You sized for peak. You pay for peak continuously. Cloud absorbs this; owned hardware does not.
- Update friction on air-gapped sites. A security patch that takes minutes elsewhere becomes a scheduled visit with signed media and a validation pass.
The honest test for on-premises is whether you would keep the hardware genuinely busy for two years. Most teams asking the question would not, and their real requirement is isolation rather than custody.
How teams over-specify
The most common pattern we see is a requirement that begins as 'our data cannot go to a foreign hyperscaler' and arrives at the procurement stage as 'we need an air-gapped installation'. Those are very different requirements, and the second one is perhaps ten times the work.
Over-specification usually happens for understandable reasons. Nobody is criticised for being too careful. The infrastructure decision is made early, by people who will not operate it, against a requirement nobody has read in the original.
The cost lands later: a system that is slower to update, harder to support, and expensive to run, for controls that were never required. The way to avoid it is to write down the constraint in one sentence with a citation. If the citation cannot be produced, the requirement deserves re-examination before it becomes a rack.
Platform parity is the property that matters
Whatever you choose, the property worth insisting on is that the platform is the same across all deployment models.
If the on-premises edition is a reduced version with different APIs, you have not chosen a deployment model — you have chosen a different product. Workloads developed on shared cloud will not lift into it, and every migration becomes a rebuild rather than a redeployment.
Parity is also what makes the sensible path viable: build and validate on shared infrastructure where iteration is cheap and fast, then move the proven workload into a private deployment once you know what it needs. Without parity, you have to commit to the boundary before you understand the workload, which is the worst possible ordering.
A sequence that works
- Establish the boundary. Regulator, contract, classification or internal policy — and get the citation.
- Prototype on shared cloud. Iteration is cheap and you learn what the workload actually needs before you buy anything.
- Measure utilisation honestly. Weeks of real data, not an estimate from a planning session.
- Choose the lightest model that satisfies the boundary. Dedicated before on-premises, on-premises before air-gapped.
- Design against measured requirements. Size the deployment from what the workload does, not from what was available in the catalogue.
- Commission with acceptance criteria. Written entry and exit criteria per phase, so 'done' is verifiable rather than a matter of opinion.
When air-gapping genuinely is the answer
It is worth being clear that sometimes it is. Classified government work, certain defence applications, and some critical national infrastructure have rules that make connectivity itself the prohibited thing. No amount of encryption or network segmentation substitutes.
In those cases, plan for the operational reality rather than treating it as an inconvenience: a defined update cadence with signed media, local staff trained to a real standard, on-site spares, and documented runbooks that assume no remote assistance is available. Air-gapped systems that are run as though they were connected systems tend to fall behind on patching, which eventually undermines the security posture the air gap was there to protect.
What to take away
- Dedicated gives isolation, on-premises gives custody, air-gapped removes connectivity — three different controls.
- Staff, facilities, lifecycle and idle capacity usually exceed the hardware cost of owning infrastructure.
- Write the constraint in one sentence with a citation; unciteable requirements are usually internal policy.
- Insist on platform parity, or migration between models becomes a rebuild.
- Prototype on shared cloud first — commit to a boundary after you understand the workload, not before.
On the platform
Deployment options on NXAARA
Private AI
Dedicated single-tenant, on-premises and air-gapped deployment with the same console, APIs and audit log across all three.
Read more →GPU Cloud
Prototype on shared or reserved capacity before committing to a private deployment.
Read more →Security
Isolation, encryption, access control, audit logging and our current certification position.
Read more →FAQ
Related questions
How long does an on-premises deployment take?
It depends far more on your procurement, power and network readiness than on the software. Dedicated single-tenant is quickest because we are not waiting on your facility. Air-gapped adds validation and logistics time. A phased schedule during scoping is more useful than a single number.
Can we run a hybrid — sensitive workloads private, everything else on shared cloud?
Yes, and it is often the most sensible arrangement. Keep the regulated workload inside the boundary and run experimentation, non-sensitive inference and development on shared capacity where iteration is cheaper.
Who operates the system after installation?
Your choice. Handover with documented runbooks and training, or a managed agreement. Air-gapped sites commonly run a hybrid: your staff day to day, provider engineers in scheduled maintenance windows.
Keep reading
More from NXAARA Insights
Fine-tuning vs RAG
Most teams reach for fine-tuning when they have a knowledge problem, and for retrieval when they have a behaviour problem. Both are expensive mistakes, and both are avoidable with one question.
Read the article →GPU cloud pricing explained
The hourly rate on a pricing page is the smallest part of an AI infrastructure bill. Here is what the rest of it consists of, and which parts you can control.
Read the article →UAE AI data residency
Residency, sovereignty and isolation get used interchangeably in procurement conversations. They are three different things, and only one of them is usually what a regulator asked for.
Read the article →Scope it before you buy hardware
Tell us what the boundary is and where it comes from. We will tell you the lightest deployment model that satisfies it.